A surface water anomaly sampling triggering method and system based on multi-source information fusion

CN122329755BActive Publication Date: 2026-09-04JIANGXI ESUN ENVIRONMENTAL PROTECTION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610796625.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-09-04
Estimated Expiration
2046-06-04

AI Technical Summary

Technical Problem

然而,这些方法主要局限于对监测数据本身的孤立分析,缺乏对环境背景、历史规律及外部因素的综合考量,导致报警逻辑相对简单,难以适应复杂多变的自然水体环境

Benefits of technology

本发明通过融合在线监测、环境工况、历史水质及外部预警四类数据,构建了全方位的异常识别体系,显著提升了地表水异常检测的准确率与响应速度。相较于传统单一阈值报警,本方案利用数字孪生模型生成动态水质基线进行残差分析,能有效剥离自然环境波动干扰,精准捕捉隐蔽及突发污染;结合图神经网络对流域拓扑与异常标记进行聚合分析,实现了污染溯源与传播路径识别,并据此自动生成靶向采样指令,解决了现有技术误报率高、响应滞后的问题,提高了监测资源的配置效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122329755B_ABST
    Figure CN122329755B_ABST
Patent Text Reader

Abstract

The application discloses a surface water abnormal sampling triggering method and system based on multi-source information fusion, and relates to the technical field of data processing.The method comprises the following steps: collecting surface water online monitoring, environmental working conditions, historical water quality and external early warning data, and constructing a multi-source associated data set; obtaining a standard data set after preprocessing; identifying data association abnormalities to generate a first abnormality mark, and combining with space-time comparison to identify implicit abnormalities to generate a second abnormality mark; constructing a digital twin model of a river basin water environment, and generating a third abnormality mark through residual analysis of real-time monitoring data and a dynamic water quality baseline; using a graph neural network to aggregate and analyze the three kinds of marks, identifying abnormal propagation paths and tracing directions, and outputting tracing results and comprehensive abnormal confidence; and generating a targeted sampling instruction to automatically trigger associated point abnormal sampling.The application fuses multi-source data and digital twinning, accurately identifies water quality abnormalities and targets tracing sampling, and improves monitoring efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a method and system for triggering surface water anomaly sampling based on multi-source information fusion. Background Technology

[0002] Traditional surface water monitoring systems primarily rely on a combination of online monitoring of routine indicators at fixed locations and periodic manual sampling. This passive regulatory model often exhibits shortcomings in response to concealed pollution or rapidly spreading sudden pollution events, including delayed response and insufficient coverage. Therefore, constructing a surface water environmental monitoring system capable of real-time anomaly detection, intelligent risk assessment, and automatic emergency response has become a key technological area urgently needing breakthroughs in the fields of smart water conservancy and ecological environmental protection.

[0003] Currently, most existing water quality anomaly monitoring technologies are based on single-dimensional threshold alarm mechanisms. Specifically, monitoring systems typically only collect online monitoring data of surface water (such as pH, dissolved oxygen, conductivity, etc.), and issue an alarm signal when a certain indicator exceeds a preset fixed upper or lower limit. Some advanced systems combine simple trend analysis, using a sliding time window to calculate the rate of change of the indicator to identify rapid water quality deterioration. However, these methods are mainly limited to the isolated analysis of the monitoring data itself, lacking a comprehensive consideration of environmental background, historical patterns, and external factors. This results in relatively simple alarm logic, making it difficult to adapt to the complex and ever-changing natural aquatic environment.

[0004] However, existing technologies suffer from a significant dual contradiction: high false alarm rates and high false negative rates. First, because they do not integrate environmental condition data (such as rainfall and hydrological flow velocity) and historical water quality baseline data, the system cannot effectively distinguish between changes in indicators caused by natural environmental fluctuations (such as increased sediment content due to heavy rain) and genuine human-caused pollution events, leading to a large number of false alarms and wasting regulatory resources. Second, traditional fixed thresholds or simple trend analysis are insufficient to identify hidden anomalies and slowly accumulating pollution processes, exhibiting extremely low sensitivity to non-point source pollution or trace, continuous leaks. Furthermore, after issuing an alarm, existing systems typically lack automated decision support and linked sampling mechanisms, failing to accurately trigger targeted sampling based on the anomaly's propagation path and source tracing direction, significantly compromising the timeliness and accuracy of emergency response. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a method and system for triggering surface water anomaly sampling based on multi-source information fusion, thereby solving the aforementioned technical problems.

[0006] The first aspect of the present invention is to provide a method for triggering surface water anomaly sampling based on multi-source information fusion, the method comprising: Collect online monitoring data, environmental condition data, historical water quality data, and external environmental early warning data of surface water, and construct corresponding multi-source correlation datasets; The multi-source associated dataset is preprocessed by sequentially performing data cleaning, spatiotemporal alignment, and standardization to obtain a noise-free, same-dimensional standard dataset. Identify the correlation anomalies between different categories of data in the standard dataset, generate a first anomaly label, and at the same time, combine spatiotemporal comparison to analyze the changing trends and deviations of water quality indicators, identify hidden anomalies, and generate a second anomaly label. A digital twin model of the watershed water environment is constructed to simulate the water quality background value under different working conditions to generate a dynamic water quality baseline. Through residual analysis of real-time monitoring data and the dynamic water quality baseline, abnormal residual features are extracted and a third abnormal label is generated. A watershed topology map containing monitoring points and water flow direction is constructed. A graph neural network algorithm is used to perform aggregation analysis on three types of anomaly markers to identify the anomaly propagation path and determine the source direction. The anomaly source tracing results and comprehensive anomaly confidence are output. Based on the comprehensive anomaly confidence level and anomaly tracing results, a targeted sampling instruction is generated to automatically trigger anomaly sampling at associated monitoring points in the water quality sampling network.

[0007] According to one aspect of the above technical solution, the steps of identifying correlation anomalies between different categories of data in the standard dataset, generating a first anomaly label, and simultaneously analyzing the changing trends and deviations of water quality indicators by combining spatiotemporal comparison to identify latent anomalies and generate a second anomaly label include: The standard dataset is subjected to logical consistency verification, and abnormal conflicts between parameters are identified based on preset association rules to generate a first abnormal label. The sliding time window algorithm is used to calculate the trend slope of water quality indicators, and the historical average and spatial neighbor data are compared differentially. When the trend or spatial deviation exceeds the dynamic threshold, a second anomaly marker is generated.

[0008] According to one aspect of the above technical solution, the step of calculating the trend slope of water quality indicators using a sliding time window algorithm and comparing the historical average with spatially adjacent data points, and generating a second anomaly marker when the trend or spatial deviation exceeds a dynamic threshold, includes: The core water quality index values ​​of the current moment and the previous N moments are extracted based on the sliding time window. The trend slope is calculated by fitting using the least squares method. If the absolute value of the trend slope exceeds the preset first dynamic threshold, it is determined that there is an abnormal trend in the time dimension. Obtain the historical average data of the current monitoring point and the real-time data of the upstream and downstream adjacent points. Calculate the difference between the current data and the historical average, as well as the spatial gradient difference with the upstream and downstream data. If any difference exceeds the corresponding second dynamic threshold, it is determined that there is a spatial dimension deviation. If an abnormal trend is determined to exist in both the time dimension and the spatial dimension, a second anomaly marker is generated.

[0009] According to one aspect of the above technical solution, a digital twin model of the watershed water environment is constructed to simulate the water quality background values ​​under different operating conditions to generate a dynamic water quality baseline. The steps include extracting abnormal residual features and generating a third anomaly marker through residual analysis of real-time monitoring data and the dynamic water quality baseline. By combining watershed topography, hydrological parameters, pollution source distribution and historical water quality data, a digital twin model of watershed water environment is constructed. By inputting different meteorological and hydrological parameters, the water quality baseline value of surface water under the corresponding conditions is simulated, and a water quality baseline that changes dynamically with the conditions is generated. Calculate the residual value between real-time monitoring data and the dynamic water quality baseline under the corresponding operating conditions, extract the fluctuation range, duration and abrupt change characteristics of the residual value, and generate a third anomaly marker when the residual value exceeds the preset residual threshold and the anomaly characteristics meet the judgment conditions.

[0010] According to one aspect of the above technical solution, the step of calculating the residual value between real-time monitoring data and the dynamic water quality baseline under the corresponding operating conditions, extracting the fluctuation amplitude, duration, and abrupt change characteristics of the residual value, and generating a third anomaly marker when the residual value exceeds a preset residual threshold and the anomaly characteristics meet the judgment conditions includes: Obtain the dynamic water quality baseline and real-time monitoring data corresponding to the current operating conditions. Classify the water quality indicators and calculate the residual values ​​between the real-time values ​​of each indicator and the corresponding baseline values ​​to obtain the residual sequence of each indicator. Feature extraction is performed on the residual sequence to calculate the maximum fluctuation range of the residual value and the duration of continuous exceedance of the preset basic threshold. At the same time, the mutation nodes and mutation range in the residual sequence are identified to form an abnormal residual feature set. Based on the preset residual threshold and the abnormal feature judgment conditions, the extracted abnormal residual features are compared with the judgment conditions. If the residual value exceeds the preset threshold and the abnormal feature meets the judgment conditions, a third abnormal label is generated.

[0011] According to one aspect of the above technical solution, the steps of constructing a watershed topology map including monitoring points and water flow direction, using a graph neural network algorithm to aggregate and analyze three types of anomaly markers, identifying anomaly propagation paths and determining the source direction, and outputting anomaly source tracing results and comprehensive anomaly confidence levels include: The three types of anomaly markers at each monitoring point are used as node features input into the watershed topology map. The graph attention network is used to calculate the anomaly association weights between the current node and its upstream and downstream neighboring nodes. The spatial distribution features of anomalies are extracted through multi-layer feature aggregation. Based on the anomaly correlation weight ranking, the upstream path with increasing anomaly signal intensity is identified, and it is determined as the source direction of the abnormal pollution source, generating anomaly source tracing results; The spatial distribution features of anomalies are input into a fully connected neural network, and the comprehensive anomaly confidence score in the 0-1 interval is output.

[0012] According to one aspect of the above technical solution, the steps of inputting three anomaly markers from each monitoring point as node features into the watershed topology map, calculating the anomaly association weights between the current node and its upstream and downstream neighboring nodes using a graph attention network, and extracting the spatial distribution features of anomalies through multi-layer feature aggregation include: The first, second, and third anomaly markers are numerically encoded to generate a multidimensional anomaly feature vector, which is then used as the initial feature input for the corresponding monitoring point nodes in the watershed topology map. By aggregating the features of neighboring nodes through the graph attention network, the abnormal association weights between the current node and its upstream and downstream neighboring nodes are calculated. The abnormal feature vectors of the neighboring nodes are then aggregated using a weighted aggregation function to generate the aggregated feature vector of the current node. The aggregated feature vector is input into a nonlinear activation function for feature transformation. The feature aggregation and transformation operations are iteratively executed through a multi-layer graph attention network, and finally the abnormal spatial distribution features that characterize the propagation law of abnormal space are output.

[0013] A second aspect of the present invention is to provide a surface water anomaly sampling triggering system based on multi-source information fusion, applied to the method described in the above-mentioned technical solution, the system comprising: The data acquisition module is used to collect online monitoring data, environmental condition data, historical water quality data, and external environmental early warning data of surface water, and to construct a corresponding multi-source associated dataset; The data processing module is used to preprocess the multi-source associated dataset, sequentially performing data cleaning, spatiotemporal alignment, and standardization to obtain a noise-free, same-dimensional standard dataset. The anomaly labeling module is used to identify the correlation anomalies between different categories of data in the standard dataset, generate the first anomaly label, and at the same time, combine spatiotemporal comparison to analyze the changing trends and deviations of water quality indicators, identify hidden anomalies, and generate the second anomaly label. The feature extraction module is used to construct a digital twin model of the watershed water environment, simulate the water quality background value under different working conditions to generate a dynamic water quality baseline, and extract abnormal residual features and generate a third abnormal label through residual analysis of real-time monitoring data and the dynamic water quality baseline. The anomaly analysis module is used to construct a watershed topology map that includes monitoring points and water flow direction. It uses a graph neural network algorithm to perform aggregate analysis on three types of anomaly markers, identify anomaly propagation paths and determine the source direction, and output anomaly source tracing results and comprehensive anomaly confidence. The instruction triggering module is used to generate targeted sampling instructions based on the comprehensive anomaly confidence level and anomaly tracing results, and automatically trigger anomaly sampling at associated monitoring points in the water quality sampling network.

[0014] A third aspect of the present invention is to provide a readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method described in the above-described technical solution.

[0015] A fourth aspect of the present invention is to provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described in the above technical solutions.

[0016] Compared with existing technologies, the surface water anomaly sampling triggering method and system based on multi-source information fusion shown in this invention has the following advantages: This invention integrates four types of data: online monitoring, environmental conditions, historical water quality, and external early warning data, constructing a comprehensive anomaly identification system that significantly improves the accuracy and response speed of surface water anomaly detection. Compared to traditional single-threshold alarms, this solution utilizes a digital twin model to generate a dynamic water quality baseline for residual analysis, effectively eliminating interference from natural environmental fluctuations and accurately capturing hidden and sudden pollution. By combining graph neural networks to aggregate and analyze watershed topology and anomaly markers, it achieves pollution source tracing and propagation path identification, and automatically generates targeted sampling instructions accordingly. This solves the problems of high false alarm rates and slow response times in existing technologies, improving the efficiency of monitoring resource allocation. Attached Figure Description

[0017] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 A flowchart illustrating the surface water anomaly sampling triggering method based on multi-source information fusion provided in this embodiment of the invention; Figure 2 The structural block diagram of the surface water anomaly sampling triggering system based on multi-source information fusion provided in the embodiments of the present invention is shown. Detailed Implementation

[0018] To make the objectives, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of the present invention will be more thorough and complete.

[0019] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0021] Example 1 Please see Figure 1 The first embodiment of the present invention provides a method for triggering surface water anomaly sampling based on multi-source information fusion, the method comprising steps S10-S60: Step S10: Collect online monitoring data, environmental condition data, historical water quality data, and external environmental early warning data of surface water, and construct a corresponding multi-source associated dataset.

[0022] Specifically, the method shown in this embodiment first collects data, specifically online monitoring data of surface water, environmental condition data, historical water quality data, and external environmental early warning data, and then constructs a multi-source associated dataset accordingly.

[0023] Among them, online monitoring data can reflect the current water quality status of surface water in real time and is the core basis for anomaly identification, covering key water quality indicators such as pH value, dissolved oxygen, ammonia nitrogen, and turbidity; while environmental operating condition data, such as hydrological flow, meteorological conditions, and discharge data from surrounding sewage outlets, are used to supplement external influencing factors of water quality changes, effectively avoiding misjudgments of anomalies caused by single water quality data; historical water quality data can provide a benchmark for normal fluctuations in water quality under different time periods and operating conditions, providing reliable support for subsequent spatiotemporal comparative analysis; and external environmental early warning data can capture potential pollution anomalies in advance, significantly improving the timeliness of anomaly identification.

[0024] Step S20: Preprocess the multi-source associated dataset by sequentially performing data cleaning, spatiotemporal alignment and standardization to obtain a noise-free, same-dimensional standard dataset. Specifically, since multi-source data comes from different monitoring devices and data sources, there are differences in acquisition frequency, units of measurement, and spatial locations. Inevitably, there are problems such as noisy data, missing values, spatiotemporal asynchrony, and inconsistent units of measurement. If the raw multi-source data is directly used for subsequent anomaly identification and analysis, it will seriously affect the accuracy and reliability of the analysis results. Therefore, this preprocessing step is a key link and important guarantee connecting data acquisition and subsequent anomaly identification.

[0025] Data cleaning primarily removes noise, outliers, and missing values ​​from the data, ensuring data integrity and accuracy through interpolation and elimination. Spatiotemporal alignment addresses the asynchrony between different data sources in terms of temporal acquisition frequency and spatial monitoring locations, achieving precise matching of various data in time and space dimensions, facilitating subsequent correlation analysis. Standardization eliminates dimensional and numerical range differences between different water quality indicators and operating conditions, ensuring data comparability and avoiding analytical biases caused by dimensional issues. The synergistic effect of these three processes provides a high-quality data foundation for subsequent anomaly identification and other tasks.

[0026] Step S30: Identify the correlation anomalies between different categories of data in the standard dataset, generate a first anomaly label, and at the same time, combine spatiotemporal comparison to analyze the changing trends and deviations of water quality indicators, identify hidden anomalies, and generate a second anomaly label.

[0027] Specifically, identifying anomalies in the correlation between different categories of data in the standard dataset mainly targets logical conflicts between different categories of data, such as a sudden increase in rainfall without a corresponding increase in turbidity, or an increase in sewage discharge without a significant change in pollution indicators. A first anomaly marker is generated through logical verification, primarily used to capture readily observable explicit anomalies. Latent anomaly identification, on the other hand, uses spatiotemporal comparative analysis to delve into the temporal trends and spatial distribution deviations of water quality indicators, capturing early, difficult-to-detect latent anomalies, such as a slow, continuous increase in water quality indicators, or significant differences between data from a monitoring point and surrounding points. This generates a second anomaly marker. These two types of anomaly markers complement and verify each other, forming a multi-dimensional anomaly identification system that ensures the comprehensiveness and accuracy of anomaly identification.

[0028] In this embodiment, the steps of identifying abnormal correlations between different categories of data in the standard dataset, generating a first anomaly label, and simultaneously analyzing the changing trends and deviations of water quality indicators by combining spatiotemporal comparison to identify latent anomalies and generate a second anomaly label include: The standard dataset is subjected to logical consistency verification, and abnormal conflicts between parameters are identified based on preset association rules to generate a first abnormal label. The sliding time window algorithm is used to calculate the trend slope of water quality indicators, and the historical average and spatial neighbor data are compared differentially. When the trend or spatial deviation exceeds the dynamic threshold, a second anomaly marker is generated.

[0029] Specifically, after obtaining the standard dataset, the first step is to perform a logical consistency check on the standard dataset, identify abnormal conflicts between parameters based on preset association rules, and then generate the first abnormal label.

[0030] The core of logical consistency verification is to check whether the logical relationship between different categories of data is reasonable and whether there are any contradictions through preset logical rules, thereby eliminating abnormal misjudgments caused by data logic errors. The preset association rules are set in advance based on the objective laws of surface water quality changes, historical monitoring data statistics, and industry professional experience. They cover the inherent relationships between different data categories, such as the positive correlation between rainfall and turbidity, and the positive correlation between sewage flow and ammonia nitrogen concentration. By comparing each type of data in the standard dataset with the preset association rules one by one, if the relationship between the data is found to violate the preset rules, that is, if there is an abnormal conflict, it is judged as an association anomaly, and the first anomaly label is generated.

[0031] Then, a sliding time window algorithm is used to calculate the trend slope of water quality indicators. This is combined with a differential comparison of historical averages and data from spatially adjacent points. When the trend or spatial deviation exceeds a preset dynamic threshold, a second anomaly marker is generated. The sliding time window algorithm effectively captures the temporal trends of water quality indicators. By setting a reasonable window size—specifically determined based on monitoring frequency and actual needs—it extracts the core water quality indicator values ​​for the current and historical moments, calculates the trend slope, and clearly reflects the rate and direction of change of water quality indicators. Historical average comparison is used to analyze the deviation between current monitoring data and historical data from the same period, determining whether the current data is within the normal historical fluctuation range. Spatially adjacent point data comparison is used to analyze the differences between current monitoring point data and surrounding adjacent monitoring point data, capturing abnormal changes in local areas. The dynamic threshold is pre-set according to different water quality indicators and operating conditions, adapting to the characteristics of water quality fluctuations in different scenarios and avoiding the problem of fixed thresholds being unable to adapt to changes in operating conditions. When the trend slope or spatial deviation of a water quality indicator exceeds the corresponding dynamic threshold, a second anomaly marker is generated.

[0032] In this embodiment, the step of calculating the trend slope of water quality indicators using a sliding time window algorithm and comparing the historical average with spatially adjacent data points, and generating a second anomaly marker when the trend or spatial deviation exceeds a dynamic threshold, includes: The core water quality index values ​​of the current moment and the previous N moments are extracted based on the sliding time window. The trend slope is calculated by fitting using the least squares method. If the absolute value of the trend slope exceeds the preset first dynamic threshold, it is determined that there is an abnormal trend in the time dimension. Obtain the historical average data of the current monitoring point and the real-time data of the upstream and downstream adjacent points. Calculate the difference between the current data and the historical average, as well as the spatial gradient difference with the upstream and downstream data. If any difference exceeds the corresponding second dynamic threshold, it is determined that there is a spatial dimension deviation. If an abnormal trend is determined to exist in both the time dimension and the spatial dimension, a second anomaly marker is generated.

[0033] More specifically, in this embodiment, the core water quality index values ​​for the current moment and the previous N moments are extracted based on a sliding time window. The trend slope is calculated using the least squares method. If the absolute value of the trend slope exceeds a preset first dynamic threshold, an abnormal trend in the time dimension is determined. The size N of the sliding time window can be flexibly set according to actual monitoring needs, typically 6-12 collection cycles, i.e., 1-2 hours, to ensure effective capture of short-term temporal trends in water quality indicators. The first dynamic threshold is pre-set based on the normal rate of change of different water quality indicators. The first dynamic threshold varies for different water quality indicators; for example, the first dynamic threshold for ammonia nitrogen can be set to 0.05 mg / (L·h). When the absolute value of the trend slope exceeds this threshold, it indicates that the rate of change of the water quality indicator exceeds the normal range, indicating an abnormal trend in the time dimension, such as a continuous and rapid increase or decrease in the water quality indicator.

[0034] Then, the historical average data of the current monitoring point and the real-time data of adjacent upstream and downstream points are obtained. The difference between the current data and the historical average, as well as the spatial gradient difference with the upstream and downstream data, are calculated. If any difference exceeds the corresponding second dynamic threshold, it is determined that there is a spatial dimensional deviation. Among them, the historical average data refers to the average water quality index of the same monitoring point at the same time in recent years, which can truly reflect the historical normal water quality level of the point at that time and provide a reliable benchmark for the comparison of the current data; the real-time data of adjacent upstream and downstream points can reflect the water quality differences of different points in the same watershed, which is convenient for capturing abnormal changes in local areas; the spatial gradient difference refers to the difference between the current monitoring point data and the data of adjacent upstream and downstream points, which can clearly reflect the spatial distribution differences and change gradients of water quality; the second dynamic threshold is set in advance according to the normal fluctuation range of different water quality indicators, and is used to judge whether the deviation of the current data from the historical average and the data of surrounding points exceeds the reasonable range. If any difference exceeds the corresponding second dynamic threshold, it indicates that the current data has an abnormal spatial dimensional deviation, such as the current point data being significantly higher than the historical data of the same period or the upstream point data.

[0035] Finally, if both an abnormal trend in the time dimension and a deviation in the spatial dimension are determined, a second anomaly marker is generated. For example, if the slope of a certain water quality indicator exceeds the first dynamic threshold (an abnormal trend in the time dimension exists), but the difference between it and the historical average and the data from surrounding points does not exceed the second dynamic threshold (no deviation in the spatial dimension), then the anomaly is likely a normal temporal fluctuation, and no second anomaly marker is generated. Only when both the abnormal trend in the time dimension and the deviation in the spatial dimension are satisfied is it determined to be a latent anomaly, and a second anomaly marker is generated, ensuring the accuracy of latent anomaly identification.

[0036] Step S40: Construct a digital twin model of the watershed water environment, simulate the water quality background value under different working conditions to generate a dynamic water quality baseline, and extract abnormal residual features and generate a third abnormal marker by residual analysis of real-time monitoring data and the dynamic water quality baseline. It should be noted that the watershed water environment digital twin model is a simulation model built based on the actual situation of the watershed, which can realistically restore the topography, hydrological characteristics, and pollutant migration and transformation patterns of the watershed.

[0037] Specifically, in this embodiment, by inputting different meteorological and hydrological parameters such as rainstorms, sunny days, flood season, and dry season, the above-mentioned watershed water environment digital twin model can accurately simulate the water quality baseline value under the corresponding conditions, generate a water quality baseline that can be dynamically adjusted according to the conditions, and ensure that the baseline is accurately matched with the actual water quality.

[0038] Based on the above, by calculating the residual value between real-time monitoring data and the dynamic water quality baseline under the corresponding operating conditions, abnormal information such as the fluctuation range, duration, and abrupt change characteristics of the residual value is extracted to generate a third anomaly marker, which further supplements the dimensions of water quality anomaly identification. By complementing the first and second anomaly markers, the accuracy of water quality anomaly identification is improved.

[0039] Step S50: Construct a watershed topology map containing monitoring points and water flow direction; use graph neural network algorithm to perform aggregation analysis on three types of anomaly markers; identify anomaly propagation paths and determine source tracing directions; and output anomaly source tracing results and comprehensive anomaly confidence level.

[0040] It should be noted that the constructed watershed topology map needs to clearly show the spatial distribution of all monitoring points, the monitoring range, and the direction of water flow within the watershed, providing a clear spatial framework for identifying abnormal propagation paths.

[0041] In this embodiment, the graph neural network algorithm, preferably the graph attention network, can accurately capture the spatial correlation and anomaly propagation patterns between different monitoring points. Through deep aggregation analysis of three types of anomaly markers, it can effectively identify the propagation path of the anomaly. At the same time, the algorithm calculates and outputs the comprehensive anomaly confidence level in the 0-1 interval, quantifies the severity and probability of the anomaly, and provides data support for the formulation of subsequent sampling strategies.

[0042] Step S60: Generate a targeted sampling instruction based on the comprehensive anomaly confidence level and anomaly tracing results, and automatically trigger anomaly sampling at associated monitoring points in the water quality sampling network.

[0043] Specifically, by combining the overall anomaly confidence level (i.e., quantifying the severity of the anomaly) with the anomaly tracing results, the source and spread of the anomaly can be clearly identified. This allows for the development of differentiated sampling strategies, specifying core parameters such as sampling locations, sampling frequency, sampling quantity, and sampling time. The system automatically triggers anomaly sampling at corresponding monitoring points in the water quality sampling network, achieving closed-loop management of anomaly identification, source tracing analysis, and targeted sampling. This not only provides accurate water sample data support for subsequent pollution analysis and pollution control but also effectively saves human, material, and financial resources, avoiding resource waste caused by blind sampling.

[0044] Compared with existing technologies, the surface water anomaly sampling triggering method based on multi-source information fusion shown in this embodiment has the following advantages: This embodiment integrates four types of data: online monitoring, environmental conditions, historical water quality, and external early warning data, constructing a comprehensive anomaly identification system that significantly improves the accuracy and response speed of surface water anomaly detection. Compared to traditional single-threshold alarms, this solution utilizes a digital twin model to generate a dynamic water quality baseline for residual analysis, effectively eliminating interference from natural environmental fluctuations and accurately capturing hidden and sudden pollution. Combined with graph neural networks for aggregated analysis of watershed topology and anomaly markers, it achieves pollution source tracing and propagation path identification, and automatically generates targeted sampling instructions accordingly. This solves the problems of high false alarm rates and slow response times in existing technologies, optimizing the allocation efficiency of monitoring resources.

[0045] Example 2 The second embodiment of the present invention also provides a surface water anomaly sampling triggering method based on multi-source information fusion. The method shown in this embodiment is basically the same as the method shown in the first embodiment, except that: In this embodiment, the steps of constructing a digital twin model of the watershed water environment, simulating water quality background values ​​under different operating conditions to generate a dynamic water quality baseline, and extracting abnormal residual features and generating a third anomaly marker through residual analysis of real-time monitoring data and the dynamic water quality baseline include: By combining watershed topography, hydrological parameters, pollution source distribution and historical water quality data, a digital twin model of watershed water environment is constructed. By inputting different meteorological and hydrological parameters, the water quality baseline value of surface water under the corresponding conditions is simulated, and a water quality baseline that changes dynamically with the conditions is generated. Calculate the residual value between real-time monitoring data and the dynamic water quality baseline under the corresponding operating conditions, extract the fluctuation range, duration and abrupt change characteristics of the residual value, and generate a third anomaly marker when the residual value exceeds the preset residual threshold and the anomaly characteristics meet the judgment conditions.

[0046] Specifically, by combining watershed topography, hydrological parameters, pollution source distribution, and historical water quality data, a digital twin model of the watershed's water environment, also known as a simulation model, is constructed. By inputting different meteorological and hydrological parameters, the model simulates the baseline water quality of surface water under corresponding conditions, generating a water quality baseline that dynamically changes with the conditions. Watershed topographic data (such as watershed area, slope, and river course), hydrological parameters (such as flow rate, velocity, and water level), pollution source distribution data (such as the location of sewage outlets, discharge volume, and pollutant type), and historical water quality data are the foundational data for constructing the digital twin model of the watershed's water environment. This ensures that the model can realistically and accurately reproduce the actual situation of the watershed's water environment and the patterns of water quality changes. By inputting different meteorological and hydrological parameters such as heavy rain, sunny days, flood season, and dry season, the model can accurately simulate the baseline water quality under the corresponding conditions, generating a water quality baseline that can be dynamically adjusted according to the conditions, ensuring that the baseline accurately matches the actual conditions.

[0047] Then, the residual value between the real-time monitoring data and the dynamic water quality baseline under the corresponding operating conditions is calculated. The fluctuation amplitude, duration, and abrupt change characteristics of the residual value are extracted. When the residual value exceeds the preset residual threshold and the abnormal characteristics meet the judgment conditions, a third anomaly marker is generated. Among them, the residual value is the difference between the real-time monitoring data and the dynamic water quality baseline under the corresponding operating conditions, which can directly and intuitively reflect the degree of deviation between the real-time water quality and the normal background value. The larger the residual value, the more serious the deviation. The fluctuation amplitude of the residual value reflects the instability of the deviation. The larger the fluctuation amplitude, the more drastic the fluctuation of the water quality from the normal level. The duration reflects the persistence of the deviation. The longer the duration, the more stable the abnormal situation is, and it is not a random fluctuation. The abrupt change characteristics reflect the suddenness of the deviation. The more obvious the abrupt change, the more likely a sudden pollution event may have occurred. These three characteristics together constitute the abnormal residual characteristics, which can comprehensively and objectively reflect the deviation and degree of abnormality between the real-time water quality and the dynamic baseline.

[0048] The preset residual threshold is set in advance based on the normal fluctuation range of different water quality indicators to determine whether the residual value exceeds the reasonable range. The abnormal feature judgment condition is set in advance based on the fluctuation range, duration and abrupt change characteristics of the residual value. Only when the residual value exceeds the preset residual threshold and the abnormal feature meets the judgment condition can the third abnormal label be generated, which effectively avoids misjudging normal minor deviations as abnormal and ensures the accuracy of the third abnormal label generation.

[0049] In this embodiment, the steps of calculating the residual value between real-time monitoring data and the dynamic water quality baseline under corresponding operating conditions, extracting the fluctuation amplitude, duration, and abrupt change characteristics of the residual value, and generating a third anomaly marker when the residual value exceeds a preset residual threshold and the anomaly characteristics meet the judgment conditions include: Obtain the dynamic water quality baseline and real-time monitoring data corresponding to the current operating conditions. Classify the water quality indicators and calculate the residual values ​​between the real-time values ​​of each indicator and the corresponding baseline values ​​to obtain the residual sequence of each indicator. Feature extraction is performed on the residual sequence to calculate the maximum fluctuation range of the residual value and the duration of continuous exceedance of the preset basic threshold. At the same time, the mutation nodes and mutation range in the residual sequence are identified to form an abnormal residual feature set. Based on the preset residual threshold and the abnormal feature judgment conditions, the extracted abnormal residual features are compared with the judgment conditions. If the residual value exceeds the preset threshold and the abnormal feature meets the judgment conditions, a third abnormal label is generated.

[0050] Specifically, the process involves acquiring the dynamic water quality baseline and concurrent real-time monitoring data corresponding to the current operating conditions. Water quality indicators are categorized, and the residual values ​​between the real-time values ​​and the corresponding baseline values ​​for each indicator are calculated one by one to obtain the residual sequence for each indicator. First, it is crucial to accurately match the current operating conditions with the dynamic water quality baseline to avoid calculation errors caused by incorrect matching of operating conditions—for example, if the current operating condition is "heavy rain + flood season," a digital twin model must be used to simulate the dynamic water quality baseline under this condition, rather than the baseline for other operating conditions. Second, water quality indicators are categorized, such as pH, dissolved oxygen, ammonia nitrogen, and turbidity. Real-time monitoring data is matched one-to-one with the baseline values ​​of the corresponding indicators, and the residual values ​​are calculated one by one to avoid calculation errors caused by confusion between different indicators. Finally, the residual values ​​of the same indicator at different times are arranged in chronological order to form a residual sequence, clearly presenting the temporal variation pattern of the indicator's residual values.

[0051] For example, if the real-time monitoring values ​​of ammonia nitrogen are 0.8 mg / L, 0.9 mg / L, and 1.0 mg / L, and the corresponding dynamic baseline values ​​under the working conditions are 0.6 mg / L, 0.6 mg / L, and 0.6 mg / L, then the calculated residual sequence is [0.2, 0.3, 0.4], which clearly shows the time-series change trend of the residual values ​​of ammonia nitrogen.

[0052] Then, feature extraction is performed on the residual sequence to calculate the maximum fluctuation amplitude of the residual values, the duration of continuous exceedance of a preset baseline threshold, and the abrupt change nodes and amplitudes in the residual sequence, forming an abnormal residual feature set. The maximum fluctuation amplitude refers to the difference between the maximum and minimum values ​​in the residual sequence, reflecting the instability of the residual values—the larger the fluctuation amplitude, the more unstable the deviation between real-time water quality and the dynamic baseline, and the higher the probability of anomaly. The duration of continuous exceedance of the preset baseline threshold refers to the time during which the residual values ​​are continuously greater than or less than the preset baseline threshold (the preset baseline threshold is the upper limit of normal fluctuation of the residual values ​​of each indicator, set according to historical residual data statistics). The longer the duration, the longer the water quality has deviated from the normal level, and the higher the probability of anomaly. Abrupt change nodes refer to the moments when the residual values ​​in the residual sequence suddenly change significantly. The abrupt change amplitude refers to the difference between the residual values ​​before and after the abrupt change node. The larger the abrupt change amplitude and the more sudden the change, the more likely a sudden pollution anomaly has occurred in the water quality, requiring close attention. By extracting these four types of abnormal residual features, a complete abnormal residual feature set is formed, which can comprehensively and objectively reflect the degree of anomaly and the pattern of change in the residual values.

[0053] Finally, based on the preset residual threshold and anomaly feature judgment conditions, the extracted abnormal residual features are compared with the judgment conditions. If the residual value exceeds the preset threshold and the abnormal features meet the judgment conditions, a third anomaly label is generated. The preset residual threshold is a more stringent judgment standard than the preset base threshold, used to determine whether the residual value exceeds a reasonable range. For example, the preset residual threshold for ammonia nitrogen can be set to 0.8 mg / L. The anomaly feature judgment conditions are pre-set based on the set of abnormal residual features, such as "maximum fluctuation amplitude ≥ 0.8 mg / L, duration of continuous exceedance of base threshold ≥ 2 collection cycles, and sudden change amplitude ≥ 0.5 mg / L". A third anomaly label is only generated when the residual value exceeds the preset residual threshold and all abnormal residual features meet the judgment conditions.

[0054] For example, if the residual values ​​(1.2 mg / L, 1.3 mg / L) of the above ammonia nitrogen residual sequence exceed the preset residual threshold of 0.8 mg / L, and the maximum fluctuation amplitude of 1.1 mg / L, the duration of 3 collection cycles, and the mutation amplitude of 0.8 mg / L all meet the judgment conditions, then a third anomaly marker is generated; if the residual value exceeds the preset threshold, but the abnormal characteristics do not meet the judgment conditions (such as the mutation amplitude of only 0.3 mg / L), then a third anomaly marker is not generated to avoid misjudgment of anomalies.

[0055] Preferably, in this embodiment, the steps of constructing a watershed topology map including monitoring points and water flow direction, using a graph neural network algorithm to aggregate and analyze three types of anomaly markers, identifying anomaly propagation paths and determining source tracing directions, and outputting anomaly source tracing results and comprehensive anomaly confidence levels include: The three types of anomaly markers at each monitoring point are used as node features input into the watershed topology map. The graph attention network is used to calculate the anomaly association weights between the current node and its upstream and downstream neighboring nodes. The spatial distribution features of anomalies are extracted through multi-layer feature aggregation. Based on the anomaly correlation weight ranking, the upstream path with increasing anomaly signal intensity is identified, and it is determined as the source direction of the abnormal pollution source, generating anomaly source tracing results; The spatial distribution features of anomalies are input into a fully connected neural network, and the comprehensive anomaly confidence score in the 0-1 interval is output.

[0056] Specifically, the three types of anomaly markers at each monitoring point are used as node features input to the watershed topology map. The graph attention network is used to calculate the anomaly association weights between the current node and its upstream and downstream neighboring nodes. The spatial distribution features of anomalies are extracted through multi-layer feature aggregation. The watershed topology map clearly defines the spatial location, monitoring range, and water flow direction of each monitoring point, providing a clear spatial framework for anomaly propagation analysis. Each monitoring point serves as a node in the map, with the first, second, and third anomaly markers for that point used as node feature inputs to ensure that node features comprehensively and accurately reflect the anomaly situation at that point. Graph Attention Network (GAT) is used to calculate association weights, dynamically allocating association weights based on the similarity of node features, spatial distance, and water flow direction, avoiding biases in association analysis caused by fixed weights. Specifically, the higher the similarity of the current node's anomaly features with its upstream and downstream neighbors, the closer the spatial distance, and the more direct the water flow direction, the higher the association weight, and vice versa. Through multi-layer feature aggregation, the anomaly feature vectors of neighboring nodes are combined with the association weights for weighted aggregation to generate the aggregated feature vector of the current node. Through multi-layer iteration, the final output is the spatial distribution characteristics of anomalies that characterize the spatial propagation pattern of anomalies, such as the distribution range, propagation direction, anomaly intensity gradient, and anomaly concentration areas within the watershed.

[0057] Then, based on the anomaly correlation weight ranking, the upstream path of increasing anomaly signal intensity is identified and determined as the source direction of the abnormal pollution source, generating anomaly source tracing results. Since the flow of surface water has a clear directionality (from upstream to downstream), pollution anomalies usually spread gradually from the upstream point of the pollution source to the downstream point along the water flow direction. Therefore, the anomaly signal intensity will show the characteristic of "high upstream and low downstream", that is, the path of increasing anomaly signal intensity is the upstream path.

[0058] Specifically, firstly, the anomaly correlation weights of each monitoring point are sorted to identify the path where the intensity of the abnormal signal gradually increases from downstream to upstream. Secondly, the upstream starting point of this increasing path is determined as the source direction of the abnormal pollution source. Combined with the pollution source distribution data in the watershed topology map, such as the location of sewage outlets and the type of pollution source, the source tracing range is further narrowed. Finally, a complete anomaly source tracing result is generated, clarifying the source tracing direction (such as the upstream tributary area of ​​a watershed), the abnormal propagation path (such as monitoring point A→B→C), the possible pollution source types (such as industrial pollution sources and agricultural non-point source pollution), and the abnormal propagation range, providing accurate directional guidance for subsequent targeted sampling.

[0059] Finally, the spatial distribution features of the anomaly are input into a fully connected neural network, which outputs a comprehensive anomaly confidence score in the 0-1 range. The spatial distribution features include information such as the anomaly's distribution range, propagation direction, intensity gradient, and concentrated areas, comprehensively and objectively reflecting the overall situation and severity of the anomaly. By converting the spatial distribution features into quantitative values ​​in the 0-1 range—the closer the comprehensive anomaly confidence score is to 1, the higher the probability of an anomaly in the water body and the more severe the anomaly, requiring emergency sampling and response; the closer it is to 0, the lower the probability of an anomaly in the water body and the less severe the anomaly, requiring conventional encrypted sampling.

[0060] In this embodiment, the three anomaly markers of each monitoring point are used as node features input to the watershed topology map. A graph attention network is used to calculate the anomaly association weights between the current node and its upstream and downstream neighbors. The steps of extracting the spatial distribution features of anomalies through multi-layer feature aggregation include: The first, second, and third anomaly markers are numerically encoded to generate a multidimensional anomaly feature vector, which is then used as the initial feature input for the corresponding monitoring point nodes in the watershed topology map. By aggregating the features of neighboring nodes through the graph attention network, the abnormal association weights between the current node and its upstream and downstream neighboring nodes are calculated. The abnormal feature vectors of the neighboring nodes are then aggregated using a weighted aggregation function to generate the aggregated feature vector of the current node. The aggregated feature vector is input into a nonlinear activation function for feature transformation. The feature aggregation and transformation operations are iteratively executed through a multi-layer graph attention network, and finally the abnormal spatial distribution features that characterize the propagation law of abnormal space are output.

[0061] Specifically, the first, second, and third anomaly markers are numerically encoded to generate a multidimensional anomaly feature vector, which is then used as the initial feature input for the corresponding monitoring point nodes in the watershed topology map.

[0062] Specifically, the three anomaly markers are first numerically encoded. Since all three are binary variables, they can be directly used as feature dimensions to generate a three-dimensional anomaly feature vector. For example, if the first anomaly marker of a monitoring point is 1, the second anomaly marker is 1, and the third anomaly marker is 1, then the numerically encoded feature vector is [1,1,1]; if the first anomaly marker is 1, the second anomaly marker is 0, and the third anomaly marker is 1, then the feature vector is [1,0,1]. Simultaneously, to further enhance the richness and accuracy of the features, auxiliary information corresponding to each anomaly marker (such as the association rule violation type corresponding to the first anomaly marker, the trend slope corresponding to the second anomaly marker, and the residual value corresponding to the third anomaly marker) can be integrated into the feature vector to generate a higher-dimensional anomaly feature vector. This ensures that the node features can comprehensively and accurately reflect the anomaly situation at the point, adapting to the computational needs of the graph attention network. Finally, the generated multidimensional anomaly feature vector is input into the node of the monitoring point in the watershed topology map as the initial feature of the node.

[0063] Then, by aggregating the features of neighboring nodes through the graph attention network, the abnormal association weights between the current node and its upstream and downstream neighboring nodes are calculated. The abnormal feature vectors of the neighboring nodes are then aggregated using a weighted aggregation function to generate the aggregated feature vector of the current node.

[0064] Specifically, firstly, the graph attention network calculates the anomaly association weights between the current node and each of its upstream and downstream neighboring nodes through an attention mechanism. These association weights are calculated based on three core factors: node feature similarity, spatial distance, and water flow direction. Higher node feature similarity (e.g., closer anomaly feature vectors between two nodes), closer spatial distance, and more direct water flow direction (e.g., greater influence of an upstream node on a downstream node) results in higher association weights; conversely, lower values ​​lead to lower weights. The association weights are normalized using a softmax function to ensure the sum of the association weights of all neighboring nodes is 1, guaranteeing the rationality of weight allocation. Secondly, a weighted aggregation function (e.g., a weighted summation function) is used to multiply the anomaly feature vector of each neighboring node by its corresponding association weight, and then all results are summed to generate the aggregated feature vector of the current node. For example, if the neighbors of the current node A are B (upstream) and C (downstream), and B's feature vector is [1,1,0] with a correlation weight of 0.7, and C's feature vector is [0,1,1] with a correlation weight of 0.3, then A's aggregated feature vector = ([1,1,0]×0.7) + ([0,1,1]×0.3) = [0.7,0.7+0.3,0.3] = [0.7,1.0,0.3]. Through this weighted aggregation, the aggregated feature vector of the current node can comprehensively reflect its own abnormal situation and the abnormal influence of neighboring nodes, accurately capturing the spatial correlation of anomalies.

[0065] Finally, the aggregated feature vector is input into a nonlinear activation function for feature transformation. The feature aggregation and transformation operations are iteratively executed through a multi-layer graph attention network, ultimately outputting anomaly spatial distribution features that characterize the spatial propagation patterns of anomalies. The role of the nonlinear activation function (such as the ReLU function) is to introduce nonlinear factors, enhance the expressive power of the aggregated feature vector, avoid insufficient feature representation caused by linear transformation, and better capture the complex relationships between anomalies. For example, through a three-layer graph attention network iteration, the first layer aggregates basic anomaly features from neighboring nodes, the second layer aggregates anomaly features from a wider range of nodes, and the third layer mines the spatial propagation gradient and distribution patterns of anomalies. The final output includes anomaly spatial distribution features containing information such as the anomaly distribution range, propagation direction, anomaly intensity gradient, and potential pollution source location, providing feature support for subsequent anomaly propagation path identification and source tracing direction determination.

[0066] Example 3 Please see Figure 2 The third embodiment of the present invention provides a surface water anomaly sampling triggering system based on multi-source information fusion, applied to the method described in any of the above embodiments, the system comprising: The data acquisition module 10 is used to collect online monitoring data, environmental condition data, historical water quality data and external environmental early warning data of surface water, and to construct a multi-source associated dataset accordingly. Data processing module 20 is used to preprocess the multi-source associated dataset, sequentially completing data cleaning, spatiotemporal alignment and standardization processing to obtain a noise-free, same-dimensional standard dataset; The anomaly labeling module 30 is used to identify the correlation anomalies between different categories of data in the standard dataset, generate a first anomaly label, and at the same time, combine spatiotemporal comparison to analyze the changing trends and deviations of water quality indicators, identify hidden anomalies, and generate a second anomaly label. The feature extraction module 40 is used to construct a digital twin model of the watershed water environment, simulate the water quality background value under different working conditions to generate a dynamic water quality baseline, and extract abnormal residual features and generate a third abnormal label through residual analysis of real-time monitoring data and the dynamic water quality baseline. The anomaly analysis module 50 is used to construct a watershed topology map that includes monitoring points and water flow direction. It uses a graph neural network algorithm to perform aggregate analysis on three types of anomaly markers, identify the anomaly propagation path and determine the source direction, and output the anomaly source tracing results and comprehensive anomaly confidence. The instruction triggering module 60 is used to generate a targeted sampling instruction based on the comprehensive anomaly confidence level and anomaly tracing results, and automatically trigger anomaly sampling at associated monitoring points in the water quality sampling network.

[0067] Compared with existing technologies, the surface water anomaly sampling triggering system based on multi-source information fusion shown in this embodiment has the following advantages: This embodiment integrates four types of data: online monitoring, environmental conditions, historical water quality, and external early warning data, constructing a comprehensive anomaly identification system that significantly improves the accuracy and response speed of surface water anomaly detection. Compared to traditional single-threshold alarms, this solution utilizes a digital twin model to generate a dynamic water quality baseline for residual analysis, effectively eliminating interference from natural environmental fluctuations and accurately capturing hidden and sudden pollution. Combined with graph neural networks for aggregated analysis of watershed topology and anomaly markers, it achieves pollution source tracing and propagation path identification, and automatically generates targeted sampling instructions accordingly. This solves the problems of high false alarm rates and slow response times in existing technologies, optimizing the allocation efficiency of monitoring resources.

[0068] Example 4 A fourth embodiment of the present invention provides a readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method described in any of the above embodiments.

[0069] Example 5 A fifth embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described in any of the above embodiments.

[0070] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0071] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this patent should be determined by the appended claims.

Claims

1. A surface water anomaly sampling triggering method based on multi-source information fusion, characterized in that, The method includes: Collect online monitoring data, environmental condition data, historical water quality data, and external environmental early warning data of surface water, and construct corresponding multi-source correlation datasets; The multi-source associated dataset is preprocessed by sequentially performing data cleaning, spatiotemporal alignment, and standardization to obtain a noise-free, same-dimensional standard dataset. The process involves identifying abnormal correlations between different categories of data in the standard dataset, generating a first anomaly marker, and simultaneously analyzing the changing trends and deviations of water quality indicators through spatiotemporal comparison to identify latent anomalies and generate a second anomaly marker. Specifically, this includes: performing logical consistency checks on the standard dataset, identifying abnormal conflicts between parameters based on preset association rules, and generating a first anomaly marker; calculating the trend slope of water quality indicators using a sliding time window algorithm, and performing a differential comparison with historical averages and spatially adjacent point data to generate a second anomaly marker. A digital twin model of the watershed water environment is constructed to simulate the water quality background value under different working conditions to generate a dynamic water quality baseline. Through residual analysis of real-time monitoring data and the dynamic water quality baseline, abnormal residual features are extracted and a third abnormal label is generated. A watershed topology map containing monitoring points and water flow direction is constructed. A graph neural network algorithm is used to perform aggregation analysis on three types of anomaly markers to identify the anomaly propagation path and determine the source direction. The anomaly source tracing results and comprehensive anomaly confidence are output. Based on the comprehensive anomaly confidence level and anomaly tracing results, a targeted sampling instruction is generated to automatically trigger anomaly sampling at associated monitoring points in the water quality sampling network; The step of calculating the trend slope of water quality indicators using a sliding time window algorithm and generating a second anomaly marker by comparing the historical average with data from spatially adjacent points includes: extracting the core water quality indicator values ​​for the current moment and the previous N moments based on the sliding time window; calculating the trend slope using the least squares method; if the absolute value of the trend slope exceeds a preset first dynamic threshold, an abnormal trend in the time dimension is determined; acquiring the historical average data of the current monitoring point and the real-time data of upstream and downstream adjacent points; calculating the difference between the current data and the historical average, and the spatial gradient difference with the upstream and downstream data, respectively; if any difference exceeds the corresponding second dynamic threshold, a spatial deviation is determined; if both an abnormal trend in the time dimension and a spatial deviation are determined, a second anomaly marker is generated.

2. The surface water anomaly sampling triggering method based on multi-source information fusion according to claim 1, characterized in that, The steps of constructing a digital twin model of the watershed water environment, simulating water quality background values ​​under different operating conditions to generate a dynamic water quality baseline, and extracting abnormal residual features and generating a third anomaly marker through residual analysis of real-time monitoring data and the dynamic water quality baseline include: By combining watershed topography, hydrological parameters, pollution source distribution and historical water quality data, a digital twin model of watershed water environment is constructed. By inputting different meteorological and hydrological parameters, the water quality baseline value of surface water under the corresponding conditions is simulated, and a water quality baseline that changes dynamically with the conditions is generated. Calculate the residual value between real-time monitoring data and the dynamic water quality baseline under the corresponding operating conditions, extract the fluctuation range, duration and abrupt change characteristics of the residual value, and generate a third anomaly marker when the residual value exceeds the preset residual threshold and the anomaly characteristics meet the judgment conditions.

3. The surface water anomaly sampling triggering method based on multi-source information fusion according to claim 2, characterized in that, The steps of calculating the residual value between real-time monitoring data and the dynamic water quality baseline under corresponding operating conditions, extracting the fluctuation amplitude, duration, and abrupt change characteristics of the residual value, and generating a third anomaly marker when the residual value exceeds a preset residual threshold and the anomaly characteristics meet the judgment conditions include: Obtain the dynamic water quality baseline and real-time monitoring data corresponding to the current operating conditions. Classify the water quality indicators and calculate the residual values ​​between the real-time values ​​of each indicator and the corresponding baseline values ​​to obtain the residual sequence of each indicator. Feature extraction is performed on the residual sequence to calculate the maximum fluctuation range of the residual value and the duration of continuous exceedance of the preset basic threshold. At the same time, the mutation nodes and mutation range in the residual sequence are identified to form an abnormal residual feature set. Based on the preset residual threshold and the abnormal feature judgment conditions, the extracted abnormal residual features are compared with the judgment conditions. If the residual value exceeds the preset threshold and the abnormal feature meets the judgment conditions, a third abnormal label is generated.

4. The surface water anomaly sampling triggering method based on multi-source information fusion according to any one of claims 1-3, characterized in that, The steps involved in constructing a watershed topology map including monitoring points and water flow direction, using a graph neural network algorithm to aggregate and analyze three types of anomaly markers, identifying anomaly propagation paths and determining source tracing directions, and outputting anomaly source tracing results and comprehensive anomaly confidence levels include: The three types of anomaly markers at each monitoring point are used as node features input into the watershed topology map. The graph attention network is used to calculate the anomaly association weights between the current node and its upstream and downstream neighboring nodes. The spatial distribution features of anomalies are extracted through multi-layer feature aggregation. Based on the anomaly correlation weight ranking, the upstream path with increasing anomaly signal intensity is identified, and it is determined as the source direction of the abnormal pollution source, generating anomaly source tracing results; The spatial distribution features of anomalies are input into a fully connected neural network, and the comprehensive anomaly confidence score in the 0-1 interval is output.

5. The surface water anomaly sampling triggering method based on multi-source information fusion according to claim 4, characterized in that, The steps include: inputting three anomaly markers from each monitoring point as node features into the watershed topology map; using a graph attention network to calculate the anomaly association weights between the current node and its upstream and downstream neighbors; and extracting the spatial distribution features of anomalies through multi-layer feature aggregation. The first, second, and third anomaly markers are numerically encoded to generate a multidimensional anomaly feature vector, which is then used as the initial feature input for the corresponding monitoring point nodes in the watershed topology map. By aggregating the features of neighboring nodes through the graph attention network, the abnormal association weights between the current node and its upstream and downstream neighboring nodes are calculated. The abnormal feature vectors of the neighboring nodes are then aggregated using a weighted aggregation function to generate the aggregated feature vector of the current node. The aggregated feature vector is input into a nonlinear activation function for feature transformation. The feature aggregation and transformation operations are iteratively executed through a multi-layer graph attention network, and finally the abnormal spatial distribution features that characterize the propagation law of abnormal space are output.

6. A surface water anomaly sampling triggering system based on multi-source information fusion, characterized in that, The system, applicable to the method of any one of claims 1-5, comprises: The data acquisition module is used to collect online monitoring data, environmental condition data, historical water quality data, and external environmental early warning data of surface water, and to construct a corresponding multi-source associated dataset; The data processing module is used to preprocess the multi-source associated dataset, sequentially performing data cleaning, spatiotemporal alignment, and standardization to obtain a noise-free, same-dimensional standard dataset. The anomaly labeling module is used to identify the correlation anomalies between different categories of data in the standard dataset, generate the first anomaly label, and at the same time, combine spatiotemporal comparison to analyze the changing trends and deviations of water quality indicators, identify hidden anomalies, and generate the second anomaly label. The feature extraction module is used to construct a digital twin model of the watershed water environment, simulate the water quality background value under different working conditions to generate a dynamic water quality baseline, and extract abnormal residual features and generate a third abnormal label through residual analysis of real-time monitoring data and the dynamic water quality baseline. The anomaly analysis module is used to construct a watershed topology map that includes monitoring points and water flow direction. It uses a graph neural network algorithm to perform aggregate analysis on three types of anomaly markers, identify anomaly propagation paths and determine the source direction, and output anomaly source tracing results and comprehensive anomaly confidence. The instruction triggering module is used to generate targeted sampling instructions based on the comprehensive anomaly confidence level and anomaly tracing results, and automatically trigger anomaly sampling at associated monitoring points in the water quality sampling network.

7. A readable storage medium having computer instructions stored thereon, characterized in that, When executed by the processor, this instruction implements the steps of the method as described in any one of claims 1-5.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Intelligent water affair monitoring management system based on digital twinning

    CN121169186A

  • Lake water quality abnormity monitoring method and system based on multi-source sensing data fusion

    CN121682603A