A watershed water quality anomaly tracing method and system based on timing fluctuation characteristics

By constructing a watershed map network and utilizing temporal fluctuation characteristics and response lag dimension analysis, the problem of rapid and accurate location of water quality anomalies in the watershed was solved, achieving efficient and accurate location of water quality anomaly sources and improving the real-time performance and coverage of water environment monitoring.

CN121393620BActive Publication Date: 2026-03-27SICHUANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies are insufficient to quickly and accurately determine the source of water quality anomalies in watershed environments. Traditional methods suffer from limited coverage, insufficient real-time performance, high operation and maintenance costs, and long analysis time.

Method used

By constructing a watershed map network and utilizing the fusion analysis of temporal fluctuation characteristics and response lag dimensions, the anomaly correlation coefficient is calculated to accurately locate the source of water quality anomalies.

Benefits of technology

It improves the efficiency and accuracy of water quality anomaly analysis, enables rapid and accurate location of anomaly sources, reduces manpower and material costs, and enhances the real-time performance and coverage of water environment monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121393620B_ABST
    Figure CN121393620B_ABST
Patent Text Reader

Abstract

The application provides a watershed water quality anomaly tracing method and system based on time sequence fluctuation characteristics, which comprises the following steps: taking each monitoring point in the watershed as a graph node, constructing a directed edge connecting the graph nodes according to the spatial relationship and water flow direction between the monitoring points, and forming a watershed graph network; obtaining historical monitoring data, extracting amplitude and fluctuation characteristics from the historical monitoring data, and calculating a time sequence fluctuation characteristic vector; traversing the graph nodes connected by the directed edges, constructing an index pair corresponding to the graph nodes, analyzing the correlation of the index pair under different lag dimensions, and calculating a response lag dimension coefficient; calculating an anomaly correlation coefficient between the graph nodes according to the time sequence fluctuation characteristic vector and the response lag dimension coefficient, and constructing an anomaly transmission path; and dividing a pollution investigation area according to the anomaly transmission path, so as to provide the staff with investigation and disposal. The application can accurately locate the position of a water quality anomaly fluctuation excitation source, and improve the water quality anomaly analysis efficiency and the anomaly source positioning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of water environment pollution monitoring, and in particular relates to a watershed water quality anomaly source tracing method and system based on time series fluctuation characteristics. BACKGROUND

[0002] Due to the complexity of the watershed water environment system, the water environment can be affected by multiple factors. Water quality dynamic analysis is a key means for assessing water body health, identifying pollution transmission paths, and taking timely remedial measures, which is of great significance for maintaining ecological balance. Through water quality dynamic analysis, the propagation path, source and evolution trend of pollutants or abnormal indicators in the water body over time and space are identified and deduced in real time or near real time, so as to realize rapid tracing of abnormal water quality.

[0003] However, traditional water quality analysis focuses on static analysis of water quality indicators, mainly relying on fixed sites or manual sampling, and then comparing the measured values at the site with the assessment standard, which is usually a fixed threshold or multiple classification threshold values. When the water quality is obviously abnormal, manual inspection is carried out. This has many limitations in practical application, such as under the influence of external factors such as river management engineering, seasonal changes, regional industrial adjustment, etc. The overall monitoring data shows systematic deviation, making it difficult to reflect the true abnormality of data changes, so there are generally problems such as limited coverage, insufficient real-time performance, high operation and maintenance costs, etc., making it difficult to achieve dynamic tracking and tracing of pollution.

[0004] On this basis, water quality anomaly analysis technology has gradually become intelligent, including algorithms and fingerprint methods. The algorithm requires sufficient hydrodynamic data to estimate the pollution source by simulating the migration of pollutants, which has high requirements for hydrodynamic data and computing resources. The fingerprint method locates the pollution based on the similarity of the spectral characteristics of pollutants and the spectral characteristics of pollution sources. Due to the complexity of the characteristic signals of pollutants, the number of sewage sample libraries is required to be large. Both have high manpower and material costs, and the analysis takes a long time. There is a significant lag from the occurrence of the anomaly to the response and disposal, and it is difficult to accurately determine the location of the anomaly source for short-term sudden changes in water quality.

[0005] Therefore, there is a lack of a method that can accurately determine the source of water quality anomalies. SUMMARY

[0006] In order to solve the above problems of the prior art, the present application provides a watershed water quality anomaly source tracing method and system based on time series fluctuation characteristics, which can accurately locate the position of the water quality anomaly source.

[0007] In order to achieve the above purpose, the technical solution adopted by the present application is as follows:

[0008] In a first aspect, the application provides a watershed water quality anomaly tracing method based on time series fluctuation characteristics, comprising:

[0009] Step S1, regarding each monitoring point in the watershed as a graph node, constructing a directed edge connecting the graph nodes according to the spatial relationship and water flow direction between the monitoring points, and forming a watershed graph network;

[0010] Step S2, obtaining historical monitoring data, extracting amplitude and fluctuation characteristics from the historical monitoring data, and calculating a time series fluctuation characteristic vector;

[0011] Step S3, traversing the graph nodes connected by the directed edges, constructing an index pair corresponding to the graph nodes, analyzing the correlation of the index pair under different lag dimensions, and calculating a response lag dimension coefficient;

[0012] Step S4, calculating an anomaly correlation coefficient between the graph nodes according to the time series fluctuation characteristic vector and the response lag dimension coefficient, and constructing an anomaly transmission path;

[0013] Step S5, dividing a pollution investigation area according to the anomaly transmission path, and providing the staff for investigation and disposal.

[0014] The application has the beneficial effects that in the case of obvious changes in the watershed water environment, the application adjusts the normal fluctuation range in combination with the measured historical monitoring data, analyzes the time series fluctuation characteristics and the lag dimension, accurately locates the water quality anomaly fluctuation excitation source position, and improves the water quality anomaly analysis efficiency and the anomaly source positioning accuracy.

[0015] Optionally, the step S4 comprises:

[0016] Step S41, calculating an anomaly fluctuation interval coefficient and a fluctuation characteristic similarity for the time series fluctuation characteristic vector of the node pair;

[0017] Step S42, calculating an anomaly correlation coefficient between the graph nodes according to the response lag dimension coefficient, the anomaly fluctuation interval coefficient and the fluctuation characteristic similarity;

[0018] Step S43, locating an anomaly fluctuation excitation node according to the anomaly correlation coefficient between the graph nodes, and constructing a complete anomaly transmission path.

[0019] Optionally, the step S41 comprises:

[0020] Step S411, selecting an index whose time series fluctuation characteristic value exceeds a fluctuation threshold as a core anomaly index in the two nodes, and simplifying the time series fluctuation characteristic vector group of each graph node according to the core anomaly index;

[0021] Step S412, based on the simplified time sequence fluctuation feature vector group, the difference value of each graph node at the maximum fluctuation time of the time sequence fluctuation feature vector is calculated, and a node abnormal fluctuation interval coefficient is obtained;

[0022] Step S413, according to the node abnormal fluctuation interval coefficient, the time sequence fluctuation feature vector group of all graph nodes is translated and aligned;

[0023] Step S414, according to the aligned time sequence fluctuation feature vector group, the fluctuation feature similarity between each graph node is calculated.

[0024] Optionally, the step S42 comprises:

[0025] Step S421, according to the response lag dimension coefficient D and the abnormal fluctuation interval coefficient d, a time decay coefficient W between each graph node is calculated:

[0026] ;

[0027] In the formula, T is the maximum response time range of the two graph node corresponding indexes;

[0028] Step S422, according to the fluctuation feature similarity B and the time decay coefficient W, an abnormal correlation coefficient R between each graph node is calculated:

[0029] R=W×B;

[0030] ;

[0031] In the formula, cos(·) represents the cosine similarity, E (1) represents the long vector after splicing of the first node, E (2) represents the long vector after splicing of the second node, and dis(·) represents the Euclidean distance.

[0032] Optionally, the step S2 comprises:

[0033] Step S21, historical monitoring data is obtained and deperiodization preprocessing is performed;

[0034] Step S22, the preprocessed historical monitoring data is calculated by difference to obtain an amplitude coefficient;

[0035] Step S23, the historical monitoring data of each index is collected to perform percentile regression fitting to obtain a dynamic percentile regression curve, and based on the dynamic percentile regression curve and the current monitoring value, a fluctuation coefficient is calculated;

[0036] Step S24, the amplitude coefficient and the fluctuation coefficient are normalized;

[0037] Step S25, calculating the time series fluctuation feature according to the normalized amplitude coefficient and the fluctuation coefficient.

[0038] Optionally, the formula of the fluctuation coefficient is:

[0039] ;

[0040] In the formula, denotes the fluctuation coefficient of index i at time t, denotes the monitoring value of index i at the time t, , , respectively represent the fitting values of the low, medium and high quantile regression lines of index i at time t;

[0041] The formula of the time series fluctuation feature is:

[0042] ;

[0043] In the formula, denotes the time series fluctuation feature value of index i at time t, denotes the amplitude coefficient of index i at the same time t, denotes the fluctuation coefficient of index i at the same time t.

[0044] Optionally, the step S3 comprises:

[0045] Step S31, constructing two graph nodes with a connection edge as a node pair, and constructing an index pair according to the monitoring index types of the graph nodes;

[0046] Step S32, taking a single data monitoring interval as a lag period, and calculating the correlation coefficients of each index pair at different lag periods within a preset maximum lag period range;

[0047] Step S33, for each index pair, screening the maximum correlation coefficient thereof at each lag period, and determining the lag period corresponding to the maximum correlation coefficient as the response lag dimension of the index pair;

[0048] Step S34, summarizing the response lag dimensions of all index pairs under the same node pair to obtain the response lag dimension coefficient of the node pair.

[0049] Optionally, the step S1 of constructing a directed edge connecting the graph nodes according to the spatial relationship and the water flow direction between the monitoring points comprises:

[0050] When it is a river water quality monitoring point or a hydrological monitoring point, the connection edges between the graph nodes are obtained based on the adjacent relationship between the monitoring points, and the direction of the connection edge is added according to the upstream to downstream manner on the river network.

[0051] When it is a lake water quality monitoring point, a lake area grid is constructed with the lake water quality monitoring point as the center, connection edges between various graph nodes are obtained according to the lake area grid, and the direction of the connection edges is added according to the mode of adjacent grid transmission, and if the lake area grid has no determined flow direction, a bidirectional connection edge is constructed;

[0052] When it is a river node and a lake area node, a connection edge of a river inlet into a lake and an adjacent lake area grid is constructed, and the direction of the connection edge is added according to the mode of the river inlet into the lake to the adjacent lake area grid;

[0053] When it is a river outside monitoring point, a connection edge is added according to a monitoring point on an adjacent river section of a catchment area, and the direction of the connection edge is added according to a transmission link of a river outside variable affected by a catchment area to adjacent river section monitoring point data.

[0054] Optionally, the step S5 comprises:

[0055] Step S51, obtaining a time sequence fluctuation feature vector group of all monitoring points on the abnormal transmission path, determining a monitoring point needing to be investigated according to a maximum value of the time sequence fluctuation feature vector, and listing a catchment area corresponding to the monitoring point needing to be investigated as an investigation area;

[0056] Step S52, traversing each monitoring point of the abnormal transmission path, splicing a time sequence fluctuation feature vector group of the monitoring point into a long vector according to a fixed index order, matching a historical event of a same type site of the monitoring point, and matching a reference event from the historical event according to a cosine similarity between the long vectors;

[0057] Step S53, according to an event record of the reference event, providing a staff to go to the to-be-investigated area for investigation and disposal, and recording and arranging as water quality abnormality early warning event information.

[0058] In a second aspect, the present application provides a watershed water quality anomaly tracing system based on time sequence fluctuation characteristics, comprising a memory, a processor and a computer program stored on the memory and executable on the processor, and the processor implements the first aspect of the watershed water quality anomaly tracing method based on time sequence fluctuation characteristics when executing the computer program.

[0059] The technical effects of the watershed water quality anomaly tracing system based on time sequence fluctuation characteristics provided in the second aspect are referred to the related description of the watershed water quality anomaly tracing method based on time sequence fluctuation characteristics provided in the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0060] Figure 1 It is a main flow diagram of the watershed water quality anomaly tracing method based on time sequence fluctuation characteristics of the embodiments of the present application.

[0061] Figure 2 A schematic diagram of a dynamic quantile regression curve involved in an embodiment of the present application;

[0062] Figure 3 A schematic diagram of a river basin graph network involved in an embodiment of the present application;

[0063] Figure 4 A structural schematic diagram of a river basin water quality anomaly tracing system based on time series fluctuation characteristics in an embodiment of the present application.

[0064] Explanation of reference signs:

[0065] 1. A river basin water quality anomaly tracing system based on time series fluctuation characteristics;

[0066] 2. A processor;

[0067] 3. A memory. DETAILED DESCRIPTION

[0068] In order to better understand the above technical solutions, exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a clearer, more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0069] The present embodiment is applicable to scenarios where water quality monitoring of a river basin water environment system is required. The prior art cannot adapt to scenarios where the river basin water environment changes rapidly, making it difficult to accurately determine the location of an anomaly source.

[0070] Therefore, the present application provides a river basin water quality anomaly tracing method based on time series fluctuation characteristics. Each monitoring point in the river basin is taken as a graph node, and a directed edge connecting the graph nodes is constructed according to the spatial relationship and water flow direction between the monitoring points, forming a river basin graph network. Historical monitoring data is obtained, and amplitude and fluctuation characteristics are extracted from the historical monitoring data to calculate a time series fluctuation characteristic vector. The graph nodes connected by the directed edge are traversed, and an index pair corresponding to the graph nodes is constructed. The correlation of the index pair is analyzed under different lag dimensions, and a response lag dimension coefficient is calculated. An anomaly correlation coefficient between the graph nodes is calculated according to the time series fluctuation characteristic vector and the response lag dimension coefficient, and an anomaly transmission path is constructed. The anomaly transmission path is used to divide a pollution investigation area for staff to investigate and dispose. Thus, the present application adjusts the normal fluctuation range by combining the measured historical monitoring data, and accurately determines the location of the anomaly source by fusion analysis of the time series fluctuation characteristics and the lag dimension.

[0071] The present application will be described below in conjunction with specific embodiments.

[0072] Referring to Figure 1 It can be known that the embodiment of the application provides a watershed water quality anomaly tracing method based on timing fluctuation characteristics, comprising:

[0073] Step S1, regarding each monitoring point in the watershed as a graph node, constructing a directed edge connecting the graph nodes according to the spatial relationship and water flow direction between the monitoring points, and forming a watershed graph network.

[0074] Among them, in the watershed to be monitored, there are two types of monitoring points, namely water quality monitoring points and hydrological monitoring points, the water quality monitoring indexes include total nitrogen, total phosphorus, dissolved oxygen, water temperature, ammonia nitrogen, chemical oxygen demand, etc., and the hydrological monitoring indexes include flow, flow rate, water level.

[0075] At this time, each monitoring point is regarded as a graph node of the watershed graph network, then the connection edges between the corresponding graph nodes can be constructed based on the spatial relationship of the monitoring points in the watershed, and finally the directions of the connection edges are given according to the flow direction between the monitoring points, forming a directed edge.

[0076] Therefore, in an embodiment, the step S1 of constructing a directed edge connecting the graph nodes according to the spatial relationship and water flow direction between the monitoring points comprises:

[0077] Step S11, when it is a river water quality monitoring point or a hydrological monitoring point, obtaining the connection edges between the graph nodes based on the adjacent relationship between the monitoring points, and adding the direction of the connection edges according to the way from the upstream to the downstream of the river network.

[0078] Step S12, when it is a lake water quality monitoring point, constructing a lake area grid centered on the lake water quality monitoring point, obtaining the connection edges between the graph nodes according to the lake area grid, and adding the direction of the connection edges according to the way of adjacent grid transmission, if there is no lake area grid with a certain flow direction, then a bidirectional connection edge is constructed.

[0079] Step S13, when it is a river node and a lake node, constructing the connection edges of the river inlet into the lake and the adjacent lake area grid, and adding the direction of the connection edges according to the way from the river inlet into the lake to the adjacent lake area grid.

[0080] Step S14, when it is an out-of-river monitoring point, adding the connection edges according to the monitoring points on the adjacent river sections of the catchment area, forming the transmission link of the out-of-river variable affected by the catchment area to the adjacent river section monitoring point data to add the direction of the connection edges.

[0081] That is, the steps S11 to S14 described above describe the construction of the connection edges and the assignment of the directions on the connection edges by different monitoring points, and the directed edges between all the monitoring points in the watershed are constructed.

[0082] Step S2, obtaining historical monitoring data, extracting amplitude features and fluctuation features from the historical monitoring data, and calculating a time series fluctuation feature vector.

[0083] The historical monitoring data is the real data actually measured by the monitoring point equipped with a sensor. The amplitude features and fluctuation features between the data are extracted from the historical real monitoring data, so as to obtain the fluctuation range of the monitoring point, thereby adapting to the case where the water environment of the basin changes obviously.

[0084] In an embodiment, step S2 comprises:

[0085] Step S21, obtaining historical monitoring data and performing decyclic preprocessing;

[0086] In order to avoid similar periodic changes interfering with time series feature recognition, the monitoring index data containing periodic changes, such as dissolved oxygen, water temperature and other water quality indicators, are decycled in this embodiment, and the overall periodic change trend is removed through modeling.

[0087] Meanwhile, in the basin water environment monitoring system, the automatic monitoring station continuously collects water quality index data including total nitrogen, total phosphorus, dissolved oxygen, water temperature, and hydrological index data such as flow and flow rate, meteorological index data such as rainfall and air temperature, and other water environment related monitoring data. Due to the obvious diurnal and seasonal change law of the natural environment, the dissolved oxygen concentration usually increases in the daytime due to the enhancement of photosynthesis, and decreases at night due to the dominance of respiration, showing a significant daily periodic fluctuation. The water temperature also shows a similar daily change trend with the intensity of sunlight. If the original data is directly used for abnormal event detection or pollution early warning model analysis, these inherent periodic fluctuations may mask the real abnormal trend signal, leading to misjudgment or missed detection.

[0088] The periodic fluctuation signal repeats with time, and the periodic change type includes daily, weekly, monthly, yearly and other cases, which is also called seasonality in time series analysis, that is, the regularity of the repeated variation of the time series data within a year or a shorter period. The common methods to eliminate seasonality include modeling methods such as difference calculation, curve fitting and moving average. Taking the classic seasonal decomposition hypothesis as an example: time series = trend + seasonality + residual, the seasonal signal needs to be decomposed from the original time series, and the data after removing the periodic fluctuation is obtained by the original sequence-seasonality.

[0089] (1) Calculate and remove the trend component using the moving average method

[0090] The trend is calculated by the moving average method, and the moving average with a window size equal to the length of the seasonal period is used, such as 24-hour periodic change data using 24-window moving average. This moving average will smooth out the seasonal fluctuations and retain the long-term trend of the time series.

[0091] (2) Average the seasonal cycle of the detrended sequence to extract the seasonal pattern

[0092] The seasonal pattern is extracted from the detrended sequence, detrended sequence = original sequence - trend. The seasonal pattern is extracted from the detrended sequence, that is, the average value is calculated at the corresponding position of each seasonal cycle, such as the 24-hour cycle change data is the average value of the numerical value at the same time point every day. This average value represents the typical pattern of the seasonal cycle, indicating the fluctuation value of each time in the 24-hour monitoring, and the time series calculated is the cycle change signal, that is, the periodic fluctuation value repeated every day.

[0093] (3) Calculate the residual sequence

[0094] Residual = original sequence - trend - seasonality, the residual part reflects the abnormal fluctuation beyond the seasonal pattern. That is, a time series, excluding short-term seasonal fluctuations and long-term trend changes, the remaining signal indicates an abnormal signal, that is, the monitoring value change under the influence of external factors.

[0095] Step S22, difference calculation is used in the pretreated historical monitoring data to obtain the amplitude coefficient:

[0096] ;

[0097] In the formula, i represents the i-th index, t represents the t-th time, represents the amplitude coefficient of index i at time t, represents the monitoring value of index i at time t.

[0098] Among them, based on the time series after the decoupling of the previous step, the change coefficient of the current time compared with the previous time is obtained by subtracting the monitoring value of the previous time from the monitoring value of the current time. In this way, the original monitoring value is converted into the change amount of each time, highlighting the short-term fluctuation characteristics. At the same time, the amplitude coefficient can be further used to construct an anomaly detection model, thereby improving the sensitivity of water environment supervision.

[0099] Step S23, collect the historical monitoring data of each index to perform percentile regression fitting to obtain a dynamic percentile regression curve, and based on the dynamic percentile regression curve and the current monitoring value, calculate the fluctuation coefficient, and the formula of the fluctuation coefficient is:

[0100] ;

[0101] In the formula, represents the fluctuation coefficient of index i at time t, represents the monitoring value of index i at the th time, , , represent the fitted value of the low, medium, high quantile regression line of indicator i at time t, respectively.

[0102] Specifically, in step S23, first, the monitoring data of the monitoring point in the past one month is collected, and the systematic changes of external factors such as the local possible transition from the rainy season to the dry season, water mobility, illumination conditions, and the like are considered. In order to more accurately depict the normal fluctuation range, the quantile regression method is used in the present application to respectively fit the quantile regression curves of the 20% low quantile line , the 50% medium quantile line , and the 80% high quantile line The numerical values of these curves dynamically change with time t, reflecting the typical fluctuation interval under different seasons, weather, or tourist peak seasons and the like.

[0103] In a specific example, the value of a certain indicator at time t is = 25 μg / L, and the fitted values of the quantile regression at the corresponding time are respectively: the low quantile line = 12 μg / L, the medium quantile line = 18 μg / L, and the high quantile line = 28 μg / L. Since , the fluctuation coefficient is calculated using the high quantile interval according to the formula:

[0104] = (25-18) / (28-18) = 0.7;

[0105] Therefore, as shown in Figure 2 , the fluctuation coefficient = 0.7 indicates that the current observation value is not yet above the normal fluctuation upper limit of the historical 80% quantile, but is already in a higher position above the median line, suggesting that the indicator may start to actively grow. If the value at the subsequent time continues to approach or exceed 1.0, it represents that the indicator has a significant anomaly. Conversely, if the value at a certain time is = 10 μg / L , the low quantile interval is used to calculate = (10-18) / (12-18) ≈ 1.33 At this time, the absolute value of the negative fluctuation coefficient is greater than 1, indicating that the value of the indicator is significantly lower than the historical normal fluctuation level, reflecting that the water environment at the site is significantly disturbed, and further verification is required. In other embodiments, the selection of low, medium, and high percentiles is based on demand, and is further adjusted according to the actual application effect.

[0106] Among them, if the calculated fluctuation coefficient , it is usually considered that , the indicator monitoring data is in the normal fluctuation range, and when When the index monitoring data is in an abnormal numerical range, it may be disturbed by other factors, and obvious high or low values may appear. By traversing all indexes and calculating the amplitude coefficient and the fluctuation coefficient of each index at each time point The original monitoring value is converted into the fluctuation characteristic quantity at each time point, and the difference between the fluctuation value at the current time point and the normal fluctuation range is quantified while focusing on the long-time fluctuation characteristics.

[0107] Step S24, normalizing the amplitude coefficient and the fluctuation coefficient.

[0108] Specifically, the amplitude coefficient and the fluctuation coefficient are scaled linearly to fall within [-1, 1], and the calculation formula is:

[0109] ;

[0110] The amplitude coefficient and the fluctuation coefficient of an index at multiple time points have been calculated through the foregoing steps, but since different monitoring indexes have different numerical magnitudes, the calculated coefficients also have different magnitudes, so linear scaling is performed to map them to a specified numerical interval.

[0111] In a specific example, for the monitoring values of an index at the last 5 time points, the calculated amplitude coefficient A = [-3.2, 0.8, -1.4, 4.0, 1.5], the maximum absolute value max(|A|) = max(3.2, 0.8, 1.4, 4.0, 1.5) = 4.0, and the new sequence A' = [3.2 / 4.0, 0.8 / 4.0, (-1.4) / 4.0, 4.0 / 4.0, 1.5 / 4.0] = [-0.8, 0.2, -0.35, 1.0, 0.375] is obtained by applying the scaling formula. Similarly, for the fluctuation coefficient S = [1.3, 0.6, 0.1, 2.1, 2.5], the scaled sequence S' = [1.3 / 2.5, 0.6 / 2.5, 0.1 / 2.5, 2.1 / 2.5, 2.5 / 2.5] = [0.52, 0.24, 0.04, 0.84, 1.0] is obtained.

[0112] After the above processing, all feature coefficients are normalized to the range [-1, 1], which not only preserves the original positive / negative direction and relative intensity of change, but also eliminates the magnitude difference, helping to improve the stability of subsequent data analysis and the accuracy of abnormal identification. It should be noted that the scaled coefficients may differ greatly when intercepting time series of different lengths, because the historical monitoring data may contain large abnormal values, making the numerator max(|X|) large, thereby reducing the current coefficient value. For water quality mutation events, only short-term data features are usually concerned, such as the presence of obvious abnormalities in the current time of the indicator. In this case, the sequence 12 hours or 24 hours before the abnormality occurs is intercepted for feature scaling, focusing on recent features of the monitoring indicator. The length of the time series also needs to be adjusted and optimized in combination with the monitoring frequency of the site, manual analysis experience, and actual application effect. In this embodiment, it is not repeated.

[0113] Step S25, according to the normalized amplitude coefficient and the fluctuation coefficient, the time series fluctuation feature is calculated, and the formula of the time series fluctuation feature is:

[0114] ;

[0115] In the formula, Ei(t) represents the time series fluctuation feature value of the indicator i at time t, Ai(t) represents the amplitude coefficient of the indicator i at the same time t, Si(t) represents the fluctuation coefficient of the indicator i at the same time t.

[0116] In this embodiment, the time series fluctuation feature value at each time is calculated according to the formula, to comprehensively reflect the coupling abnormal signal of the indicator in the two dimensions of change amplitude and fluctuation deviation.

[0117] In a specific example, for the monitoring values of a certain indicator in the last 5 time points, the amplitude coefficient A = [-0.8, 0.2, -0.35, 1.0, 0.375], the fluctuation coefficient S = [0.52, 0.24, 0.04, 0.84, 1.0], and the time series fluctuation feature E = [-0.416, 0.048, -0.14, 0.84, 0.375] is further obtained.

[0118] Therefore, by multiplying the amplitude signal and the fluctuation signal, the time series fluctuation feature can effectively suppress the situation of only large change but still within the normal range, and the situation of deviating from the normal but changing gently, highlighting the high-risk situation of rapid change and significant deviation, quickly locking the abnormal triggering time point, and significantly improving the precision and robustness of the water environment intelligent analysis system.

[0119] Step S3, traversing the graph nodes connected by the directed edges, constructing the index pair corresponding to the graph nodes, analyzing the correlation of the index pair under different lag dimensions, and calculating the response lag dimension coefficient.

[0120] In an embodiment, step S3 comprises:

[0121] Step S31, constructing two graph nodes connected by the edges as a node pair, and constructing the index pair according to the monitoring index types of the graph nodes.

[0122] Among them, for the node pair connected by the edges, the correlation coefficient of the monitoring indexes of the two graph nodes under different lags is calculated. Because it involves the site type corresponding to the graph node, in step S31 of the embodiment, the index pair is constructed according to the site type corresponding to the graph node, including:

[0123] If the monitoring indexes corresponding to the two graph nodes are indexes of the same type, the lag correlation coefficient between the same indexes is calculated, and if the monitoring indexes corresponding to the two graph nodes are indexes of different types, the lag correlation coefficient between the two groups of indexes corresponding to the two graph nodes is calculated.

[0124] The following is illustrated by an example: in a river basin graph network, there are multiple monitoring points corresponding to multiple graph nodes in the graph. The first node is located upstream and the second node is located downstream, and there is a water flow connection relationship between them. At this time:

[0125] Suppose the first node and the second node are both water quality sites, monitoring total phosphorus, total nitrogen and ammonia nitrogen. Since the water flows from the first node to the second node, the pollutants will usually appear downstream after a few minutes or hours. In order to capture this transmission delay, the same indexes of the first node and the second node are matched and analyzed, and the lag correlation coefficients of the total phosphorus sequence, the lag correlation coefficients of the total nitrogen sequence, and the lag correlation coefficients of the ammonia nitrogen sequence are calculated.

[0126] Suppose the first node is a hydrological station monitoring water level and flow rate, and the second node is a water quality site monitoring total phosphorus, total nitrogen and ammonia nitrogen. Since the water flows from the first node to the second node, the influence of water level and flow rate on the water quality of the second node may be reflected after a few minutes or hours. In order to capture this lag effect, the indexes between the first node and the second node are cross-matched and analyzed, and the lag correlation coefficients of the water level sequence with the total phosphorus, total nitrogen and ammonia nitrogen sequences, and the lag correlation coefficients of the flow rate sequence with the total phosphorus, total nitrogen and ammonia nitrogen sequences are calculated.

[0127] Step S32, taking a single data monitoring interval as a lag period, and calculating the correlation coefficients of each index pair under different lag periods within a preset maximum lag period range.

[0128] wherein the correlation coefficient of the calculation index at different lag periods, such as lag 1 period, lag 2 period, etc., is a data monitoring interval p. When the first node has an edge pointing to the second node, the lag correlation coefficient between the index i of the first node and the index j of the second node is calculated, and the specific formula is:

[0129] ;

[0130] wherein r h is the correlation coefficient at lag h period, Corr() represents the correlation coefficient calculation, commonly used statistical methods include Pearson, Spearman, etc., represents the value of the index j at time t, and N represents that the correlation analysis data contains historical monitoring data of the previous N periods.

[0131] In a specific example, such as between the index i of the first node and the index j of the second node , the correlation analysis uses N=24 periods of historical monitoring data, and the monitoring interval p is 1 hour, so 24 hours of monitoring data are intercepted, and the correlation at lag h=0, 1, 2, …, 6 periods, i.e., the lag influence relationship of 0 to 6 hours, is investigated. Corresponding to: , , …, . r0~r6 respectively describe whether the index i of the first node at 0 to 6 hours before the current time has a significant correlation with the index j of the second node.

[0132] Step S33, for each index pair, the maximum correlation coefficient at each lag period is screened, and the lag period number corresponding to the maximum correlation coefficient is determined as the response lag dimension of the index pair.

[0133] wherein the lag period h starts from 0, at this time, the maximum lag period number needs to be considered to determine the range of the lag period h. Specifically, assuming that from the abnormal fluctuation of the index i of the first node, the time to affect the index j of the second node does not exceed K i,j hours, then the maximum lag period number max(L i,j ) is calculated as follows:

[0134] ;

[0135] In the formula, represents the upward rounding.

[0136] Therefore, the lag period h represents the fluctuation of the index i value of the first node, which affects the time of the index j of the second node. For example, if the total phosphorus of a certain upstream site has an abnormally high value, it is transmitted to the downstream site after 3 hours, and the concentration value of the total phosphorus may be different, but it can also be observed that the total phosphorus has a significant increase. In addition, if the rainfall data of a certain urban area is monitored, the water quality station of the corresponding river section in the catchment area observes the overall change in water quality after 2 hours, which reflects the interval time from the start of the rainfall, the washing and gathering into the river, and finally the observation of the water quality fluctuation by the water quality station. Similarly, the r h There may be significant differences, such as seasonal changes in aquatic plants in the river, and reconstruction of rain and sewage pipe networks in the town, which will significantly affect the lag influence relationship between the indicators of the two nodes. Therefore, the upper limit K (i,j) of the lag period h can be obtained according to the modeling, such as the transmission of the upstream and downstream of the river, and the simulation of the pollution concentration diffusion rate by constructing a mechanism model, but there are many complex scenarios in the basin, and the requirements for data quality and modeling accuracy are high. Considering the actual application and management needs, a time range can also be determined based on the summary of historical water quality abnormal records and the actual management experience. The range of the lag period h needs to be considered in combination with the actual situation of the basin, and the r h is calculated by traversing different lag periods h, which can automatically identify the lag response relationship between cross-node and cross-indicators, so as to more accurately depict the dynamic coupling characteristics of the water environment system.

[0137] Step S34, aggregating the response lag dimensions of all index pairs of the same node to obtain the response lag dimension coefficient of the node pair.

[0138] At this time, in the above step S33, the maximum correlation coefficient between the indicators i and j is selected, and the lag period number corresponding to the maximum correlation coefficient is taken as the response lag dimension . In step S34, the average value of the u response lag dimensions of all indicators between the first node and the second node is obtained, and the response lag dimension coefficient D between the nodes is calculated, and the specific formula is:

[0139] ;

[0140] That is, in step S3 of the embodiment, step S31 obtains a plurality of index pairs between two nodes, traverses each index pair (i, j), and obtains r h through step S32 to calculate the r h corresponding to the maximum r The response lag dimension of each index pair can be different due to the differences in physical and chemical properties and differences in measurement techniques. Therefore, the response lag dimensions of all index pairs between the two nodes are averaged by step S34 to obtain a response lag dimension coefficient D, which represents the overall response lag level between the monitoring indexes of the two nodes.

[0141] Step S4: Calculate the abnormal correlation coefficient between the graph nodes according to the time series fluctuation feature vector and the response lag dimension coefficient, and construct the abnormal transmission path.

[0142] In an embodiment, step S4 includes:

[0143] Step S41: Calculate the abnormal fluctuation interval coefficient and the fluctuation feature similarity for the time series fluctuation feature vector of the node pair.

[0144] Specifically, in this embodiment, step S41 includes:

[0145] Step S411: Select the indexes of the node whose time series fluctuation feature value exceeds the fluctuation threshold as the core abnormal indexes, and simplify the time series fluctuation feature vector group of each graph node according to the core abnormal indexes.

[0146] In one specific example, the abnormal traceability rules corresponding to different types of monitoring points are different. When the monitoring index types of the first node and the second node are different, the number of indexes of the node with the least abnormal fluctuation indexes among the two graph nodes is taken as the reference, and the indexes of the other node are screened, and the first n indexes with the largest E value are selected as the core abnormal indexes. When the monitoring indexes of the first node and the second node are of the same type, the same indexes are selected as the core abnormal indexes.

[0147] In one specific example, in a certain river basin monitoring network, two connected monitoring nodes are analyzed, such as a first node being a water quality station, and monitoring indexes including dissolved oxygen, ammonia nitrogen, total phosphorus, and total nitrogen, wherein the dissolved oxygen, total phosphorus, and total nitrogen have abnormal fluctuation moments; a second node also being a water quality station, and monitoring indexes including ammonia nitrogen, total phosphorus, total nitrogen, and chemical oxygen demand, wherein the ammonia nitrogen, total phosphorus, and total nitrogen have abnormal fluctuation moments. The same indexes after screening are total phosphorus and total nitrogen, and the time series fluctuation feature vector group of the first node and the second node only retains total phosphorus and total nitrogen. In another case, such as the second node being a hydrological station, and monitoring indexes including flow rate, flow, and water level, wherein the flow rate and water level have abnormal fluctuation moments. At this time, the first node has 3 indexes with abnormal fluctuation moments, and the second node has 2 indexes with abnormal fluctuation moments, so that based on the index number 2, the first node selects the first 2 indexes with the maximum E value, such as total phosphorus and total nitrogen, and the time series fluctuation feature vector group of the first node only retains total phosphorus and total nitrogen, and the time series fluctuation feature vector group of the second node only retains flow rate and water level.

[0148] When the core abnormal index is obtained, the time series fluctuation feature vector group of the first node and the second node is obtained according to step S25. The length of the time series fluctuation feature vector E is the maximum lag number of step S33 The interception can also be set according to actual needs, and the fluctuation threshold φ e The screening E≥φ e marks the abnormal fluctuation moment of the index.

[0149] In one specific example, the time series fluctuation feature vector E is intercepted according to the maximum lag number =6 of the certain nodes, that is, the of the index i is intercepted, and whether it is≥φ e is judged one by one, if the ≥φ e at any moment, that is, it is determined that there is an index abnormal fluctuation moment, otherwise it is determined that there is no index abnormal fluctuation moment.

[0150] Step S412, based on the simplified time series fluctuation feature vector group, the difference value of each graph node at the maximum fluctuation moment of the time series fluctuation feature vector is calculated, and a node abnormal fluctuation interval coefficient is obtained.

[0151] Among them, based on the simplified time series fluctuation feature vector group, the moment with the maximum E value of the index is selected, and the difference value between the moment with the maximum E value of the index of the first node i and the moment with the maximum E value of the index of the second node j When the monitoring indicators of the first node and the second node are of the same type, the difference between the same indicators is matched, and when the monitoring indicator types of the first node and the second node are different, the difference is calculated between the two groups of indicators corresponding to the nodes. Further, the g abnormal fluctuation time difference values of all fluctuation indicator pairs are obtained, and the abnormal fluctuation interval coefficient d between the nodes is calculated, and the specific formula is:

[0152] ;

[0153] In a specific example, the characteristic value Ei(t-1) of the indicator i of the first node is the largest, and the characteristic value Ej(t-4) of the indicator j of the second node is the largest, and then In a specific example, the characteristic value Ei(t-1) of the indicator i of the first node is the largest, and the characteristic value Ej(t-4) of the indicator j of the second node is the largest, and then In a specific example, the characteristic value Ei(t-1) of the indicator i of the first node is the largest, and the characteristic value Ej(t-4) of the indicator j of the second node is the largest, and then = |(t-1)-(t-4)| = 3. The of all indicator pairs between the first node and the second node is calculated, and the average of all is taken to obtain the abnormal fluctuation interval coefficient d.

[0154] Step S413, according to the node abnormal fluctuation interval coefficient, the time sequence fluctuation feature vector group of all graph nodes is translated and aligned.

[0155] Among the simplified time sequence fluctuation feature vector group, the time of the largest E value of all indicators of the first node is obtained, and the indicator i farthest from the current time is screened, that is, the time of the largest E value is closest to the current time. Similarly, the second node is screened to obtain the indicator j closest to the current time of the largest E value of the second node. For example, the characteristic value Ei(t-1) of the indicator i of the first node is the largest, and the characteristic value Ej(t-4) of the indicator j of the second node is the largest. In a specific example, the characteristic value Ei(t-1) of the indicator i of the first node is the largest, and the characteristic value Ej(t-4) of the indicator j of the second node is the largest, and then The time difference between the two is 3, and the time sequence fluctuation feature vector group of all indicators of the first node is translated by 3, such as the time sequence fluctuation feature vector of the indicator i after translation At this time, the E value maximum time of the first node i indicator and the second node j indicator, and are all in the second last position of the vector, and the E value maximum time of the remaining indicator pairs is not necessarily aligned.

[0156] Step S414, according to the aligned time sequence fluctuation feature vector group, the fluctuation feature similarity between each graph node is calculated.

[0157] Among them, according to the aligned new vector group, the order of the time sequence fluctuation feature vector is adjusted, and the fluctuation feature similarity is further calculated.

[0158] ​​​​Specifically, when the first node and the second node monitor different types of indicators, the time series fluctuation feature vectors of the two nodes are arranged in the order of E

[0159] , ;

[0160] In the formula, E (1) represents the first node, E (2) represents the second node after splicing. By calculating the cosine similarity cos(E (1) ,E (2) ) and the Euclidean distance dis(E (1) ,E (2) ) between the long vectors, the fluctuation feature similarity B is further obtained. When the length of the long vector is q, the specific formula is:

[0161] ;

[0162] ;

[0163] .

[0164] In a specific example, in a certain river basin monitoring network, two connected monitoring nodes are analyzed. The first node is a hydrological station, and the time series fluctuation feature vector group after simplification and alignment includes flow [0.33, -0.07, 0.45, -0.19, 0.12, 0.77, 0.63] and water level [0.11, 0.28, 0.19, 0.35, 0.89, 0.53, -0.22]. The second node is a water quality station, and the time series fluctuation feature vector group after simplification and alignment includes total phosphorus [0.09, 0.32, -0.05, 0.22, 0.12, 0.72, 0.45] and total nitrogen [0.03, -0.07, -0.20, 0.38, 0.58, 0.88, -0.11]. The first node and the second node monitor different types of indicators, and the time series fluctuation feature vectors of the two nodes are arranged in the order of E max value. The first node splicing order is flow and water level, and the corresponding E (1)= [0.33, -0.07, 0.45, -0.19, 0.12, 0.77, 0.63, 0.11, 0.28, 0.19, 0.35, 0.89, 0.53, -0.2], the second node splicing order is total phosphorus, total nitrogen, and the corresponding E (2) = [0.09, 0.32, -0.05, 0.22, 0.12, 0.72, 0.45, 0.03, -0.07, -0.20, 0.38, 0.58, 0.88, -0.11]. Further calculation obtains cos(E (1) , E (2) ) ≈ 0.77, dis(E (1) , E (2) ) ≈ 1.08, and the corresponding fluctuation characteristic similarity B ≈ 0.36.

[0165] Step S42, according to the response lag dimension coefficient, the abnormal fluctuation interval coefficient and the fluctuation characteristic similarity, the abnormal correlation coefficient between each graph node is calculated.

[0166] In an embodiment, step S42 comprises:

[0167] Step S421, according to the response lag dimension coefficient and the abnormal fluctuation interval coefficient, the time decay coefficient between each graph node is calculated.

[0168] Wherein, according to the response lag dimension coefficient D and the abnormal fluctuation interval coefficient d between the first node and the second node, the time decay coefficient W is calculated:

[0169] ;

[0170] In the formula, T is the maximum response time range of the corresponding indexes of the two graph nodes, and the maximum value in all maximum lag period numbers max(L i,j ) in step S33 of the corresponding node is taken.

[0171] Step S422, according to the fluctuation characteristic similarity and the time decay coefficient, the abnormal correlation coefficient between each graph node is calculated.

[0172] Wherein, according to the time sequence fluctuation characteristic similarity B and the time decay coefficient W between the first node and the second node, the abnormal correlation coefficient R between the nodes is calculated:

[0173] R = W × B;

[0174] In a specific example, if the first node is a hydrological station and the second node is a water quality station, the response lag dimension coefficient D between the two nodes is 3, the abnormal fluctuation interval coefficient d is 2, the maximum response time range T is 12, the fluctuation characteristic similarity B is 0.36, and the time decay coefficient W between the first node and the second node is calculated as e(-(2×|2-3|) / 12) ≈0.85, further get abnormal correlation coefficient R = 0.85 * 0.36 ≈ 0.31.

[0175] Step S43, according to the abnormal correlation coefficient between each graph node, the abnormal fluctuation excitation node is located, and the complete abnormal transmission path is constructed.

[0176] Wherein, according to the abnormal correlation coefficient R between the first node and the second node, it is judged whether the abnormal correlation coefficient is greater than the preset threshold value, when the abnormal correlation coefficient R of the connection edge is greater than or equal to the preset threshold value R , it is determined that there is abnormal fluctuation correlation between the nodes, otherwise it is determined that there is no abnormal fluctuation correlation;

[0177] When there is abnormal fluctuation correlation between the first node and the second node, the first node is taken as a new second node, and all nodes having connection edges with the first node are traversed, when the connection edge points to the new second node, the traversed node is taken as a new first node, and the new abnormal correlation coefficient R is further calculated.

[0178] When there is no abnormal fluctuation correlation between the first node and the second node, it is determined that it is a normal monitoring point; if all connection edges of the second node have no abnormal fluctuation correlation, the second node currently traversing the connection edge is determined as an excitation node of abnormal fluctuation.

[0179] In a specific example, in a certain basin monitoring network, the preset abnormal correlation coefficient threshold value R = 0.3. Referring to Figure 3 , it can be seen that when the water quality station 6 appears abnormal water quality early warning, the water quality station 6 is taken as a second node, and the nodes having connection edges pointing to the water quality station 6 are traversed. The abnormal correlation coefficient of water quality station 3 and water quality station 6 is 0.47 ≥ R φ, it is determined that water quality station 3 and water quality station 6 have abnormal correlation; the abnormal correlation coefficient of water quality station 4 and water quality station 6 is 0.15 < R φ, it is determined that water quality station 4 and water quality station 6 have no abnormal correlation, and water quality station 4 is a normal monitoring point.

[0180] At this time, the water quality station 3 is taken as a new second node, and the nodes having connection edges pointing to the water quality station 3 are traversed, that is, the water quality station 8 and the water quality station 2 are taken as new first nodes. The abnormal correlation coefficient of water quality station 8 and water quality station 3 is 0.59 ≥ R φ, it is determined that water quality station 8 and water quality station 3 have abnormal correlation; the abnormal correlation coefficient of water quality station 2 and water quality station 3 is 0.52 ≥ R φ, it is determined that water quality station 2 and water quality station 3 have abnormal correlation.

[0181] At this time, the water quality station 8 is taken as a new second node, and the nodes having connection edges pointing to the water quality station 8 are traversed. The abnormal correlation coefficient of water quality station 7 and water quality station 8 is 0.09 < R, it is determined that the water quality station 7 and the water quality station 8 have no abnormal correlation, and the water quality station 7 is a normal monitoring point. At this time, all the connection edges pointing to the water quality station 8 have no abnormal fluctuation correlation, and it is determined that the water quality station 8 is an abnormal fluctuation excitation node.

[0182] Similarly, taking the water quality station 2 as a new second node, traversing the nodes connected to the water quality station 2, and finally backtracking to obtain the excitation node rainfall station 1. Finally, two abnormal fluctuation excitation nodes of the water quality warning event of the water quality station 6 and the corresponding abnormal transmission path are obtained, that is, Figure 3 The black bold part in the middle.

[0183] Step S5, according to the abnormal transmission path, the pollution investigation area is divided, and the staff is arranged to investigate and dispose.

[0184] In an embodiment, step S5 comprises:

[0185] Step S51, obtain the time sequence fluctuation feature vector group of all monitoring points on the abnormal transmission path, determine the monitoring points that need to be investigated according to the maximum value of the time sequence fluctuation feature vector, and list the catchment area corresponding to the monitoring points that need to be investigated as the investigation area.

[0186] Among them, according to the maximum value of the time sequence fluctuation feature vector to determine the monitoring points that need to be investigated can be: according to the maximum value of the time sequence fluctuation feature vector of the monitoring point, that is, according to max(E) of each monitoring point in descending order, and then selecting the first m monitoring points according to the order list. The monitoring points that need to be investigated; or set a maximum threshold value φ E , select the monitoring points with max(E)>φ E as the monitoring points that need to be investigated.

[0187] Step S52, traverse each monitoring point of the abnormal transmission path, splice the time sequence fluctuation feature vector group into a long vector according to the fixed index order, match the historical events of the same type of station of the monitoring point, and match the reference event from the historical events according to the cosine similarity between the long vectors.

[0188] Among them, a similarity threshold value φ C can be set, and the historical events with cosine similarity C>similarity threshold value φ C are selected as the reference events through the cosine similarity C between the long vectors.

[0189] Step S53, according to the event record of the reference event, the staff goes to the investigation area to investigate and dispose, and records and arranges the water quality abnormal warning event information.

[0190] In the extraction step S52, the matched reference event is matched, and basic information of the event, on-site investigation, and event handling records are obtained. Then, the historical related event records are referred to, and the investigation is carried out in the to-be-investigated area, and the on-site investigation and specific management measures are recorded. Finally, the water quality abnormality early warning event information is integrated, and the event key information is input into the database, including but not limited to event basic information, time sequence fluctuation feature vector, on-site handling record and the like.

[0191] Therefore, the application discards the strong dependence on a complex physical model or a pollution source spectrum library, and instead dynamically constructs a normal fluctuation benchmark based on historical measured data of each monitoring point. The adaptive fluctuation interval based on historical monitoring data is used to replace the static threshold, which can automatically respond to systematic shifts caused by seasonal changes, engineering disturbances or industrial layout adjustments, and ensure that the abnormality determination is always based on the deviation degree relative to the current normal state, rather than the absolute value. In combination with the amplitude coefficient and the fluctuation coefficient, the time sequence fluctuation feature is calculated, which significantly reduces the data acquisition cost and the computing resource demand. Meanwhile, the lag correlation analysis and the fluctuation feature similarity calculation between nodes are introduced, and the time decay weight is fused, which further improves the sensitivity to real abnormal events and the anti-interference ability, and infers the abnormal propagation path without an explicit migration model, to realize the lightweight and data-driven rapid tracing.

[0192] In summary, the application realizes efficient identification and accurate tracing of water quality abnormalities by means of data-driven dynamic modeling, time sequence fluctuation feature extraction, cross-node lag correlation and similarity matching, without relying on high-cost external data, which effectively overcomes the core defects of the prior art in response speed, adaptability and resource dependence.

[0193] For reference Figure 4 The embodiment of the application also provides a watershed water quality abnormality tracing system 1 based on a time sequence fluctuation feature, which comprises a memory 3, a processor 2, and a computer program stored in the memory 3 and executable on the processor 2, and the processor 2 implements the steps of the above-mentioned embodiments when executing the computer program.

[0194] Since the system / device described in the above-mentioned embodiments of the application is the system / device used to implement the method of the above-mentioned embodiments of the application, the specific structure and modifications of the system / device can be understood by those skilled in the art based on the method described in the above-mentioned embodiments of the application, and thus will not be described here. Any system / device used in the method of the above-mentioned embodiments of the application belongs to the scope of protection of the application.

[0195] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, apparatus or computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer readable storage media (including, but not limited to, disk memory, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.

[0196] The embodiments of methods, apparatuses (devices) and computer program products according to the present application can be described with reference to flow charts and / or block diagrams illustrating the architecture, functionality, and operation of implementations of the present application. It will be understood that each block of the flow chart and / or block diagrams, and combinations of blocks in the flow chart and / or block diagrams, can be implemented by computer program instructions. Some embodiments can be implemented in a hardware device or other component.

[0197] It should be noted that any references made herein to elements or components should not be construed as limiting the scope of the claims. The word "comprising" does not exclude the presence of elements or steps other than those listed in a claim. The word "a" or "an" preceding the

[0198] Furthermore, it is to be understood that the use of "a" or "an", "the" or "said" employed throughout the present specification in reference to a particular feature, structure, item, component, material or the like, is not intended to be construed as limiting the scope of the application to only a single feature, structure, item, component, material or the like. Rather, such phrases are used herein to provide a broadest possible scope of the present application to encompass both singular as well as plural structures, items, components, materials and the like.

[0199] While the preferred embodiments of the application have been described, additional variations and modifications can be employed by those skilled in the art. Therefore, the claims should be interpreted as including all such variations and modifications as falling within the scope of the present application.

[0200] It will be apparent to those skilled in the art that various modifications and variations can be made to the present application without departing from the spirit or scope of the application. Thus, it is intended that the present application cover modifications and variations of this application provided they come within the scope of the appended claims and their equivalents.

Claims

1. A method for tracing abnormal water quality in a river basin based on timing fluctuation characteristics, characterized in that, The method comprises the following steps: Step S1, taking each monitoring point in a river basin as a graph node, and constructing a directed edge connecting the graph nodes according to the spatial relationship and water flow direction between the monitoring points to form a river basin graph network; Step S2, obtaining historical monitoring data, extracting amplitude features and fluctuation features from the historical monitoring data, and calculating a time series fluctuation feature vector, comprising: Step S21, obtaining historical monitoring data and performing decyclic preprocessing; Step S22, performing differential calculation on the preprocessed historical monitoring data to obtain an amplitude coefficient; Step S23, collecting historical monitoring data of each index to perform percentile regression fitting to obtain a dynamic percentile regression curve, and calculating a fluctuation coefficient based on the dynamic percentile regression curve and a current monitoring value; Step S24, performing normalization processing on the amplitude coefficient and the fluctuation coefficient; Step S25, calculating a time series fluctuation feature according to the normalized amplitude coefficient and the fluctuation coefficient, and the formula of the fluctuation coefficient is: ; wherein, denotes the volatility coefficient of indicator i at time t, denotes the monitoring value of indicator i at the time t, , , represent the fitted values of the low, medium, and high quantile regression lines of indicator i at time t, respectively. The formula of the time series fluctuation feature is: ; In the formula, denotes the time series fluctuation characteristic value of index i at time t, denotes the amplitude coefficient of index i at the same time t, denotes the fluctuation coefficient of index i at the same time t; Step S3, traversing the graph nodes connected by the directed edges, constructing an index pair corresponding to the graph nodes, for each index pair, screening a maximum correlation coefficient thereof at each lag period, determining a lag period corresponding to the maximum correlation coefficient as a response lag dimension of the index pair, and calculating a response lag dimension coefficient according to the response lag dimension; Step S4, calculating an abnormal correlation coefficient between the graph nodes according to the time series fluctuation feature vector and the response lag dimension coefficient, and constructing an abnormal transmission path, comprising: Step S41, calculating an abnormal fluctuation interval coefficient and a fluctuation feature similarity for the time series fluctuation feature vector of a node pair; Step S42, calculating an abnormal correlation coefficient between each graph node according to the response lag dimension coefficient, the abnormal fluctuation interval coefficient and the fluctuation feature similarity, comprising: Step S421, calculating a time decay coefficient W between each graph node according to the response lag dimension coefficient D and the abnormal fluctuation interval coefficient d: ; In the formula, T is the maximum response time range of the indexes corresponding to two graph nodes; Step S422, calculating an abnormal correlation coefficient R between each graph node according to the fluctuation feature similarity B and the time decay coefficient W: R=W×B; ; wherein cos(·) denotes the cosine similarity, E (1) denotes the long vector after concatenation of the first node, E (2) denotes the long vector after concatenation of the second node, dis(·) denotes the Euclidean distance; Step S43, positioning an abnormal fluctuation excitation node according to the abnormal correlation coefficient between each graph node to construct a complete abnormal transmission path; Step S5, dividing a pollution investigation area according to the abnormal transmission path for a staff to investigate and dispose.

2. The method of claim 1, wherein, The step S41 comprises: Step S411, screening indexes with time series fluctuation feature values exceeding a fluctuation threshold in both nodes as core abnormal indexes, and simplifying the time series fluctuation feature vector group of each graph node according to the core abnormal indexes; Step S412, calculating a difference value of each graph node at the maximum fluctuation time of the time series fluctuation feature vector based on the simplified time series fluctuation feature vector group to obtain a node abnormal fluctuation interval coefficient; Step S413, translating and aligning the time series fluctuation feature vector group of all graph nodes according to the node abnormal fluctuation interval coefficient; Step S414, according to the aligned time sequence fluctuation feature vector group, the fluctuation feature similarity between each graph node is calculated.

3. The method of claim 1, wherein, The step S3 comprises: Step S31, the two graph nodes with a connection edge are constructed as a node pair, and an index pair is constructed according to the monitoring index type of the graph node; Step S32, taking a single data monitoring interval as a lag period, the correlation coefficients of each index pair at different lag periods are calculated within a preset maximum lag period range; Step S33, for each index pair, the maximum correlation coefficient at each lag period is screened, and the lag period corresponding to the maximum correlation coefficient is determined as the response lag dimension of the index pair; Step S34, the response lag dimensions of all index pairs under the same node pair are summarized to obtain the response lag dimension coefficient of the node pair.

4. The method of claim 1 to 3, wherein, The step S1 comprises: When it is a river water quality monitoring point or a hydrological monitoring point, the connection edges between the graph nodes are obtained based on the adjacent relationship between the monitoring points, and the direction of the connection edge is added according to the upstream to downstream mode on the river network; When it is a lake water quality monitoring point, a lake area grid is constructed with the lake water quality monitoring point as the center, the connection edges between the graph nodes are obtained according to the lake area grid, and the direction of the connection edge is added according to the adjacent grid transmission mode, if there is no determined flow direction of the lake area grid, a bidirectional connection edge is constructed; When it is a river node and a lake area node, the connection edge of the river inlet to the adjacent lake area grid is constructed, and the direction of the connection edge is added according to the mode from the river inlet to the adjacent lake area grid; When it is a river outside monitoring point, the connection edge is added according to the monitoring points on the adjacent river sections of the catchment area, forming a transmission link of the river outside variable affecting the adjacent river section monitoring point data to add the direction of the connection edge.

5. The method of claim 1 to 3, wherein, The step S5 comprises: Step S51, the time sequence fluctuation feature vector group of all monitoring points on the abnormal transmission path is obtained, the monitoring point to be investigated is determined according to the maximum value of the time sequence fluctuation feature vector, and the catchment area corresponding to the monitoring point to be investigated is listed as the investigation area; Step S52, each monitoring point of the abnormal transmission path is traversed, the time sequence fluctuation feature vector group is spliced into a long vector according to a fixed index order, a reference event is matched from the historical events of the same type of the monitoring point according to the cosine similarity between the long vectors; Step S53, according to the event record of the reference event, the staff goes to the investigation area for investigation and disposal, and records and arranges the water quality abnormal early warning event information. 6.A system for tracing the source of water quality anomaly in a river basin based on timing fluctuation characteristics, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the water quality abnormality tracing method based on time sequence fluctuation features in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Water quality abnormity tracing method and system based on unmanned aerial vehicle monitoring

    CN115905450A

  • Box-type substation state monitoring and early warning method based on artificial intelligence

    CN120974305A