Environment monitoring data anomaly detection and traceability method and device based on edge calculation

By combining edge computing with a remote backend, the anomaly analysis algorithm and parameters are dynamically matched, solving the problems of flexibility and traceability of environmental monitoring data, and achieving efficient anomaly identification and traceability.

CN122046155APending Publication Date: 2026-05-15BI SHENGYUN (WUHAN) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BI SHENGYUN (WUHAN) INFORMATION TECH CO LTD
Filing Date
2026-02-11
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing methods for detecting anomalies in environmental monitoring data lack flexibility, cannot adapt to environmental changes in different regions and time periods, are prone to false alarms and missed alarms, and cannot accurately trace the causes of anomalies.

Method used

By adopting an edge computing-based approach, environmental monitoring data is acquired in real time through a combination of distributed sensors and an edge computing center with a remote backend. An anomaly analysis algorithm and parameters based on type matching are used to calculate anomaly path data and correlations, and the calculation parameters are dynamically adjusted to achieve multi-dimensional anomaly data tracing.

Benefits of technology

It improves the accuracy of environmental monitoring data identification and traceability, adapts to the analysis needs of different types of data, dynamically adjusts processing methods, and ensures the flexibility and comprehensiveness of anomaly identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122046155A_ABST
    Figure CN122046155A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, and discloses an environment monitoring data anomaly detection and traceability method and device based on edge computing, and the method depends on a connection architecture of a distributed multi-type detection sensor, an edge computing center and a remote background. The edge computing center acquires sensor environment monitoring data in real time, matches corresponding anomaly analysis algorithms and parameters, and outputs anomaly analysis data and geographic information to a remote background; and the remote background extracts multi-sensing dimension abnormal data, calculates an association relationship between an abnormal path and dimensions, adjusts parameter update data when a preset association strategy is met, and further calculates a traceability abnormal path, abnormal amplitude data, traceability abnormal data and traceability geographic information. According to the method, efficient anomaly detection of environment monitoring data is realized through edge calculation and background cooperation, and meanwhile, geographic information corresponding to anomaly causes is accurately traced, so that the method is adaptive to complex monitoring scene requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data analysis, and in particular to a method and apparatus for detecting and tracing anomalies in environmental monitoring data based on edge computing. Background Technology

[0002] With the rapid development of environmental monitoring sensing technology, the types of monitoring sensors have continued to expand, covering multi-dimensional monitoring equipment for air, water quality, soil, noise, and other aspects, resulting in an explosive growth in the amount of monitoring data. Simultaneously, the dimensions of data are constantly expanding, integrating real-time monitoring values, operating parameters, and environmentally related data. The data types possess both structured and unstructured characteristics, significantly increasing the complexity of environmental monitoring data and placing higher demands on the efficiency and accuracy of data processing and anomaly identification.

[0003] Currently, most environmental monitoring data anomaly detection methods employ simple threshold comparisons. These methods compare the data against preset fixed thresholds to determine if the data exceeds normal ranges, thus identifying anomalies. Other scenarios rely on manual analysis, where staff meticulously check and analyze each piece of monitoring data, filtering out anomalies and analyzing their causes. While these methods achieve basic anomaly identification, they are ill-suited to complex and ever-changing environmental monitoring scenarios. Furthermore, manual analysis is time-consuming and labor-intensive, and prone to overlooking potential anomalies due to subjective judgment biases, failing to meet the demands of large-scale, real-time monitoring.

[0004] Existing anomaly analysis methods are generally simple, with fixed threshold patterns lacking flexibility. They cannot adaptively adjust to the natural conditions and environmental fluctuations of different regions and time periods, leading to false alarms and missed alarms. Furthermore, existing technologies can only identify anomalies, failing to accurately trace the causes of abnormal data. It is difficult to pinpoint whether anomalies stem from sensor malfunctions, sudden environmental changes, or human interference, resulting in a lack of targeted anomaly handling. Summary of the Invention

[0005] To improve the ability to identify, adapt, and trace anomalies in environmental monitoring data, this application provides a method and apparatus for detecting and tracing anomalies in environmental monitoring data based on edge computing.

[0006] Firstly, this application provides a method for detecting and tracing anomalies in environmental monitoring data based on edge computing, employing the following technical solution: An edge computing-based method for detecting and tracing anomalies in environmental monitoring data, utilizing a distributed array of detection sensors (including multiple types) connected to an edge computing center, which in turn connects to a remote backend, includes the following steps: The edge computing center acquires environmental monitoring data generated by the detection sensors in real time. Based on the type of environmental monitoring data, it matches the corresponding anomaly analysis algorithm and anomaly calculation parameters from a preset type database. Based on the environmental monitoring data and anomaly calculation parameters, the anomaly analysis algorithm outputs anomaly analysis data, extracts the geographic information of the detection sensors corresponding to the anomaly analysis data, and sends the anomaly analysis data and geographic information to the remote backend. The system remotely acquires multiple anomaly analysis data sets and extracts anomaly data from multiple sensor dimensions. Based on the anomaly data of each sensor dimension and geographic information, it calculates the anomaly path data for each dimension and calculates the correlation between the anomaly path data of multiple dimensions. If the correlation matches the preset correlation strategy, it calculates the matching value between the correlation and the correlation strategy, adjusts the anomaly calculation parameters according to the matching value to output more anomaly analysis data, and updates the anomaly path data based on the latest anomaly analysis data. The source abnormal path is calculated based on multiple abnormal path data and their correlations; the abnormal magnitude data corresponding to the path is calculated based on the abnormal path data and abnormal analysis data; the source abnormal data is calculated based on the source path and abnormal magnitude data; and the source geographical information is calculated based on the source abnormal path and source abnormal data.

[0007] By adopting the above technical solution, the edge computing center acquires environmental monitoring data generated by detection sensors in real time. Based on the type of environmental monitoring data, it matches the corresponding anomaly analysis algorithm and anomaly calculation parameters from the type database, so that the anomaly identification process matches the characteristics of different types of data. The remote backend extracts sensor dimension anomaly data, calculates anomaly path data in combination with geographic information, mines the correlation between multi-dimensional anomaly path data, adjusts the anomaly calculation parameters according to the matching value of the correlation relationship and correlation strategy, and outputs more anomaly analysis data, thereby updating and obtaining more comprehensive anomaly path data. Then, by tracing the source anomaly path and anomaly magnitude data, it derives the source anomaly data and source geographic information, which is conducive to fully restoring the propagation and generation trajectory of the anomaly.

[0008] Furthermore, the step of matching the anomaly analysis algorithm and anomaly calculation parameters corresponding to the type of environmental monitoring data from a preset type database also includes the following steps: The type of environmental monitoring data is extracted based on the sensor type, and the environmental monitoring data is divided into multiple groups according to the sensor type to obtain environmental monitoring groups; Based on the sensor type, the corresponding anomaly analysis algorithm and anomaly calculation parameters are obtained from the type database. The anomaly analysis algorithm includes a first analysis algorithm, a second analysis algorithm, and a third analysis algorithm. The anomaly calculation parameters include a first anomaly value, a second anomaly range, and a third anomaly basis. The first analysis algorithm uses the first outlier to filter elements in the environmental monitoring group of the same sensor type; The second analysis algorithm filters and averages the elements in the environmental monitoring group of the same sensor type, and then extracts them through the second anomaly range. The third analysis algorithm uses the third outlier to dynamically calculate the comparison image of the elements in the environmental monitoring group of the same sensor type, and identifies the image based on the preset template comparison.

[0009] By adopting the above technical solution, environmental monitoring data is divided into environmental monitoring groups according to sensor type. Dedicated anomaly analysis algorithms and anomaly calculation parameters are matched for different sensor types, so that the analysis logic is deeply matched with the monitoring characteristics of the sensor types, and the analysis is targeted for different types of data. The three types of algorithms are combined with the corresponding outliers and anomaly ranges to carry out screening, filtering and averaging extraction, and dynamic calculation of image recognition operations, which are adapted to the analysis needs of different data such as pollution, gas, and large object monitoring. The data processing method is dynamically adjusted, which helps to improve the scenario adaptability of anomaly identification and improve the accuracy of anomaly identification of environmental monitoring group data.

[0010] Furthermore, the step of using anomaly analysis algorithms to output anomaly analysis data based on environmental monitoring data and anomaly calculation parameters also includes the following steps: Based on the first environmental monitoring group with the first type of sensing, the abnormal trend of the values ​​in the first environmental monitoring group is obtained, and the abnormal trend is verified to match the analysis trend of the first analysis algorithm. If the verification of the abnormal trend and the analysis trend fails, the values ​​in the first environmental monitoring group are converted into values ​​that conform to the analysis trend, and then the values ​​that are greater than or less than the first abnormal value are selected as the first abnormal data. Based on the second environmental monitoring group with the second type of sensing, the second analysis algorithm is used to perform de-straining filtering on the data in the second environmental monitoring group based on a preset filtering threshold, and to perform averaging calculation based on a preset average window, and to extract the values ​​within the second anomaly range as the second anomaly data. Based on the third environmental monitoring group with the third sensing type being image type, the difference between the pixels of the current frame and the previous frame in the third environmental monitoring group is calculated. If the difference between the pixels is greater than the third anomaly baseline, the pixels of the current frame corresponding to the difference are saved into the third anomaly data. The first, second, and third abnormal data are merged into anomaly analysis data.

[0011] By adopting the above technical solutions, through differential operations such as verifying and correcting abnormal trends, de-straining filtering and averaging calculation, and comparing image pixel differences, the abnormal data extraction can be accurately matched with the characteristics of various types of sensor data, and dynamically adapted to the abnormal features of different types of data; at the same time, merging and integrating various types of abnormal data is beneficial to improving the completeness of abnormal analysis data.

[0012] Furthermore, the method also includes the following steps: The quantity of values ​​in the first abnormal data is obtained as the first abnormal number, and the growth rate of the first abnormal number is calculated as the first growth rate. The filtering threshold or the width of the average window is adjusted according to the negative correlation of the growth rate.

[0013] By adopting the above technical solutions, the filtering threshold or average window width can be dynamically adjusted to adapt to the increase and change of the number of first anomalies, thereby improving the adaptability of data processing.

[0014] Furthermore, the method also includes the following steps: The quantity of values ​​obtained from the second abnormal data is the second abnormality quantity. The growth rate of the second abnormality quantity is the second growth rate. The comprehensive growth rate is calculated by weighting the first growth rate and the second growth rate. The third abnormality basis is adjusted according to the positive correlation of the comprehensive growth rate.

[0015] By adopting the above technical solution, the comprehensive growth rate is obtained by weighting the first growth rate and the second growth rate, and the third abnormal basis is dynamically adjusted. This helps to improve the real-time adaptability of the third abnormal data, match the abnormal development trend, and improve the dynamic effect of abnormal identification.

[0016] Furthermore, the method also includes the following steps: Calculate the number of pixels in each frame of the third abnormal data, and record the number of pixels corresponding to all frames as abnormal pixel data. The abnormal image data is converted into abnormal spectrum data in the frequency domain using a frequency domain conversion algorithm; Obtain the intensity value corresponding to the frequency within the preset frequency range in the abnormal spectrum data, and adjust the first abnormal value according to the negative correlation of the intensity value.

[0017] By adopting the above technical solution, the first abnormal value is adjusted by processing the relevant features of the third abnormal data, and the changes in abnormal features are dynamically adapted, which helps to maintain the matching between the abnormal identification logic and the actual abnormality.

[0018] Furthermore, the step of calculating the correlation between abnormal path data across multiple dimensions also includes the following steps: The similarity between the first type of abnormal path data and the second type of abnormal path data is calculated as the first similarity value, and the overlap between the first type of abnormal path data and the second type of abnormal path data is calculated as the first overlap degree; the first correlation value is calculated based on the first similarity value and the first overlap degree. The similarity between the first type of abnormal path data and the third type of abnormal path data is calculated as the second similarity value, and the overlap between the first type of abnormal path data and the third type of abnormal path data is calculated as the second overlap degree; the second correlation value is calculated based on the second similarity value and the second overlap degree. The similarity between the second type of abnormal path data and the third type of abnormal path data is calculated as the third similarity value, and the overlap between the second type of abnormal path data and the third type of abnormal path data is calculated as the third overlap degree; the third correlation value is calculated based on the third similarity value and the third overlap degree. The corresponding relationships are matched from the preset association database based on the first, second, and third relevance values.

[0019] By adopting the above technical solution, the similarity and overlap of different types of abnormal path data are quantified, and various correlation values ​​are calculated to match the corresponding association relationships. The differences in the characteristics of multi-dimensional abnormal paths are dynamically adapted to maintain the matching between the association relationships and the actual distribution of anomalies, which helps to improve the rationality of multi-dimensional abnormal path association analysis.

[0020] Furthermore, if the association relationship conforms to the preset association strategy, the step of calculating the matching value between the association relationship and the association strategy also includes the following steps: The path type-related value is calculated based on the association relationship, and the association strategy is the path type-related range. If the path type-related values ​​are within the path type-related range, then the association relationship conforms to the association strategy. Calculate the difference between the path type-related value and the maximum value in the path type-related range, and calculate the percentage of the position of the difference within the path type-related range as the matching value; among them, the minimum value in the path type-related range has the smallest percentage of position, and the maximum value in the path type-related range has the largest percentage of position.

[0021] By adopting the above technical solution, the relevant values ​​of path types are quantified and compared with the relevant range of association strategies, clarifying the conditions for meeting the association relationship, accurately calculating the matching value, and dynamically adapting to the range standard of the association strategy; this helps to improve the rationality of association judgment and enhance the adaptive ability of the matching judgment between association relationship and association strategy.

[0022] Furthermore, the step of calculating the source anomaly path based on multiple anomaly path data and their correlations also includes the following steps: Calculate the close or similar portions of the abnormal path data of the first, second, and third types; The part that is close or similar between the first type and the second type is taken as the first intermediate path data, and the road segment corresponding to the first correlation value in the first intermediate path data is taken as the first path data segment. The closest part between the second type and the third type is taken as the second intermediate path data, and the road segment in the second intermediate path data that corresponds to the second correlation value is taken as the second path data. The closest part between the first type and the third type is taken as the third intermediate path data, and the road segment in the third intermediate path data that corresponds to the third relevant value is taken as the third path data. The source of the abnormal path is derived by fitting the first, second, and third path data.

[0023] By adopting the above technical solution, the relevant parts of multiple types of abnormal paths are integrated, and the corresponding road segments are selected to fit the source abnormal path in combination with various relevant values, so as to maintain the adaptability of the source path to multi-dimensional abnormal features; so that the source abnormal path can better match the actual propagation trajectory of the abnormality, dynamically adapt to the association features of different types of abnormal paths, and help improve the rationality of the source path.

[0024] Secondly, this application provides an environmental monitoring data anomaly detection and tracing device based on edge computing, which adopts the following technical solution: An edge computing-based environmental monitoring data anomaly detection and tracing device includes a processor, wherein the processor executes the steps of the edge computing-based environmental monitoring data anomaly detection and tracing method as described in any one of the preceding claims. Attached Figure Description

[0025] Figure 1 This is a flowchart illustrating the steps of an environmental monitoring data anomaly detection and tracing method based on edge computing.

[0026] Figure 2 This is a flowchart illustrating the steps of matching the anomaly analysis algorithm and anomaly calculation parameters corresponding to the type of environmental monitoring data from a pre-set type database. Detailed Implementation

[0027] The embodiments of this application are described in detail below, and examples of the embodiments are shown in the accompanying drawings.

[0028] This application discloses an edge computing-based method for anomaly detection and source tracing in environmental monitoring data. It is applicable to multi-dimensional, large-scale environmental monitoring scenarios, such as cross-regional air quality monitoring, watershed water quality monitoring, and multi-pollutant monitoring in industrial parks. The method leverages the real-time processing capabilities of edge computing to meet the anomaly identification needs of different types of monitoring data, and combines multi-dimensional data correlation analysis from a remote backend to achieve anomaly source tracing. (Refer to...) Figure 1 The specific method is as follows: System deployment and initial configuration: A distributed deployment of various detection sensors, including air, water, soil, and noise sensors, is implemented. Each sensor has a built-in GPS positioning module to obtain accurate geographic information and is pre-defined with a unique sensor ID and data type identifier (e.g., "Air-PM2.5", "Water-pH"). Each sensor connects to the nearest edge computing center via 5G / NB-IoT / LoRa communication protocols. The edge computing center communicates with the remote backend (cloud server cluster) via fiber optic or 5G private network, ensuring data transmission latency ≤50ms. Simultaneously, the edge computing center has a pre-defined type database, indexed by "data type," which stores corresponding anomaly analysis algorithms and parameters: time-series data such as PM2.5 is matched with the LSTM algorithm (30-minute time window, threshold coefficient 1.2, 500 training iterations); stable data such as pH values ​​are matched with the 3σ algorithm (99.7% confidence interval, 10 data points sliding window); and unstructured data such as noise spectra are matched with the CNN algorithm (3×3 convolutional kernel, 2×2 pooling window, 64-dimensional feature extraction). This database supports remote backend updates of algorithms and parameters.

[0029] The first step involves the edge computing center generating and transmitting anomaly analysis data. The edge computing center acquires environmental monitoring data uploaded by various sensors in real time at a preset frequency (1 time / second to 1 time / minute, dynamically adjusted according to sensor type, e.g., 1 time / 10 seconds for air sensors, 1 time / 5 minutes for soil sensors). This data includes data values, data type identifiers, sensor IDs, collection timestamps, and raw GPS geographic information. The edge computing center extracts the data type identifiers, precisely matches them in the type database, and calls the corresponding anomaly analysis algorithm and parameters. It then substitutes the data values ​​and collection timestamps into the algorithm and combines them with the parameters to determine anomalies: When using the 3σ algorithm, it calculates the mean μ and standard deviation σ of the last 10 data points; data values ​​exceeding [μ-3σ, μ+3σ] are considered anomalies. When using the LSTM algorithm, it predicts real-time data through a trained model; deviations exceeding "threshold coefficient × historical average deviation" generate anomaly analysis data (including anomaly markers, anomaly values, occurrence time, and sensor IDs). Meanwhile, the edge computing center extracts the corresponding geographical information (latitude and longitude, installation location) based on the sensor ID, binds it with the anomaly analysis data, and sends it to the remote backend through encryption protocols such as TLS1.3 to ensure transmission security.

[0030] The second step involves the remote backend performing anomaly path data calculation and parameter adjustment: The remote backend receives anomaly analysis data and geographic information sent by the edge computing center, sorts it by sensor ID and collection timestamp, and filters out time-series continuous anomaly data within the same monitoring area (time interval ≤ 5 minutes, such as data from the most recent hour). Multiple sensor dimension anomaly data are then separated by data type identifier (in this example, air pollutant concentration, water quality indicators, and noise / soil categories; each dimension retains anomaly values, occurrence time, and geographic coordinates). The remote backend calls GIS tools and, based on the geographic coordinates and occurrence time of the anomaly data for each dimension, uses the Dijkstra shortest path algorithm combined with time-series weighted path fitting algorithms to calculate the anomaly path data for each dimension (including node coordinates, distance between nodes, propagation speed, and anomaly values ​​at each node). The remote backend uses a Pearson correlation coefficient analysis model to mine the correlation relationships of multi-dimensional anomaly paths, specifically including temporal correlation (node ​​time difference ≤ 15 minutes), spatial correlation (node ​​distance ≤ 500 meters), and numerical correlation (correlation coefficient of change trend ≥ 0.7).

[0031] The preset association strategy is "satisfying at least two of the above association relationships". If this strategy is met, a matching value in the range of 0 to 1 is calculated by weighting (time 0.4, space 0.3, numerical 0.3) (e.g., a matching value of 0.7 is obtained when time and space are associated). The anomaly calculation parameters of the edge computing center are adjusted based on the matching value: when the matching value is ≥0.8, the recognition range is expanded (e.g., the 3σ confidence interval is adjusted to 95%, and the LSTM threshold coefficient is adjusted to 1.0); the parameters remain unchanged between 0.5 and 0.8; and when the matching value is <0.5, the recognition range is narrowed (e.g., the 3σ confidence interval is adjusted to 99.9%). The remote backend sends the adjusted parameters to the edge computing center, which regenerates the anomaly data with the new parameters and provides feedback. The remote backend then updates the anomaly path data for each dimension accordingly.

[0032] The third step involves the remote backend performing anomaly tracing calculations: Based on the updated multi-dimensional anomaly path data and their relationships, the remote backend uses a tracing path fusion algorithm to calculate the anomaly tracing path: common nodes (appearing in at least two dimensions) are extracted from the anomaly paths of each dimension, arranged in reverse chronological order of anomaly occurrence time, and combined with the propagation speed to generate a tracing path from the end of the diffusion to the starting point (e.g., the original path A→B→C, the tracing path is C→B→A); if there are no common nodes, the dimension path with the highest matching value is selected as the basis, and the tracing path is derived by combining the spatial distribution of other dimensions.

[0033] Based on the outliers of each node and the corresponding normal thresholds preset by the remote backend, the outlier magnitude data is calculated: when the outlier is higher than the upper limit, the outlier magnitude = (outlier value - upper limit of normal threshold) / upper limit of normal threshold; when it is lower than the lower limit, the outlier magnitude = (lower limit of normal threshold - outlier value) / lower limit of normal threshold. Each node corresponds to an outlier magnitude value, forming a sequence. Combining the source tracing path with this sequence, source tracing anomaly data is calculated. The outlier data of the node with the largest outlier magnitude is selected as the core, and the average outlier magnitude and the maximum outlier duration of all nodes are integrated to form complete source tracing anomaly data.

[0034] Based on the coordinates of the starting node of the tracing path and the description of the sensor installation location, the tracing geographic information, including precise latitude and longitude, regional affiliation, and surrounding environmental characteristics, is calculated to complete the geographic tracing. Through the above steps, the edge computing center achieves adaptive anomaly identification for different types of monitoring data. The remote backend completes anomaly tracing and dynamically adjusts parameters to ensure the flexibility and comprehensiveness of identification and accurately locate the source of the anomaly. In this embodiment, the detailed steps of the edge computing center algorithm matching and execution are as follows: After acquiring the monitoring data, the sensor type in the data type identifier is first extracted. This type is divided into "monitoring object - core parameter", including air, water quality, soil, noise, and vision (each category includes specific core parameters and corresponding sensor types).

[0035] The edge computing center groups the data according to the type of sensor. It generates environmental monitoring groups by dual clustering of sensor type identifier and sensor ID to ensure that the data in the same group comes from the same core parameter collected by the same type of sensor device. Each group of data is arranged in ascending order by the collection timestamp to form a time-series queue. Each data element in the queue contains core fields such as data value (or image data), collection time, sensor ID, and GPS geographic information.

[0036] After grouping, the edge computing center accurately matches the corresponding anomaly analysis algorithm and parameters in a pre-defined type database based on the sensor type. The database pre-stores a mapping table of "sensor type-algorithm-parameter". The anomaly analysis algorithm is divided into first, second, and third analysis algorithms, and the anomaly calculation parameters are divided into first anomaly value, second anomaly range, and third anomaly basis. The specific mapping rules and execution steps are as follows: 1. Adaptation and Execution of the First Analysis Algorithm and the First Outlier (Adaptation to Time-Series Numerical Data such as Pollution Concentration and Water Quantity): For sensor types with "absolute value exceeding the standard" as the core anomaly characteristic (such as water quality-pollution concentration, water quality-water quantity, etc.), the first analysis algorithm and the first outlier are matched; the first outlier is based on national / industry standards and regional baseline presets, and is precisely configured according to the sensor type (such as COD 50mg / L, river flow bidirectional threshold 10~100m³ / s, lead 350mg / kg). Execution steps: The first analysis algorithm is called, and data elements in the same group are filtered in turn. Data values ​​are extracted and their preset relationship with the first outlier is judged. If the relationship is satisfied, it is marked as "first-level candidate outlier data" and the basis is recorded. If the relationship is not satisfied, it is marked as normal data. Finally, a list of first-level candidate outlier data is output.

[0037] 2. Adaptation and Execution of the Second Analysis Algorithm and the Second Anomaly Range (Adapted to Fluctuating Time-Series Numerical Data such as Gas Content and Air Volume): For sensor types whose values ​​are easily affected by environmental disturbances and require noise reduction, a second analysis algorithm and a second anomaly range (such as a 95% confidence interval) based on historical fluctuations are matched. The core of the algorithm consists of three steps: "filtering and denoising - averaging - anomaly extraction": The first step is filtering and denoising. Gas data is filtered using a moving average (15 data points in a window), and air volume data is filtered using a Kalman filter (preset state, observation equation, and noise variance Q=0.01, R=0.1); the second step is averaging. The arithmetic mean is calculated in segments with a 5-minute period to obtain an averaged data sequence; the third step is anomaly extraction. The averaged data is compared with the second anomaly range. If the data exceeds the range, it is marked as "secondary candidate anomaly data" and key parameters are recorded. Finally, a list of secondary candidate anomaly data is output.

[0038] 3. Adaptation and Execution of the Third Analysis Algorithm and the Third Anomaly Foundation (Adaptation to Visual Data such as Large Solid Objects and Collapses): For image / video sequence data acquired by dynamic cameras, the third analysis algorithm and the third anomaly foundation (including image preprocessing parameters, dynamic calculation parameters, and template library) are matched. Specifically, for visual data involving large solid objects, the corresponding parameters are 1920×1080 resolution, 25fps frame rate, etc., and a corresponding template library is provided; for visual data involving terrain changes, the corresponding parameters are 2560×1440 resolution, 15fps frame rate, etc., and a corresponding template library is provided. The template library supports remote, periodic updates.

[0039] The specific execution steps of the third analysis algorithm are as follows: First, image preprocessing: The edge computing center receives the video stream data uploaded by the dynamic camera, scales each frame of the image according to the resolution in the third anomaly basis, converts it into a grayscale image (grayscale value = 0.299×R + 0.587×G + 0.114×B), and then performs binarization processing through the grayscale threshold (grayscale value > threshold is set to 255, otherwise it is 0), to obtain a binarized image sequence.

[0040] The second step is to dynamically calculate the comparison image: call the frame difference method module, select the reference frame and the current frame according to the preset frame difference interval, calculate the pixel difference between the two binarized images and count the proportion; if the proportion of difference pixels is greater than or equal to the foreground extraction threshold, extract the difference region as the foreground image, superimpose the foreground images of 5 consecutive sets of frame differences, and generate a comparison image containing the location, area and contour of the abnormal region (stored as a grayscale image, with the pixel grayscale value corresponding to the cumulative number of differences).

[0041] The third step is template comparison and recognition: A corresponding normal scene template is retrieved from the template library. The SSIM algorithm is used to calculate the similarity between the compared image and the template. The SSIM formula is (2μxμy+C1)(2σxy+C2) / [(μx²+μy²+C1)(σx²+σy²+C2)] (where μx and μy are the grayscale mean, σx and σy are the standard deviations, σxy is the covariance, C1=(0.01×255)², C2=(0.03×255)²). A preset similarity threshold of 0.7 is used. If SSIM ≤ 0.7, an anomaly is determined, and the corresponding image frame and compared image are marked as "Level 3 candidate anomaly data." The similarity and coordinates of the anomaly region are recorded, and finally, a list of Level 3 candidate anomaly data is output.

[0042] After the three types of algorithms are executed, the edge computing center merges the first-level, second-level, and third-level candidate anomaly data to form complete anomaly analysis data; if there is no candidate anomaly data, it is marked as "no anomaly". Then, the corresponding sensor geographic information is extracted and sent to the remote backend. The subsequent process is the same as described above.

[0043] In this embodiment, after the edge computing center completes the grouping of environmental monitoring data and the matching of anomaly analysis algorithms and calculation parameters, it performs differentiated anomaly data extraction processes for the first type, second type, and third type of sensor data, as follows: I. Verification, correction, and screening of anomalous trends in the first type of environmental monitoring group: The first type of sensor data corresponds to the sensor types mentioned above, such as "water quality - pollution concentration (COD, ammonia nitrogen)," "water quality - water quantity (river flow, pipe network flow)," and "soil - heavy metal content (lead, cadmium)," which are characterized by "numerical data, strong time-series correlation, and predictable abnormal trends." The corresponding first environmental monitoring group is a time-series data queue sorted by timestamp under the same sensor type, such as 360 data points collected by a COD sensor at 1 time / 10 seconds per hour.

[0044] 1. Anomaly trend extraction and analysis: trend verification and edge detection: The computing center invokes the trend analysis module of the first analysis algorithm to extract abnormal trends based on the time-series data of the first environmental monitoring group: First, a trend analysis time window is set (obtained from the type database, default 30 minutes, corresponding to 36 data points). A linear fitting algorithm is used to calculate the trend slope k of the data within the window, with the formula as follows: Where n is the number of data points within the window, t_i is the timestamp of the i-th data point, converted to relative time (e.g., t1=0 for the first data point, t2=10 seconds for the second, and so on), and x_i is the monitoring value of the i-th data point. Anomaly trend types are defined based on the trend slope k: k≥0.05 (unit: monitoring value / second) represents a "monotonically increasing trend," k≤-0.05 represents a "monotonically decreasing trend," and |k|<0.05 represents a "stable fluctuation trend." Simultaneously, the pre-set analysis trend of the first analysis algorithm is obtained from the type database; this analysis trend represents typical anomalies for the first type of sensor data, such as the analysis trend for "water quality - pollution concentration" being a "monotonically increasing trend," the analysis trend for "river flow" being a "sudden increase / decrease trend," and the analysis trend for "soil moisture content" being a "stable fluctuation trend."

[0045] The trend verification rule is as follows: if the extracted abnormal trend is consistent with the analysis trend (e.g., the abnormal trend of pollution concentration data is monotonically increasing, which matches the analysis trend), then the verification is successful; if they are inconsistent (e.g., the abnormal trend of pollution concentration data is drastic fluctuation, which conflicts with the preset monotonically increasing analysis trend), then the verification fails.

[0046] 2. Data trend correction: If trend verification fails, the edge computing center calls the data correction module to convert the values ​​in the first environmental monitoring group into values ​​that conform to the analysis trend. The specific correction process is as follows: ① Based on the analysis trend type, construct a trend fitting function (e.g., when the analysis trend is monotonically increasing, the fitting function is set as a linear increasing function y=kt+b, where k is the preset trend slope obtained from the type database, such as k=0.03mg / (L·s) for COD data; b is the initial value, taken from the value of the first data point in the window); ② Calculate the deviation between the original value of each data point and the predicted value of the fitting function: Δx=x_i-(kt_i+b) ③ If the absolute value of the deviation |Δx| ≤ the preset deviation threshold (obtained from the type database, such as 5mg / L for COD data), the original value is retained; if |Δx| > the deviation threshold, the value of the data point is corrected to x_i'=kt_i+b+sign(Δx)×deviation threshold, where sign(Δx) is the deviation sign, ensuring that the corrected data still conforms to the original fluctuation direction and the analysis trend; ④ The trend extraction is re-executed on the corrected data queue until the abnormal trend and the analysis trend are verified, with a maximum of 3 iterations. If it still fails, it is marked as "trend abnormal and pending verification" and included in the candidate abnormal data.

[0047] 3. First step in filtering out abnormal data: After trend verification and correction are completed, the edge computing center calls the filtering module of the first analysis algorithm to filter the corrected data of the first environmental monitoring group based on the first outlier: if the first outlier is a one-way threshold (such as 50 mg / L for COD), then all data points x_i' > the first outlier are filtered out and marked as first outlier data; if the first outlier is a two-way threshold (such as 10 m³ / s to 100 m³ / s for river flow), then all data points x_i' < the lower limit of the first outlier or x_i' > the upper limit of the first outlier are filtered out and marked as first outlier data; each first outlier data includes the original value, the corrected value, the trend type, the verification result, the filtering basis (such as "COD corrected value 68 mg / L > first outlier 50 mg / L"), the collection timestamp, the sensor ID, and the geographic information.

[0048] II. Deburring Filtering, Averaging Calculation, and Outlier Data Extraction for the Second Type of Environmental Monitoring Group: It is clear that the second type of sensor data corresponds to the sensor types mentioned above, such as "air-gas content (SO2, VOCs)", "ventilation network-air volume" and "air-particulate matter concentration (high frequency acquisition)", which are "numerical values ​​that are easily fluctuating and contain instantaneous spike noise". The second environmental monitoring group is a time-series data queue of the same sensor type, such as 720 data points collected by a certain SO2 sensor at 1 time / 5 seconds per hour.

[0049] 1. De-aliasing filtering: The edge computing center calls the de-scratching filtering module of the second analysis algorithm. First, it obtains the preset filtering threshold from the type database. This threshold is set based on the historical normal fluctuation characteristics of the sensor data. The calculation formula is: Filtering threshold = μ + 3σ, where μ is the mean of normal data in the same period of the past 7 days, and σ is the standard deviation. For example, for the SO2 sensor, μ = 0.08 mg / m³, σ = 0.02 mg / m³, then the filtering threshold = 0.08 + 3 × 0.02 = 0.14 mg / m³; for the air volume sensor, μ = 30 m³ / h, σ = 3 m³ / h, then the filtering threshold = 30 + 3 × 3 = 39 m³ / h (lower limit filtering threshold = 30 - 3 × 3 = 21 m³ / h).

[0050] The specific steps for de-straining filtering are as follows: ① Iterate through each data point x_j in the second environmental monitoring group, where j is the data point index; ② If it is a one-way filtering threshold, such as SO2 data, only excessively high spikes need to be removed, then determine if x_j > the filtering threshold: if so, replace the data point with the arithmetic mean of the previous and next data points, i.e., x_j' = (x_{j-1} + x_{j+1}) / 2. If j = 1, then take x2; if j is the last data point, then take x_{j}. -1}); if not, retain the original data x_j'=x_j; ③ If it is a two-way filtering threshold, such as air volume data, it is necessary to remove excessively high or low spikes, then determine whether x_j < lower limit filtering threshold or x_j > upper limit filtering threshold: if yes, process according to the above mean replacement rule; if no, retain the original data; ④ Repeat the traversal once to avoid continuous spikes not being processed, and obtain the stable data sequence after spike removal [x_1',x_2',...,x_j',...].

[0051] 2. Averaging calculation: Based on the de-strained stationary data sequence, the averaging module of the second analysis algorithm is invoked. Averaging calculations are performed according to the preset averaging window (i.e., a smoothing filter window, matched to the sensor acquisition frequency; for example, if the SO2 sensor acquisition frequency is 1 time / 5 seconds, the averaging window is set to 20 data points, corresponding to a 100-second duration) of the type database. ① Divide the stationary data series into segments according to the average window length; merge the last segment with the previous segment if it has fewer than 20 data points. ② Calculate the average value of each segment: ③ The averaged data sequence is obtained, where m is the segment index and m ≥ 1. Each averaged value corresponds to a time interval, such as the first segment corresponding to 0~100 seconds, and the second segment corresponding to 101~200 seconds.

[0052] 3. Extraction of the second type of outlier data: Obtain the second anomaly range from the type database, such as the second anomaly range for SO2 being [0.15 mg / m³, +∞), and the second anomaly range for air volume being [5 m³ / h, 50 m³ / h]. Average each data sequence... Compare with the second abnormal range: If Exceeding the second abnormal range (such as SO2) =0.17mg / m³>0.15mg / m³), air volume ( =4.2m³ / h<5m³ / h), then extract the original data points (x_j' after deburring) and averaged values ​​corresponding to this segment. The time intervals are segmented and marked as the second abnormal data. The second abnormal data need to record the filtering parameters (filtering threshold, number of replacements), averaging parameters (window size, segment index), and the basis for anomaly judgment (such as "averaged value 0.17mg / m³∈second abnormal range[0.15mg / m³,+∞)").

[0053] III. Calculation of inter-frame pixel difference and extraction of abnormal data for the third type of environmental monitoring group: To clarify: The third type of sensor data corresponds to image data such as "visual-large solid objects" and "visual-terrain changes" based on dynamic cameras (following the principle of event cameras, only capturing the inter-frame change area). The third environmental monitoring group is a continuous sequence of image frames (such as 300 frames of images captured by a dynamic camera at 15fps, with frame numbers F1~F300), and the image format is RGB 24-bit or grayscale 8-bit.

[0054] 1. Image preprocessing and inter-frame pairing: The edge computing center calls the image processing module of the third analysis algorithm to preprocess the image frames of the third environmental monitoring group: all image frames are uniformly converted into grayscale images (if they are RGB images, they are converted according to the formula (grayscale value G=0.299R+0.587G+0.114B)), and the resolution is kept consistent with the third anomaly baseline (e.g., 1920×1080) to avoid errors in the difference calculation caused by size differences; then they are paired according to the rule of "current frame - previous frame", such as F2 paired with F1, F3 paired with F2, ..., Fn paired with Fn-1, where n is the total number of frames.

[0055] 2. Calculation of inter-frame pixel difference: For each pair of paired frames (let the current frame be Fn and the previous frame be Fn-1), calculate the difference between the corresponding pixels: If it is a grayscale image, the pixel difference is (D(p,q)=|G_n(p,q)-G_{n-1}(p,q)|), where (p,q) are the coordinates of the pixel (p∈[0,1919], q∈[0,1079]), Gn(p,q) is the grayscale value of the current frame Fn at coordinate (p,q), and G_{n-1}(p,q) is the grayscale value of the previous frame Fn-1 at coordinate (p,q); if it is an RGB image (not converted to grayscale), the pixel difference is D(p,q)=sqrt{ (R_n-R_{n-1})^2+(G_n-G_{n-1})^2+(B_n-B_{n-1})^2, where Rn, Gn, and Bn are the RGB channel values ​​of the current frame at that coordinate; the third anomaly base (i.e., pixel difference threshold) is obtained from the type database, and this threshold is adjusted according to the image monitoring scenario: for example, the grayscale difference threshold for the "large solid object monitoring" scenario is set to 30 (corresponding to obvious object occlusion), the grayscale difference threshold for the "terrain collapse monitoring" scenario is set to 25 (corresponding to changes in terrain contours), and the difference threshold for RGB images is set to 50 (corresponding to significant changes in color / brightness).

[0056] 3. Saving of third-degree abnormal data: Iterate through the differences D(p,q) of all pixels. If (D(p,q) > third anomaly base), then the pixel is determined to be an anomaly pixel, and its relevant information is saved as third anomaly data: Basic information: current frame number Fn, previous frame number Fn-1, pixel coordinates (p,q), pixel difference D(p,q); Auxiliary information: anomaly region marker (clustering consecutive anomaly pixels into anomaly regions, recording the region boundary coordinates (p1,q1) ~ (p2,q2), region area S = (p2-p1+1) × (q2-q1+1) pixels), image acquisition timestamp, sensor ID, and geographic information; If the number of anomaly pixels in a pair of frames accounts for less than 0.01% of the total number of pixels (to avoid misjudgment due to small noise), then third anomaly data is not generated for that frame.

[0057] IV. Merging (Packaging) of Anomaly Analysis Data: After the edge computing center extracts the first, second, and third anomaly data, it calls the data integration module to perform a merging and packaging operation. The specific process is as follows: 1. Data Format Standardization: Convert the three types of abnormal data into JSON format. Fields include "Abnormal Data Type" (First / Second / Third), "Sensor ID", "Sensor Type", "Collection Timestamp", "Geographic Information" (Latitude and Longitude, Installation Location), "Raw Data", "Processed Data" (Correction Value / Filtered Value / Averaged Value / Pixel Difference), "Abnormal Judgment Basis", and "Related Parameters" (such as Algorithm Type, Abnormal Value / Range / Basis, Filter Window, Average Window, etc.). 2. Data Deduplication: If the same sensor generates multiple types of abnormal data at the same timestamp (error ≤ 1 second) (e.g., a monitoring point simultaneously collects water quality COD (Type 1) and SO2 (Type 2) data, and all are abnormal at the same timestamp), then all data are retained and marked as "multi-dimensional collaborative anomaly"; if the same type of data from the same sensor is repeatedly marked as abnormal at 3 consecutive timestamps, then it is merged into one abnormal record, and the duration of the anomaly is recorded (e.g., "2024-05-20 10:00:00~10:00:20, lasting 20 seconds"). 3. Packaging and Transmission: The standardized and deduplicated abnormal data is sorted by "sensor ID + timestamp" to generate an abnormal analysis data packet (compressed using LZ77 compression algorithm to ensure transmission efficiency), and a data check code (such as CRC32 check to avoid transmission damage) is added. Then, it is sent to the remote backend according to the process described above.

[0058] After extracting the first, second, and third types of anomaly data, the edge computing center dynamically adjusts the filtering threshold / average window for the second type and the third type of anomaly base based on the growth trends of the two types of anomaly data. This ensures that the anomaly identification parameters are accurately adapted to the real-time anomaly development trend. The specific steps are as follows: I. Dynamic Adjustment of Filter Threshold or Average Window Based on the First Growth Rate: 1. Statistics of the First Anomaly Quantity and Calculation of the First Growth Rate. First, define the statistical period: Adapt to the average window of the second type of data and the trend analysis window of the first type of data mentioned above. The preset statistical period is 10 minutes (configurable through the type database, ranging from 5 to 30 minutes), that is, this adjustment process is executed every 10 minutes. The edge computing center calculates the first growth rate according to the following steps: ① Extract the first abnormal data within the current statistical period (denoted as Tn, such as 10:00~10:10), count its total number, and denoted as N_{1,n}. For example, if the COD sensor generates 28 first abnormal data in the current period, then N_{1,n}=28. ② Extract the number of the first anomalies in the previous statistical period (denoted as Tn-1, such as 09:50~10:00), denoted as N_{1,n-1}. For example, if there are 12 anomalies in the previous period, then N_{1,n-1}=12. ③ Calculate the first growth rate r_1 using the following formula: When N_{1,n-1}=0, if abnormal data appears in the current period, the default first growth rate is 100% (considered as significant growth); if there are still no abnormalities, the growth rate is 0%.

[0059] ④ Apply boundary constraints to the calculated r_1: If (r_1 > 300%), then take r_1 = 300% to avoid excessive parameter adjustment due to extreme anomalies; if r_1 < -50%, then take r_1 = -50% to consider the number of anomalies as significantly reduced.

[0060] 2. Adjust the target selection (filter threshold or average window): Retrieve preset adjustment rules for the second type of sensor data from the type database: If the second type is a sensor type such as "air-gas content (SO2, VOCs)" or "air-particulate matter concentration" where "noise is mainly instantaneous spikes", the adjustment object is the filter threshold; if the second type is a sensor type such as "ventilation network-air volume" or "river ventilation volume" where "data fluctuations are gradual but require a rapid response trend", the adjustment object is the average window width.

[0061] 3. Negative correlation dynamic adjustment logic "negative correlation adjustment" definition: The higher the first growth rate r_1 (the faster the number of anomalies grows), the lower the filtering threshold (or the narrower the average window) to improve the sensitivity of anomaly identification; the lower the first growth rate (or negative growth), the higher the filtering threshold (or the wider the average window) to reduce the false alarm rate.

[0062] (1) Filter threshold adjustment (applicable to gas content and particulate matter concentration): ① Obtain the initial filtering threshold Th_0 (e.g., Th_0=0.14, mg / m³ for SO2) and adjustment coefficient k_{Th}=0.2 (fixed coefficient, which can be fine-tuned remotely from the type database) for the second type. ② Calculate the adjusted filter threshold (Th_{new}), the formula is: ; ③ Apply reasonable constraints to (Th_{new}): Lower limit constraint: Th_{new}≥μ+1σ (μ is the mean of normal data in the same period of the past 7 days, and σ is the standard deviation, to avoid a large number of false alarms due to an excessively low threshold); Upper limit constraint: Th_{new}≤μ+5σ (to avoid false alarms due to an excessively high threshold). ④ Example: The SO2 sensor has Th_0 = 0.14 mg / m³, μ = 0.08 mg / m³, σ = 0.02 mg / m³, and the current r_1 = 150%, then: Th_{new}=0.14×(1-150%×0.2)=0.14×0.7=0.098mg / m³, and μ+1σ=0.10mg / m³, therefore Th_{new}=0.10mg / m³ (triggering the lower limit constraint).

[0063] (2) Average window width adjustment (applicable to air volume and ventilation volume): ① Obtain the initial average window width W_0 (e.g., W_0 = 20 data points for airflow, corresponding to 100 seconds) and adjustment coefficient k_W = 0.3 from the type database; ② Calculate the adjusted average window width W_{new} using the following formula: ; where, round() is the rounding function (the window width must be an integer number of data points); ③ Constrain W_{new}: Lower limit constraint: W_{new}≥5 data points (to avoid excessive data fluctuations due to a narrow window); Upper limit constraint: W_{new}≤50 data points (to avoid response delays due to a wide window); ④ Example: The wind speed sensor has W_0=20 data points, and the current r_1=-30% anomaly count decreases, then: W_{new}=round(20(1-(-30%)×0.3))=round(20×1.09)=22 data points (the window is widened, reducing sensitivity).

[0064] 4. The adjusted filter threshold Th_{new} or average window W_{new} is updated in real time to the cache database of the edge computing center, replacing the original initial parameters for the second type of data processing in the next statistical period. At the same time, the adjustment process (including N_{1,n}, r_1, parameters before adjustment, parameters after adjustment, and adjustment coefficients) is recorded in the "Parameter Adjustment Log" field of the anomaly analysis data and sent to the remote backend with the data packet.

[0065] II. Dynamic Adjustment Based on the Third Abnormality of the Comprehensive Growth Rate: 1. The statistical period for the second abnormal quantity statistics and the second growth rate calculation is consistent with that of the first growth rate (default 10 minutes). The steps are as follows: ① Extract the total number of second abnormal data within the current statistical period Tn, denoted as N_{2,n}. For example, if the SO2 sensor generates 16 second abnormal data in the current period, then N_{2,n}=16; ② Extract the number of second abnormal data within the previous statistical period Tn-1, denoted as N_{2,n-1}. For example, if there were 8 in the previous period, then N_{2,n-1}=8; ③ Calculate the second growth rate r_2, the formula of which is completely consistent with the first growth rate r_1, including boundary constraints: r_2in[-50%,300%]; ④ Example: N_{2,n}=16, N_{2,n-1}=8, then r_2=(16-8) / 8×100%=100%.

[0066] 2. Weighted calculation of the overall growth rate: ① Obtain preset weights from the type database: The first growth rate weight w_1 = 0.4. The first type consists of core abnormal indicators such as pollution concentration and water volume, and the weight is relatively low to avoid excessive influence from a single indicator. The second growth rate weight w_2 = 0.6. The second type consists of diffusion indicators such as gas and air volume, which are more correlated with visual anomalies and have a higher weight. Also, w_1 + w_2 = 1. The weights can be adjusted remotely according to the monitoring scenario. ② Calculate the comprehensive growth rate (r_{total}) using the following formula: ; where max(0%,·) means that if the weighted result is negative, the overall growth rate is taken as 0% (to avoid reverse adjustment of the third abnormal basis when the number of abnormalities decreases); ③ Example: r_1=150%, r_2=100%, then r_{total}=0.4×150%+0.6×100%=60%+60%=120%.

[0067] 3. Positively correlated dynamic regulation of the third abnormal basis: The definition of "positive correlation adjustment" is as follows: the higher the overall growth rate r_{total}, the faster the combined growth of the number of first and second type anomalies, the lower the base of the third type of anomaly, and the lower the pixel difference threshold, so as to improve the sensitivity of visual anomaly recognition and make it easier to capture subtle terrain changes and small solid object piles; the lower the overall growth rate, the higher the base of the third type of anomaly, so as to reduce false alarms caused by image noise.

[0068] Specific adjustment steps: ① Obtain the initial value Base_0 (e.g., the grayscale difference threshold for monitoring large solid objects (Base_0=30)) and adjustment coefficient k_{Base}=0.25 from the type database; ② Calculate the adjusted third anomaly base Base_{new} using the following formula: ; where, round() is the rounding function (the pixel difference threshold must be an integer); ③ Apply reasonable constraints to Base_{new}: Lower limit constraint: Base_{new}≥10 (to avoid the image noise being misjudged as abnormal due to the threshold being too low); Upper limit constraint: Base_{new}≤50 (to avoid the obvious visual anomalies being missed due to the threshold being too high); ④ Example: For monitoring large solid objects (Base_0=30), currently (r_{total}=120%), then: Base_{new}=round(30×(1-120%×0.25))=round(30×0.7)=21, the threshold is reduced and the sensitivity is improved.

[0069] 4. Parameter activation and recording: The adjusted Base_{new} is updated in real time to the type database cache of the edge computing center for the processing of the third type of image data in the next statistical period. At the same time, the comprehensive growth rate calculation process (r_1, r_2, weight, r_{total}), the values ​​before and after the adjustment of the third anomaly base, the adjustment coefficient, and other information are included in the "parameter adjustment log" and sent to the remote backend synchronously with the anomaly analysis data, so as to facilitate the tracing of the impact of parameter adjustment on anomaly identification when tracing the source of anomalies.

[0070] In this embodiment, the detailed steps for dynamically adjusting the first outlier and matching the multi-dimensional outlier path correlation are as follows: I. Dynamic adjustment based on the negative correlation of the first outlier in the third outlier data: This step involves mining the pixel and frequency domain features of the third type of anomalous data (visual anomalies) to dynamically optimize the first type of anomaly (adapting to the first type of sensor data, such as pollution concentration and water volume), thereby achieving the linkage and adaptation between visual anomalies and numerical anomalies. The specific steps are as follows: 1. Abnormal Pixel Data Statistics: The edge computing center extracts all abnormal image frames from the third-level abnormal data, i.e., image frames marked as "third-level candidate abnormal data". For each abnormal image frame, abnormal pixel count is performed: The statistical object is defined as the total number of pixels identified as "abnormal pixels" (i.e., pixels with a pixel difference greater than the third-level abnormal baseline) in each frame, denoted as the single-frame abnormal pixel count P_f, where f is the frame number of the abnormal frame, such as F1, F2, ..., Fm, and m is the total number of abnormal frames in the current statistical period. Abnormal pixel data construction: The single-frame abnormal pixel counts of all abnormal frames are arranged chronologically by frame number to generate a time-series abnormal pixel data sequence P=[P_{F1},P_{F2},...,P_{Fm}], and the cumulative total abnormal pixels of this sequence are calculated simultaneously. This is used for the rationality verification of subsequent intensity value calculations; Data preprocessing: If the number of abnormal pixels P_f in a certain frame is less than 0.01% of the total number of pixels, the total number of pixels = resolution width × height, such as 1920 × 1080 = 2073600 pixels, 0.01% is 207 pixels, then it is determined to be a noise frame and removed from sequence P to avoid interfering with frequency domain analysis.

[0071] 2. Frequency domain transformation and anomalous spectrum data generation: The edge computing center calls the frequency domain transformation module and uses Fast Fourier Transform (FFT) to perform frequency domain transformation on the temporally anomalous pixel data sequence P. The specific process is as follows: Sequence preprocessing: Zero-padding is performed on the abnormal pixel data sequence P after removing noise frames (padding to 2^k data points, where k is the smallest integer and satisfies 2^k ≥ sequence length; for example, if the sequence length is 80, padding is performed to 128 data points) to eliminate spectral leakage. At the same time, detrending processing is performed using a linear detrending algorithm to subtract the linear trend component of the sequence, with the formula P'_i = P_i - (a × i + b), where a is the trend slope and b is the intercept, which is obtained by least squares fitting. FFT Transform Execution: The standard FFT algorithm (such as the Cooley-Tukey algorithm) is called to perform a transformation on the preprocessed sequence P'=[P'_1,P'_2,...,P'_{2^k}], resulting in the complex form of frequency domain data F=[F_1,F_2,...,F_{2^k}], where each complex number F_i=A_i+jB_i, A_i is the real part, and B_i is the imaginary part; Abnormal spectrum data extraction: Calculate the spectral intensity (amplitude) of each frequency point I_i=sqrt{A_i^2+B_i^2}, forming abnormal spectrum data I=[I_1,I_2,...,I_{2^k}], and simultaneously record the corresponding frequency sequence f=[f_1,f_2,...,f_{2^k}], with a frequency resolution Δf=1 / T, where T is the duration of the statistical period, such as 10 minutes = 600 seconds, Δf=1 / 600≈0.0017Hz.

[0072] 3. Target spectral range and intensity value extraction: Preset spectrum range definition: Obtain a preset spectrum range matching the third type of sensor data from the type database. This range is set based on the physical change characteristics of visual anomalies (e.g., the visual change frequency of large solid object stacks and terrain collapses is concentrated in 0.01~0.5Hz, excluding high-frequency noise (>0.5Hz) and low-frequency baseline drift (<0.01Hz)); for example, the preset spectrum range for the "large solid object monitoring" scenario is [0.02Hz, 0.3Hz], and for the "terrain collapse monitoring" scenario it is [0.01Hz, 0.1Hz]; Frequency point filtering: Traverse the frequency sequence f of the abnormal spectrum data and filter out all frequency points f_j that fall within the preset spectrum range, where j is the index of the filtered frequency point and its corresponding intensity value I_j; Intensity value calculation: Calculate the weighted average intensity of the filtered intensity values ​​as the final intensity value I_{avg}, where the weight is the period length corresponding to the frequency point (the longer the period, the higher the weight, to avoid the influence of short-period noise), the formula is: Where n is the number of frequency points within the preset spectrum range; if n=0 (no effective frequency points), then I_{avg}=0 (considered as no effective visual abnormality features).

[0073] 4. Dynamic Adjustment of First Outlier Negative Correlation: Definition of "Negative Correlation Adjustment": The higher the intensity value I_{avg} (the more drastic and wider the range of visual anomalies), the lower the first outlier (improving the sensitivity of anomaly identification for the first type of data and timely capturing associated numerical anomalies); the lower the intensity value, the higher the first outlier (reducing the false alarm rate). Specific steps: Basic parameter acquisition: Obtain the initial value V_{1,0} of the first outlier from the type database (e.g., (V_{1,0}=50mg / L for COD)), adjustment coefficient k_V=0.3 (can be adjusted remotely from the backend, range 0.1~0.5), lower threshold V_{1,min} of the first outlier (set according to national / industry standards, e.g., V_{1,min}=30mg / L for COD), and upper threshold V_{1,max} (e.g., (V_{1,max}=80mg / L for COD)). The first outlier after adjustment is calculated using the formula: V_{1,new}=V_{1,0}×(1-I_{avg} / I_{max}×k_V); where I_{max} is the maximum reference intensity value preset in the type database, set based on the maximum intensity value of historical visual anomaly data, such as I_{max}=200; if I_{avg}=0, then V_{1,new}=V_{1,0} is not adjusted. Reasonableness constraint: If V_{1,new} < V_{1,min}, then V_{1,new} = V_{1,min}; if V_{1,new} > V_{1,max}, then V_{1,new} = V_{1,max}. Example: For a COD sensor, V_{1,0} = 50 mg / L, I_{avg} = 120, I_{max} = 200, k_V = 0.3, and V_{1,min} = 30 mg / L. Then: V_{1,new} = 50 × (1 - 120 / 200 × 0.3) = 41 mg / L, which meets the constraint and is effective. Parameter activation and recording: The adjusted V_{1,new} is updated in real time to the type database cache of the edge computing center, replacing the original initial value for the first type of data processing in the next statistical cycle; at the same time, the abnormal pixel statistics results, frequency domain transformation parameters, intensity value I_{avg}, the first abnormal value before and after adjustment, and other information are recorded in the "parameter adjustment log" and sent to the remote backend along with the abnormal analysis data.

[0074] II. Multi-dimensional anomaly path association matching: This step targets the three types of abnormal path data: the first type (numerical), the second type (fluctuating numerical), and the third type (visual). It calculates correlation values ​​based on quantified similarity and spatial overlap to achieve accurate matching of associations. The specific steps are as follows: 1. Preprocessing of abnormal path data: The remote backend first performs unified preprocessing on the three types of abnormal path data to ensure consistent calculation benchmarks: Data structure standardization: The unified format for all types of abnormal path data is defined as "path node set + timestamp range," where path nodes are in Cartesian coordinates (the original latitude and longitude coordinates are converted to x / y coordinates via UTM projection, unit: meters, accuracy ±0.1 meters). Each node contains (x, y, t) (x / y are the Cartesian coordinates, and t is the time of the anomaly corresponding to the node); for example, the first type of abnormal path S_1=[(x_{11},y_{11},t_{11}),(x_{12},y_{12},t_{12}),...,(x_{1n},y_{1n},t_{1n})] The second type of path S_2 and the third type of path S_3 have the same format; Time window alignment: Extract the intersection range of the timestamps of the three types of paths [t_{start},t_{end}]. For example, the time range of S_1 is 10:00~10:30, S_2 is 10:05~10:35, and S_3 is 10:02~10:28. Then the time window after alignment is 10:05~10:28. Path nodes that exceed this window are pruned, and the time-aligned paths S_1', S_2', and S_3' are retained; Path smoothing: The B-spline interpolation algorithm is used to smooth the pruned paths to eliminate the calculation error caused by node discreteness. After interpolation, the density of path nodes is uniformly 1 / 5 meters, that is, the distance between adjacent nodes is ≤5 meters.

[0075] 2. Cross-type anomaly path similarity calculation (first / second / third similarity value): To quantify the "spatial morphological similarity" between two types of paths, an improved Hausdorff distance calculation is used, with the following steps: Core formula definition: For two types of paths S_a and S_b, the Hausdorff distance H(S_a,S_b)=max(h(S_a,S_b),h(S_b,S_a)), where: h(S_a,S_b) represents the maximum value of the minimum distance from each node in S_a to S_b; similarity value conversion: normalize the Hausdorff distance to a similarity value Sim in the interval [0,1], the formula is: Sim=1-H(S_a,S_b) / H_{max}, where H_{max} is the maximum reference distance preset by the type database (based on the scale of the monitoring area, such as H_{max}=1000 meters for industrial park monitoring, H_{max}=5000 meters for watershed monitoring), the closer Sim is to 1, the more similar the path morphology; three types of similarity value calculation: ① First similarity value Sim_1: based on S_1' (first type) The first similarity value is calculated based on S_1' and S_2' (second type of path). For example, if H(S_1',S_2') = 350 meters and H_{max} = 1000 meters, then Sim_1 = 1 - 350 / 1000 = 0.65. The second similarity value Sim_2 is calculated based on S_1' and S_3' (third type of path). For example, if H(S_1',S_3') = 420 meters, then Sim_2 = 1 - 420 / 1000 = 0.58. The third similarity value Sim_3 is calculated based on S_2' and S_3'. For example, if H(S_2',S_3') = 280 meters, then Sim_3 = 1 - 280 / 1000 = 0.72.

[0076] 3. Calculation of cross-type anomaly path overlap (first / second / third overlap): Overlap ratio is used to quantify the degree of spatial overlap between two types of paths. It is calculated using the intersection-union ratio (IoU), and the steps are as follows: Path buffer generation: For each preprocessed path S_a', generate a spatial buffer polygon Polygon_a; using the path centerline as a reference, expand the preset buffer radius R to both sides (based on the sensor monitoring range, such as R=50 meters for air sensors and R=100 meters for water quality sensors) to form a closed polygon; Overlap formula: the overlap between the two types of paths Where Area(·) represents the polygon area (unit: square meters), and Overlapin[0,1], the closer to 1, the more spatial overlap there is; three types of overlap calculation: ① First overlap degree Overlap_1: IoU between Polygon_1 (S_1' buffer) and Polygon_2 (S_2' buffer), example: Area(Polygon_1∩Polygon_2)=150000㎡, Area(Polygon_1∪Polygon_2)=400000㎡, then Overlap_1=150000 / 400000=0.375; ② Second overlap degree Overlap_2: IoU between Polygon_1 and Polygon_3 (S_1' buffer) The IoU of the _3' buffer is as follows: Example: Area(Polygon_1capPolygon_3)=120000㎡, Area(Polygon_1cupPolygon_3)=380000㎡, then Overlap_2=120000 / 380000≈0.316; ③ Third overlap degree Overlap_3: IoU of Polygon_2 and Polygon_3, Example: Area(Polygon_2∩Polygon_3)=220000㎡, Area(Polygon_2∪Polygon_3)=350000㎡, then Overlap_3=220000 / 350000≈0.629.

[0077] 4. Weighted calculation of correlation values ​​(first / second / third correlation values): The correlation value is a weighted fusion of similarity value and overlap degree, which comprehensively reflects the association strength between the two types of paths. The steps are as follows: Weight Acquisition: Retrieve preset weights w_{Sim}=0.5 (similarity weight) and w_{Overlap}=0.5 (overlap weight) from the type database. The sum of the weights is 1. The weights can be adjusted remotely in the background according to the monitoring scenario. For example, when focusing on spatial morphology, w_{Sim}=0.6, and when focusing on regional coverage, w_{Overlap}=0.6. Correlation formula: Corr = w_{Sim} × Sim + w_{Overlap} × Overlap; The three correlation values ​​are calculated as follows: ① First correlation value Corr_1 = 0.5 × Sim_1 + 0.5 × Overlap_1 = 0.5 × 0.65 + 0.5 × 0.375 = 0.5125; ② Second correlation value Corr_2 = 0.5 × Sim_2 + 0.5 × Overlap_2 = 0.5 × 0.58 + 0.5 × 0.316 = 0.448; ③ Third correlation value Corr_3 = 0.5 × Sim_3 + 0.5 × Overlap_3 = 0.5 × 0.72 + 0.5 × 0.629 = 0.6745.

[0078] 5. Relationship matching based on relational databases: The remote backend pre-sets a related database, which uses "first related value interval + second related value interval + third related value interval" as an index to map the corresponding relationships (including relationship strength and relationship type description). Specific steps are as follows: Example of a relational database structure (taking an industrial park monitoring scenario as an example):

[0079] Matching logic: The calculated Corr_1=0.5125, Corr_2=0.448, and Corr_3=0.6745 are compared with the intervals in the database one by one. It is found that they all fall within the range of R2. Therefore, the association relationship R2 (medium association) is matched. Matching result recording: The matched associations (including association strength and description), three types of correlation values, interval thresholds and other information are included in the abnormal path association analysis report and stored in association with the abnormal path data and abnormal analysis data, providing a basis for subsequent source tracing and abnormal path calculation.

[0080] This step further refines the judgment logic and matching value calculation for "the association relationship conforms to the preset association strategy". The specific steps are as follows: I. Weighted calculation of path type-related values ​​(weights are strongly correlated with type): The path type correlation value is a core indicator that integrates the correlation values ​​of three types of cross-dimensional abnormal paths (Corr_1, Corr_2, and Corr_3). Its weighted average is based on the monitoring priority of the sensor type, data reliability, and pre-defined anomaly correlations to ensure that the contribution of different types of abnormal paths matches the actual monitoring scenario. 1. Acquisition and Definition of Type Weights: The remote backend extracts type-related weights from the type database. These weights are strongly correlated with the core attributes of the sensor type, and the specific rules for setting them are as follows: Weighting criteria: The first type (numerical core indicators such as pollution concentration and water volume) directly reflects the severity of environmental anomalies and has the highest reliability, with a weight of w_1=0.4; the second type (easily fluctuating indicators such as gas content and air volume) is the key carrier of anomaly propagation, with a weight of w_2=0.3; the third type (visual indicators) provides intuitive evidence of anomalies, with a weight of w_3=0.3. Weight constraints: The sum of the three types of weights satisfies w_1+w_2+w_3=1, and supports dynamic adjustment by the remote backend according to the monitoring scenario. For example, in watershed water quality monitoring, the weight of the first type can be adjusted to w_1=0.5, and the weight of the third type can be adjusted to w_3=0.2; in industrial park monitoring, the weight of the third type can be increased to w_3=0.4. Weight storage: The type database stores data using "monitoring scenario - weight combination" as the index. For example, "industrial park monitoring" corresponds to the weight combination (0.3, 0.3, 0.4), and "watershed water quality monitoring" corresponds to (0.5, 0.3, 0.2).

[0081] 2. Calculation logic for path type-related values: The path type correlation value (R_{type}) is a weighted sum of the three correlation values, as shown in the following formula: R_{type}=w_1×Corr_1+w_2×Corr_2+w_3×Corr_3; where: Corr_1, Corr_2, and Corr_3 are the first, second, and third correlation values ​​calculated above (within the range [0,1]); the value range of R_{type} is [0,1]. If the calculation result exceeds this range (such as due to weight adjustment or abnormal correlation value), it will be automatically truncated to the [0,1] interval, that is, 0 is taken when R_{type}<0, and 1 is taken when R_{type}>1.

[0082] 3. Calculation Example (based on the relevant values ​​from the previous text): Given the previously calculated values: Corr_1=0.5125, Corr_2=0.448, Corr_3=0.6745, and using the default weight combination (w_1=0.4, w_2=0.3, w_3=0.3), then: R_{type}=0.4×0.5125+0.3×0.448+0.3×0.6745=0.54175. The calculated result R_{type}=0.54175 is in the interval [0,1], so no truncation is needed.

[0083] II. Definition and Matching Decision of Association Strategy: The preset association strategy uses "path type related range" as the core judgment criterion. This range corresponds one-to-one with the association relationships in the association database, clearly defining the R_{type} threshold range corresponding to different association strengths. The specific steps are as follows: 1. The preset path type related range is stored in the association database of the remote backend. Each association relationship (such as R1~R4 above) is bound to a unique path type related range [R_{min}, R_{max}]. The range boundary is marked based on historical association analysis data to ensure coverage of the actual R_{type} distribution range of the corresponding association strength. An example is shown below:

[0084] Range description: The interval is left-closed and right-open (except for the maximum value of R1, 1.0, which is a closed interval) to avoid repeated matching of boundary values; R_{min} is the minimum R_{type} threshold for the corresponding association strength, and R_{max} is the maximum threshold.

[0085] 2. Matching and judgment logic for association relationships and association strategies: The judgment steps are as follows: ① The remote backend obtains the current matching relationship (such as the matched R2 in the previous example) and extracts its corresponding path type related range [R_{min}, R_{max}=[0.4,0.7); ② Compare the relationship between the path type related value R_{type} and the range: if R_{min}≤R_{type}<R_{max} or R_{type}≤R_{max}, for the end point of the closed interval, it is determined that "the relationship conforms to the preset association strategy"; if (R_{type}) exceeds the range, it is determined that "it does not conform" and the subsequent matching value calculation is not performed; ③ Example verification: The previously calculated R_{type}=0.54175 falls within the range [0.4,0.7) of R2, so it is determined that it conforms to the association strategy and enters the matching value calculation step.

[0086] III. Calculation of Match Values ​​(Based on Position Percentage): The matching value is an indicator that quantifies the degree to which the association relationship conforms to the association strategy. It is obtained by calculating the relative position of R_{type} within the path type relevance range. The closer the position is to the maximum value of the range, the higher the matching value. The specific steps are as follows: 1. Core parameter definition: Difference ΔR: The difference between the path type-related value and the maximum value of the range, and the formula is ΔR=R_{max}-R_{type}; Range span S_R: The numerical span of the path type-related range, the formula is S_R=R_{max}-R_{min}; Position percentage: The relative position of R_{type} within [R_{min}, R_{max}], with a value range of [0%, 100%]. When R_{type} = R_{min}, the position percentage is 0% (minimum), and when R_{type} = R_{max}, it is 100% (maximum).

[0087] 2. Formula for calculating the matching value: The calculation logic for the match value (Match_val) (in percentage form, rounded to two decimal places) is as follows: Formula derivation explanation: The position of R_{type} within the range is quantified by the proportion of the difference between R_{type} and R_{min} to the range span. The higher the proportion, the closer R_{type} is to the upper limit of the association strength, and the higher the association consistency.

[0088] 3. Calculation Example (Based on the data above) Given: R_{type}=0.54175, R2 range [R_{min}=0.4, R_{max}=0.7], then: ① Calculate the range span: S_R=0.7-0.4=0.3; ② Calculate the numerator difference: R_{type}-R_{min}=0.54175-0.4=0.14175; ③ Calculate the matching value: Match_val=(0.14175 / 0.3)×100%≈47.25%.

[0089] 4. Boundary Case Handling: If R_{type} = R_{min}: Match_val = 0.00% (meets the minimum standard of the association strategy); If R_{type} = R_{max}, such as R_{type} = 1.0 in R1: Match_val = 100.00%, meets the highest standard of the association strategy; If R_{type} < R_{min} or R_{type} > R_{max} due to data anomalies: it has been determined as "not conforming to the association strategy" in the previous matching judgment, and no matching value is calculated.

[0090] 5. Validation and recording of matching values: The calculated matching value Match_val will be directly used for the "anomaly calculation parameter adjustment" mentioned above. For example, when the matching value is ≥80%, the anomaly identification range will be expanded. At the same time, information such as the path type related value R_{type}, the path type related range [R_{min}, R_{max}], the difference ΔR, the range span S_R, the matching value Match_val, and the weight combination will be recorded in the "association strategy matching log" and stored in association with the anomaly analysis data and the association relationship results. This will facilitate the tracing of the matching logic and parameter adjustment basis during subsequent anomaly tracing.

[0091] This step, based on the correlation relationships (first / second / third correlation values) of the first, second, and third types of abnormal path data, extracts the close / similar parts of cross-type paths, filters core road segments, and fuses and removes duplicates to fit and generate an accurate source-tracing abnormal path that reflects the source of the anomaly. This ensures that the path is highly adapted to the multi-dimensional anomaly propagation characteristics. The specific steps are as follows: I. Quantitative calculation of the close / similar parts of multiple types of anomaly paths: Define "proximate or similar parts" as follows: These refer to continuous road segments (each segment consisting of at least three consecutive path nodes to avoid interference from isolated nodes) in both types of abnormal paths, where the spatial distance is ≤ a preset proximity threshold and the morphological similarity is ≥ a preset similarity threshold. The calculation is based on the preprocessed standardized path data (UTM projection plane coordinates, time alignment, and smoothed paths S_1', S_2', S_3'). The specific process is as follows: 1. Preset Threshold Acquisition: Core thresholds for path proximity / similarity determination are extracted from the type database. These thresholds are strongly correlated with the monitoring scenario: Spatial proximity threshold D_{th}: The maximum allowable distance between two types of path nodes, such as D_{th}=50 meters for industrial park monitoring and D_{th}=200 meters for watershed monitoring; Morphological similarity threshold Sim_{th}: The minimum similarity of road segment morphology, calculated based on the improved Hausdorff distance, Sim_{th}=0.6, meaning a similarity ≥60% is considered morphologically similar; Road segment length threshold L_{th}: The minimum length of an effective road segment, such as L_{th}=100 meters, to avoid short-distance segments affecting fitting accuracy.

[0092] 2. Extraction process for similar / closer parts of cross-type paths: Taking "first type and second type paths ((S_1') and (S_2'))" as an example, the process is the same for other cross-type combinations (S_2' and S_3', S_1' and S_3'): ① Segment division: Divide S_1' and S_2' into several continuous segments according to the node density (1 node / 5 meters). Each segment contains (k) nodes (k≥3). Let the set of segments of S_1' be L_1={L_{11},L_{12},...,L_ {1p}}, the set of road segments of S_2' is L_2={L_{21},L_{22},...,L_{2q}} (p, q are the number of road segments); ② Spatial proximity determination: for each pair of road segments (L_{1i},L_{2j}), i∈[1,p], j∈[1,q], calculate the center distance of the road segments D_{center}=sqrt{(x_{1i}-x_{2j})^2+(y_{1i}-y_{2j})^2}); ((x_{1i},y_{1...})^2 Let (i) be the center coordinates of (L_{1i}, and (x_{2j}, y_{2j}) be the center coordinates of L_{2j}). If D_{center} ≤ D_{th}, then they are considered spatially close. ③ Morphological similarity calculation: For spatially close road segment pairs, the morphological similarity is calculated using the road segment-level improved Hausdorff distance. Sim_{seg} = 1 - frac{H(L_{1i}, L_{2j})}{H_{max}}), (H(L_{1i}, L_{2j})}{H_{max}}). {2j}) is the Hausdorff distance between road segments, H_{max}=2×D_{th}. If Sim_{seg}≥Sim_{th}, then it is determined to be morphologically similar; ④ Valid road segment screening: retain road segment pairs that simultaneously satisfy "spatial proximity" and "morphological similarity", and calculate the actual length of the road segment (distance between the first and last nodes of the road segment). If the length ≥ (L_{th}), then the road segment pair is "close / similar part" and is merged into a continuous road segment (take the average coordinates of the two path nodes to form a merged road segment).

[0093] 3. Calculation results for the three types of cross-type close / similar parts: Type 1 and Type 2: Extract 1 to n fused road segments, combine them into first intermediate path data M_1={M_{11},M_{12},...,M_{1n}}, each M_{1k} contains fused node coordinates, road segment length, and morphological similarity Sim_{seg}; Type 2 and Type 3: Extract 1 to m fused road segments, combine them into second intermediate path data M_2={M_{21},M_{22},...,M_{2m}}; Type 1 and Type 3: Extract 1 to l fused road segments, combine them into third intermediate path data M_3={M_{31},M_{32},...,M_{3l}}.

[0094] II. Core Road Segment Filtering (Road Segment Matching Based on Relevance Values): The core logic for core road segment selection is as follows: the higher the correlation value, the greater the weight of the road segment in the intermediate path data, and complete road segments are selected first; the lower the correlation value, the more core nodes in the road segment that are most closely related to the anomaly are selected, ensuring that the road segment selection is accurately matched with the correlation strength.

[0095] 1. First segment of path data (based on the first relevance value Corr_1): ① Weight allocation: For each segment M_{1k} in the first intermediate path data M_1, assign a weight W_{1k} = Corr_1 × Sim_{seg,k}, where Sim_{seg,k} is the morphological similarity of the segment, amplifying the priority of highly similar segments; ② Segment sorting: Sort the segments in descending order of weight W_{1k}, selecting the top 1-2 segments with the highest weight (to avoid path redundancy due to too many segments); ③ Segment selection rules: If Corr_1 ≥ 0.6 (high relevance): Select the complete segment as a candidate segment, retaining all nodes; if 0.3 ≤ C orr_1 < 0.6 (moderate correlation): Select the core part of the road segment (the middle 60% of the road segment length, i.e., remove the first and last 20% of the edge nodes); If Corr_1 < 0.3 (low correlation): Only retain the 2~3 core nodes with the highest outliers in the road segment (core nodes are defined as the top 3 nodes in the road segment with the highest outlier amplitude); ④ Example: Corr_1 = 0.5125 (moderate correlation) in the previous text, the road segment M_{11} with the highest weight in M_1 has a length of 300 meters, and the middle 60% (180 meters) of nodes are extracted to form the first segment of path data P_1, which contains 36 consecutive nodes (180 meters ÷ 5 meters / node = 36).

[0096] 2. Second segment of path data (based on the second correlation value Corr_2): ① Weight allocation: For the road segment M_{2k} in the second intermediate path data M_2, the weight W_{2k} = Corr_2 × Sim_{seg,k}; ② "Closest part" determination: Select the road segment with the largest weight W_{2k} (i.e., the road segment with the strongest "spatial + morphological" correlation between the second and third types of paths) as the core candidate; ③ Road segment truncation rule: The interval determination is consistent with the first segment of path data (based on Corr_2); ④ Example: Corr_2 = 0.448 (medium correlation) mentioned above. The road segment M_{21} with the highest weight in M_2 has a length of 250 meters. Truncate 60% (150 meters) of the middle nodes to form the second segment of path data P_2, which contains 30 consecutive nodes.

[0097] 3. Third segment path data (based on the third correlation value Corr_3): ① Weight allocation: For the road segment M_{3k} in the third intermediate path data M_3, the weight W_{3k} = Corr_3 × Sim_{seg,k}; ② "Closest part" determination: Select the road segment with the largest weight W_{3k}; ③ Road segment extraction rules: consistent with the previous text; ④ Example: In the previous text, Corr_3 = 0.6745 (high correlation), the road segment M_{31} with the highest weight in M_3 has a length of 400 meters. Select the complete road segment to form the third segment path data P_3, which contains 80 consecutive nodes.

[0098] III. Fitting the Source Tracing Path of Anomalies (Connection + Deduplication + Smoothing) The core objective of fitting is to connect the three path data segments in the "reverse time sequence" of anomaly propagation (from the end of the anomaly to the source), remove spatial overlap, and form a continuous, non-redundant source tracing path. The specific steps are as follows: 1. The path nodes are arranged in reverse order of time to trace the source path. The path needs to reflect the reverse trajectory of the anomaly "from spread to the source". Therefore, the nodes of the three path data are first rearranged in reverse order by timestamp: the anomaly occurrence time (t) of each node in each path data is extracted; the nodes are sorted in descending order of (t) (latest anomaly → earliest anomaly) to obtain the reversed path P_1', P_2', P_3'.

[0099] 2. Identification and Removal of Spatial Overlap: ① Overlap Determination: For any two path segments (e.g., P_1' and P_2'), generate spatial buffers for each segment (buffer radius = R in the path overlap calculation above, e.g., 50 meters); ② Overlap Area Calculation: If the intersection area of ​​the buffers of two path segments is ≥ 30% of the buffer area of ​​a single path segment, it is determined as an "overlapping part"; ③ Deduplication Rules: Retain the "path segments with higher correlation values" in the overlapping parts. For example, if P_1' corresponds to Corr_1 = 0.5125 and P_2' corresponds to Corr_2 = 0.448, then retain the overlapping nodes of P_1' and remove the overlapping nodes of P_2'; if the correlation values ​​are close (difference ≤ 0.1), then retain the path segments with higher node density. ④ Example: P_1' and P_3' have a 100-meter overlapping section (the area of ​​the buffer zone intersection accounts for 45% of the area of ​​the overlapping part of P_1'). Since Corr_3=0.6745>Corr_1=0.5125, the overlapping nodes of P_3' are retained and the overlapping nodes of P_1' are removed.

[0100] 3. Path Connection and Smoothing Fit: ① Connection Logic: Connect the three deduplicated path nodes in the order of "earliest anomaly occurrence node → latest occurrence node" to form the initial tracing path P_{trace,init}; ensure that the distance between adjacent nodes is ≤10 meters (if the distance is too large, use linear interpolation to supplement transition nodes); ② Smoothing Processing: Use B-spline interpolation algorithm to smooth the initial path and eliminate abrupt changes in the broken line: set the interpolation node density to 1 / 2 meters to ensure path continuity; set the spline curve order to 3 (to balance smoothness and path fit), and the interpolation formula is: ; where N_{i,k}(u) is the k-th order B-spline basis function, P_i is the deduplicated path node, u∈[0,1]; ③ Path pruning: Remove isolated branches with a length <50 meters in the smoothed path (considered as noise segments), and retain the main path; ④ Final source tracing path: Obtain a continuous, smooth, and non-redundant source tracing path P_{trace}, which includes node coordinates (x / y), corresponding timestamps, associated anomaly types (first / second / third type), and other information.

[0101] 4. Fitting result verification: If the number of nodes in the final source tracing abnormal path is less than 50 (or the total length is less than 250 meters), it is judged as "insufficient path fitting". The road segment with the second highest correlation strength is automatically added (such as selecting the road segment with the second highest weight from M_1) and refitted. If it still does not meet the requirements, it is marked as "requires manual review" and all intermediate calculation data are retained.

[0102] 5. Results Recording: Information such as the source anomaly path P_{trace}, the three core path data (P_1, P_2, P_3), the calculation results of the overlapping area, the smoothing parameters, and the fitting verification results are stored together with the anomaly analysis data and the correlation results to provide core path support for subsequent anomaly magnitude data calculation and source geographic information positioning.

[0103] This application also discloses an environmental monitoring data anomaly detection and tracing device based on edge computing, including a processor, wherein the processor executes the steps of the environmental monitoring data anomaly detection and tracing method based on edge computing as described in any of the above embodiments.

Claims

1. A method for anomaly detection and source tracing in environmental monitoring data based on edge computing, characterized in that, The detection sensors are based on a distributed setup, including multiple types of sensors. These sensors are connected to an edge computing center, which in turn is connected to a remote backend. The method includes the following steps: The edge computing center acquires environmental monitoring data generated by the detection sensors in real time, and matches the anomaly analysis algorithm and anomaly calculation parameters corresponding to the type of environmental monitoring data from a preset type database. Based on environmental monitoring data and anomaly calculation parameters, anomaly analysis algorithms are used to output anomaly analysis data. Geographic information of the detection sensors corresponding to the anomaly analysis data is extracted, and the edge computing center sends the anomaly analysis data and geographic information to the remote backend. The system remotely acquires multiple anomaly analysis data sets and extracts anomaly data from multiple sensor dimensions. Based on the anomaly data of each sensor dimension and geographic information, it calculates the anomaly path data for each dimension and calculates the correlation between the anomaly path data of multiple dimensions. If the correlation matches the preset correlation strategy, it calculates the matching value between the correlation and the correlation strategy, adjusts the anomaly calculation parameters according to the matching value to output more anomaly analysis data, and updates the anomaly path data based on the latest anomaly analysis data. The source abnormal path is calculated based on multiple abnormal path data and their correlations; Based on the abnormal path data and abnormal analysis data, the abnormal magnitude data corresponding to the path is calculated. Based on the source tracing path and abnormal magnitude data, the source tracing abnormal data is calculated. Based on the source tracing abnormal path and source tracing abnormal data, the source tracing geographical information is calculated.

2. The method for anomaly detection and tracing of environmental monitoring data based on edge computing according to claim 1, characterized in that, The step of matching the anomaly analysis algorithm and anomaly calculation parameters corresponding to the type of environmental monitoring data from a preset type database also includes the following steps: The type of environmental monitoring data is extracted based on the sensor type, and the environmental monitoring data is divided into multiple groups according to the sensor type to obtain environmental monitoring groups; Based on the sensor type, the corresponding anomaly analysis algorithm and anomaly calculation parameters are obtained from the type database. The anomaly analysis algorithm includes a first analysis algorithm, a second analysis algorithm, and a third analysis algorithm. The anomaly calculation parameters include a first anomaly value, a second anomaly range, and a third anomaly basis. The first analysis algorithm uses the first outlier to filter elements in the environmental monitoring group of the same sensor type; The second analysis algorithm filters and averages the elements in the environmental monitoring group of the same sensor type, and then extracts them through the second anomaly range. The third analysis algorithm uses the third outlier to dynamically calculate the comparison image of the elements in the environmental monitoring group of the same sensor type, and identifies the image based on the preset template comparison.

3. The method for anomaly detection and tracing of environmental monitoring data based on edge computing according to claim 2, characterized in that, The steps of using anomaly analysis algorithms to output anomaly analysis data based on environmental monitoring data and anomaly calculation parameters also include the following steps: Based on the first environmental monitoring group with the first type of sensing, the abnormal trend of the values ​​in the first environmental monitoring group is obtained, and the abnormal trend is verified to match the analysis trend of the first analysis algorithm. If the verification of the abnormal trend and the analysis trend fails, the values ​​in the first environmental monitoring group are converted into values ​​that conform to the analysis trend, and then the values ​​that are greater than or less than the first abnormal value are selected as the first abnormal data. Based on the second environmental monitoring group with the second type of sensing, the second analysis algorithm is used to perform de-straining filtering on the data in the second environmental monitoring group based on a preset filtering threshold, and to perform averaging calculation based on a preset average window, and to extract the values ​​within the second anomaly range as the second anomaly data. Based on the third environmental monitoring group with the third sensing type being image type, the difference between the pixels of the current frame and the previous frame in the third environmental monitoring group is calculated. If the difference between the pixels is greater than the third anomaly baseline, the pixels of the current frame corresponding to the difference are saved into the third anomaly data. The first, second, and third abnormal data are merged into anomaly analysis data.

4. The method for anomaly detection and tracing of environmental monitoring data based on edge computing according to claim 3, characterized in that, The method also includes the following steps: The quantity of values ​​in the first abnormal data is obtained as the first abnormal number, and the growth rate of the first abnormal number is calculated as the first growth rate. The filtering threshold or the width of the average window is adjusted according to the negative correlation of the growth rate.

5. The method for anomaly detection and tracing of environmental monitoring data based on edge computing according to claim 4, characterized in that, The method also includes the following steps: The quantity of values ​​obtained from the second abnormal data is the second abnormality quantity. The growth rate of the second abnormality quantity is the second growth rate. The comprehensive growth rate is calculated by weighting the first growth rate and the second growth rate. The third abnormality basis is adjusted according to the positive correlation of the comprehensive growth rate.

6. The method for anomaly detection and tracing of environmental monitoring data based on edge computing according to claim 3, characterized in that, The method also includes the following steps: Calculate the number of pixels in each frame of the third abnormal data, and record the number of pixels corresponding to all frames as abnormal pixel data. The abnormal image data is converted into abnormal spectrum data in the frequency domain using a frequency domain conversion algorithm; Obtain the intensity value corresponding to the frequency within the preset frequency range in the abnormal spectrum data, and adjust the first abnormal value according to the negative correlation of the intensity value.

7. The method for anomaly detection and tracing of environmental monitoring data based on edge computing according to claim 3, characterized in that, The steps for calculating the correlation between outlier path data across multiple dimensions also include the following: The similarity between the first type of abnormal path data and the second type of abnormal path data is calculated as the first similarity value, and the overlap between the first type of abnormal path data and the second type of abnormal path data is calculated as the first overlap degree; the first correlation value is calculated based on the first similarity value and the first overlap degree. The similarity between the first type of abnormal path data and the third type of abnormal path data is calculated as the second similarity value, and the overlap between the first type of abnormal path data and the third type of abnormal path data is calculated as the second overlap degree; the second correlation value is calculated based on the second similarity value and the second overlap degree. The similarity between the second type of abnormal path data and the third type of abnormal path data is calculated as the third similarity value, and the overlap between the second type of abnormal path data and the third type of abnormal path data is calculated as the third overlap degree; the third correlation value is calculated based on the third similarity value and the third overlap degree. The corresponding relationships are matched from the preset association database based on the first, second, and third relevance values.

8. The method for anomaly detection and tracing of environmental monitoring data based on edge computing according to claim 7, characterized in that, If the association relationship conforms to the preset association strategy, the step of calculating the matching value between the association relationship and the association strategy also includes the following steps: The path type-related value is calculated based on the association relationship, and the association strategy is the path type-related range. If the path type-related values ​​are within the path type-related range, then the association relationship conforms to the association strategy. Calculate the difference between the path type-related value and the maximum value in the path type-related range, and calculate the percentage of the position of the difference within the path type-related range as the matching value; among them, the minimum value in the path type-related range has the smallest percentage of position, and the maximum value in the path type-related range has the largest percentage of position.

9. The method for anomaly detection and tracing of environmental monitoring data based on edge computing according to claim 7, characterized in that, The steps for calculating the source anomaly path based on multiple anomaly path data and their correlations also include the following steps: Calculate the close or similar portions of the abnormal path data of the first, second, and third types; The part that is close or similar between the first type and the second type is taken as the first intermediate path data, and the road segment corresponding to the first correlation value in the first intermediate path data is taken as the first path data segment. The closest part between the second type and the third type is taken as the second intermediate path data, and the road segment in the second intermediate path data that corresponds to the second correlation value is taken as the second path data. The closest part between the first type and the third type is taken as the third intermediate path data, and the road segment in the third intermediate path data that corresponds to the third relevant value is taken as the third path data. The source of the abnormal path is derived by fitting the first, second, and third path data.

10. A device for detecting and tracing anomalies in environmental monitoring data based on edge computing, characterized in that, The method includes a processor that performs the steps of the edge computing-based environmental monitoring data anomaly detection and tracing method as described in any one of claims 1-9.