Water level data anomaly detection processing method based on multi-method fusion

By combining the median filter, box plot and isolation forest model to detect water level data anomalies, the problem of insufficient adaptability caused by noise interference and data missing in water conservancy projects is solved, outlier detection in high-precision and high-risk scenarios is achieved, and data quality is improved.

CN120671047APending Publication Date: 2025-09-19UNIV OF JINAN
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510777445.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing water level data anomaly detection methods lack adaptability when faced with noise interference, data missing, etc., resulting in misjudgment of outliers and poor real-time performance, making it difficult to meet the needs of data cleaning and processing in water conservancy projects.

Method used

A multi-method fusion water level data anomaly detection and processing method is adopted, including the combination of median filtering + 3σ model, box plot model and isolation forest model. Through data collection, preprocessing, multi-method fusion anomaly detection, visualization and result evaluation, an adapted detection method is selected for outlier processing.

Benefits of technology

The quality of water level data has been improved, and it can effectively identify and process outliers, meet the detection needs of high-precision and high-risk scenarios, and provide reliable data support for water conservancy projects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671047A_ABST
    Figure CN120671047A_ABST
Patent Text Reader

Abstract

The invention provides a water level data anomaly detection processing method based on multi-method fusion. The water level data anomaly detection processing method comprises the following steps of data acquisition and environment configuration; library importing and environment setting: importing a pandas library, a numpy library and a matplotlib library for data processing and visualization, realizing an isolated forest algorithm by sklearn.ensemble. Isolation Forest, and drawing a box graph by means of seaborn; setting Chinese font display: ensuring normal Chinese display of the chart; and collecting and classifying multi-source data. According to the water level data anomaly detection processing method based on multi-method fusion provided by the invention, the provided water level data anomaly detection processing method based on the fusion of the box graph model, the isolated forest model and the median filtering + 3sigma model is utilized to detect and clean the water level data of the selected gate station in one month, and the RMSE and the MAE are calculated to obtain the abnormal water level data of the selected gate station. A detection method suitable for three conditions of high-precision requirement, high-risk scene comprehensive screening and rapid analysis is obtained; drawing a curve graph of the original data and the processed data, marking abnormal points, and visually comparing the cleaning effects of the three methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of water conservancy automation research, and in particular to a water level data anomaly detection and processing method based on multi-method fusion. Background Art

[0002] Water level data in water conservancy projects comes from a wide range of sources, usually from a variety of monitoring equipment and platforms, and sometimes requires manual entry or processing, which leads to uneven data format, accuracy and quality, and cannot meet the needs of hydraulic numerical models such as optimization and control. In addition, the huge amount of data and the non-uniform format bring huge challenges to data cleaning and processing. Traditional data processing methods are difficult to meet the cleaning and processing needs of water level data in water conservancy projects in the existing complex data environment.

[0003] In recent years, with the development of sensor technology and machine learning, data cleaning methods have been continuously optimized. For example, online filtering technology based on Kalman filtering has been widely used in real-time water level data processing, which can quickly estimate the state of the dynamic system. In addition, intelligent water level monitoring stations collect multi-dimensional hydrological data by integrating multiple sensors and use advanced communication technology to transmit data to the processing center in real time. Currently, many methods for detecting abnormal data have been explored, which can be divided into outlier detection based on statistics, distance, density, and prediction, as well as direct anomaly detection algorithms, combined model anomaly detection and constraint-based anomaly detection. With the further increase in the demand for data, data cleaning platforms based on parallel computing have gradually emerged. For example, parallel data cleaning platforms built using algorithms such as the Laida criterion, support vector regression (SVR) and long short-term memory network (LSTM) can efficiently process water quality and water level time series data. These platforms not only improve the efficiency of data cleaning, but also provide technical support for the real-time processing of large-scale data. However, current research and applications still have some shortcomings, including data quality problems caused by noise interference, data missing, etc., as well as insufficient adaptability, misjudgment of outliers, and poor real-time performance caused by traditional water level data anomaly detection methods.

[0004] Therefore, it is necessary to provide a water level data anomaly detection and processing method based on multi-method fusion to solve the above technical problems. Summary of the Invention

[0005] The present invention provides a water level data anomaly detection and processing method based on multi-method fusion, which solves some shortcomings in current research and application, including data quality problems caused by noise interference, data missing, etc., as well as insufficient adaptability, misjudgment of outliers, and poor real-time performance caused by traditional water level data anomaly detection methods.

[0006] To solve the above technical problems, the present invention provides a water level data anomaly detection and processing method based on multi-method fusion, comprising the following steps:

[0007] S1, data collection and environment configuration;

[0008] S11. Library import and environment setup: Import pandas, numpy, and matplotlib libraries for data processing and visualization, sklearn.ensemble.IsolationForest implements the isolation forest algorithm, and seaborn draws box plots;

[0009] Set Chinese font display: ensure that Chinese characters in the chart are displayed normally;

[0010] S12. Multi-source data collection and classification: Collect water level data from typical sluice or pump stations under different environments, including data from water diversion channel monitoring stations, on-site measured data, manually reported and uploaded data, and data automatically uploaded by automated equipment;

[0011] Record data source characteristics, such as monitoring site location, measurement environment conditions, equipment number, collection frequency, and other information and store them in categories;

[0012] S2. Data reading and preprocessing: data reading and structuring, data preprocessing and format unification;

[0013] S3, multi-method fusion anomaly detection: median filter + 3σ model detection, box plot model detection and isolation forest model detection;

[0014] S4, visualization and comparison of test results;

[0015] S41. Visualization function definition: plot_water_level_comparison: plots the time-water level curve and marks the outliers of the three methods with different colors; plot_boxplot_outliers: Generates a box plot to show data distribution and outlier locations, annotating Q1, Q3, and IQR. plot_combined_results: Generates a comprehensive comparison chart showing the overlap and difference areas of the detection results of the three methods; S42. Visual content presentation: Generate multi-dimensional comparison charts and anomaly detection result comparison tables, count the number and location of outliers in different data segments for each method, and intuitively display the detection consistency and differences; S5. Method evaluation and environmental adaptation analysis: performance index calculation and environmental adaptation analysis; S6, selection of adaptation method and outlier processing; S61. Method selection and application: Based on the evaluation results, select appropriate detection methods for different environments; S62. Outlier processing: According to the anomaly type and business requirements, the detected outliers are corrected, deleted, or retained, and the data is processed in batches using the adaptation method to generate the final anomaly-processed data set; S7. Detection effect verification and method optimization: practical application verification and model optimization. Preferably, the data reading and structuring in S2: defining the read_excel_data function, reading the Excel file to specify the worksheet, parsing it into the DataFrame format, and extracting key fields such as date, pump station name, and water level value; Data preprocessing and format unification: Define the preprocess_data function to process the time row, identify the pump station name, extract the water level data, integrate it into a unified DataFrame and sort it by time; Remove noise and invalid values, unify the timestamp format, use linear interpolation or polynomial interpolation for missing data, and convert it into a unified structure that includes water level value, monitoring time, data source identification, and monitoring point information. Preferably, the median filter + 3σ model detection: define the detect_outliers_median_filter function, apply median filtering to the preprocessed data for denoising, calculate the mean μ and standard deviation σ, and mark outliers outside the range of [μ-3σ, μ+3σ]. Box plot model detection: define the detect_outliers_boxplot function, divide the data into segments by time window, calculate Q1, Q3, IQR, set thresholds Q1-1.5IQR and Q3+1.5IQR, and mark outliers; Isolation forest model detection: Define the detect_outliers_isolation_forest function, convert the water level data into a feature vector containing time series and water level values, train the model, and call predict() to mark outliers. Preferably, the performance indicator calculation in S5 is as follows: define the evaluate_outlier_detection_methods function, and calculate RMSE, MAE, accuracy, recall, and F1 score based on manual annotation; Environmental adaptation analysis: Analyze the detection performance of each method under different water diversion channel regions, water flow conditions, and data sources, evaluate the ability to resist misjudgment of sudden fluctuations in data during the flow regulation period, and determine the optimal detection method under different environments. The preferred practical application test in S7 is: define the main function, integrate the whole process, process one month's data of a typical sluice station / pump station, call each function to complete processing, detection, visualization and evaluation, apply the results to the water conservancy scheduling scenario, and compare the support effect on scheduling decision-making before and after processing; Model optimization: Collect application feedback, analyze method deficiencies, optimize model parameters or processes, regularly update the model to adapt to changes in the environment and data characteristics, and continuously improve detection capabilities. Preferably, the median filter + 3σ model detection in S3 includes the following steps: Median filter denoising: Apply sliding window median filter to the input water level data data, with the window size window_size set by default, calculate the median of the data in the window point by point, and replace the original value to eliminate local noise; Statistical feature calculation: Calculate the mean μ and standard deviation σ of the filtered data, and determine the abnormal threshold based on the 3σ principle: Upper bound: μ + sigma_threshold × σ; Lower bound: μ-sigma_threshold×σ; Dynamic optimization: If the data fluctuation coefficient exceeds 0.5, the sigma_threshold will be automatically adjusted to 2.5 to avoid misjudgment of normal fluctuations; Outlier Marking: Traverse the data points, mark the values ​​outside the range of [lower bound, upper bound] as outliers, and return the outlier index and the marked data set. Preferably, the box plot model detection in S3 includes the following steps: Time window division: According to the time_series time series, the data is divided into multiple windows with window_hours as the unit; For the data in each window, the lower quartile Q1, upper quartile Q3 and interquartile range IQR = Q3-Q1 are calculated. Dynamic threshold calculation: The basic thresholds are Q1-iqr_scale×IQR (lower bound) and Q3+iqr_scale×IQR (upper bound); Scenario adaptation: If the data in the window contains manually reported "traffic control" semantic tags, iqr_scale is automatically adjusted to 2.0, and the threshold is relaxed to avoid misjudgment of normal fluctuations; Outlier identification: Mark data points in each window that are below the lower bound or above the upper bound as outliers, and record the timestamp and corresponding watermark of the outlier. Preferably, the isolation forest model anomaly detection in S3 includes the following steps: Feature vector construction: Convert the water level data and time series time_series into a two-dimensional feature matrix; First column: water level value (normalized to [0,1]); Second column: time characteristics; Model training and optimization: Initialize the isolation forest model and set the parameters: n_estimators: The number of trees, the default is 100. If the amount of data is less than 1000, it will be automatically reduced to 50 to improve efficiency; Contamination: Expected anomaly ratio, the default value is 0.05, and it can be adjusted dynamically based on the historical anomaly rate. Use grid search to optimize the max_samples parameter. When the data volume is greater than 5000, set it to auto, otherwise set it to 1 / 2 of the data volume. Outlier prediction and marking: Call the predict() method, output 1 (normal) or -1 (abnormal), mark the data points predicted as -1 as abnormal, and return the anomaly score.

[0016] Compared with related technologies, the water level data anomaly detection and processing method based on multi-method fusion provided by the present invention has the following beneficial effects:

[0017] The present invention provides a water level data anomaly detection and processing method based on multi-method fusion, and utilizes the proposed water level data anomaly detection and processing method based on the fusion of box plot model, isolation forest model, median filter + 3sigama model to detect and clean the water level data of the selected sluice station for one month. By calculating the RMSE and MAE, a detection method suitable for three situations: high-precision requirements, comprehensive screening of high-risk scenarios, and rapid analysis is obtained; curve graphs of the original data and the processed data are drawn, abnormal points are marked, and the cleaning effects of the three methods are intuitively compared. The calculation results of the embodiment show that the water level data anomaly detection and processing method based on multi-method fusion proposed in the present invention can be used for water level data anomaly detection in water conservancy projects, and can effectively help engineers identify and process abnormal values ​​in the water level data monitoring process, thereby improving data quality and providing reliable data support for subsequent hydrodynamic analysis and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A schematic structural diagram of a first embodiment of a water level data anomaly detection and processing method based on multi-method fusion provided by the present invention;

[0019] Figure 2This is a comparison chart of the comprehensive effects before and after processing the water level data detected in May at the selected A sluice station using the proposed water level data anomaly detection processing method based on the box plot model, isolation forest model, median filter + 3sigama model fusion;

[0020] Figure 3 It is a structural schematic diagram of the water level data curve finally obtained after being processed by the method of the present invention. DETAILED DESCRIPTION

[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0022] First embodiment

[0023] Please refer to Figure 1 、 Figure 2 and Figure 3 ,in, Figure 1 A schematic structural diagram of a first embodiment of a water level data anomaly detection and processing method based on multi-method fusion provided by the present invention; Figure 2 This is a comparison chart of the comprehensive effects before and after processing the water level data detected in May at the selected A sluice station using the proposed water level data anomaly detection processing method based on the box plot model, isolation forest model, median filter + 3sigama model fusion; Figure 3 The diagram is a schematic diagram of the structure of the water level data curve finally obtained after processing by the method of the present invention. A method for detecting and processing water level data anomalies based on multi-method fusion includes the following steps:

[0024] S1, data collection and environment configuration;

[0025] S11. Library import and environment setup: Import pandas, numpy, and matplotlib libraries for data processing and visualization, sklearn.ensemble.IsolationForest implements the isolation forest algorithm, and seaborn draws box plots;

[0026] Set Chinese font display: ensure that Chinese characters in the chart are displayed normally;

[0027] S12. Multi-source data collection and classification: Collect water level data from typical sluice or pump stations under different environments, including data from water diversion channel monitoring stations, on-site measured data, manually reported and uploaded data, and data automatically uploaded by automated equipment;

[0028] Record data source characteristics, such as monitoring site location, measurement environment conditions, equipment number, collection frequency, and other information and store them in categories;

[0029] S2. Data reading and preprocessing: data reading and structuring, data preprocessing and format unification;

[0030] S3, multi-method fusion anomaly detection: median filter + 3σ model detection, box plot model detection and isolation forest model detection;

[0031] S4, visualization and comparison of test results;

[0032] S41. Visualization function definition:

[0033] plot_water_level_comparison: plots the time-water level curve and marks the outliers of the three methods with different colors;

[0034] plot_boxplot_outliers: Generates a box plot to show data distribution and outlier locations, annotating Q1, Q3, and IQR.

[0035] plot_combined_results: Generates a comprehensive comparison chart showing the overlap and difference areas of the detection results of the three methods;

[0036] S42. Visual content presentation: Generate multi-dimensional comparison charts and anomaly detection result comparison tables, count the number and location of outliers in different data segments for each method, and intuitively display the detection consistency and differences;

[0037] S5. Method evaluation and environmental adaptation analysis: performance index calculation and environmental adaptation analysis;

[0038] S6, selection of adaptation method and outlier processing;

[0039] S61. Method selection and application: Based on the evaluation results, select appropriate detection methods for different environments;

[0040] S62. Outlier processing: According to the anomaly type and business requirements, the detected outliers are corrected, deleted, or retained, and the data is processed in batches using the adaptation method to generate the final anomaly-processed data set;

[0041] S7. Detection effect verification and method optimization: practical application verification and model optimization.

[0042] Data reading and structuring in S2: define the read_excel_data function, read the specified worksheet of the Excel file, parse it into the DataFrame format, and extract key fields such as date, pump station name, and water level value;

[0043] Data preprocessing and format unification: Define the preprocess_data function to process the time row, identify the pump station name, extract the water level data, integrate it into a unified DataFrame and sort it by time;

[0044] Remove noise and invalid values, unify the timestamp format, use linear interpolation or polynomial interpolation for missing data, and convert it into a unified structure that includes water level value, monitoring time, data source identification, and monitoring point information.

[0045] The median filter + 3σ model detection: define the detect_outliers_median_filter function, apply median filtering to the preprocessed data for denoising, calculate the mean μ and standard deviation σ, and mark outliers outside the range of [μ-3σ, μ+3σ].

[0046] Box plot model detection: define the detect_outliers_boxplot function, divide the data into segments by time window, calculate Q1, Q3, IQR, set thresholds Q1-1.5IQR and Q3+1.5IQR, and mark outliers;

[0047] Isolation forest model detection: Define the detect_outliers_isolation_forest function, convert the water level data into a feature vector containing time series and water level values, train the model, and call predict() to mark outliers.

[0048] Performance indicator calculation in S5: define the evaluate_outlier_detection_methods function, calculate RMSE, MAE, accuracy, recall rate, and F1 score based on manual annotation;

[0049] Environmental adaptation analysis: Analyze the detection performance of each method under different water diversion channel regions, water flow conditions, and data sources, evaluate the ability to resist misjudgment of sudden fluctuations in data during the flow regulation period, and determine the optimal detection method under different environments.

[0050] The practical application test in S7: define the main function, integrate the entire process, process one month of data from a typical sluice station / pump station, call various functions to complete processing, detection, visualization and evaluation, apply the results to water conservancy scheduling scenarios, and compare the support effect on scheduling decisions before and after processing;

[0051] Model optimization: Collect application feedback, analyze method deficiencies, optimize model parameters or processes, regularly update the model to adapt to changes in the environment and data characteristics, and continuously improve detection capabilities.

[0052] like Figure 1As shown, the invention is mainly used to detect and process outliers in water level data and generate a visual comparison chart. The entire process uses three different outlier detection methods and compares and analyzes the detection effects of these three methods. The specific steps are as follows:

[0053] S1. By automatically identifying the sluice station name, water level location (before / after the station) and time row in the worksheet, it adapts to unstructured data and reduces manual intervention.

[0054] First, the collected water level monitoring data of sluice station A is preprocessed.

[0055] To ensure the logic and consistency of traffic data, the original data is converted into structured time series data.

[0056] It is necessary to clean the traffic at a spatial scale, standardize the date format in the raw data, find rows containing multiple time points, and determine the position of the time column.

[0057] Analyze the data structure in the water level monitoring worksheet, enter a valid sluice station name, extract the water level data before and after the station based on the label, and associate it with the sluice station name.

[0058] S2. Apply the median filter + 3sigama model to perform anomaly detection on the water level data monitored by Gate A.

[0059] Set the one-dimensional array X to store the data sequence to be tested and the median filter window size parameter W. Set the standard deviation multiple parameter K to calculate the outlier threshold. The default value is 3.

[0060] Apply a median filter to the data to be measured and calculate the median within the window.

[0061] Based on the original data sequence X and the median filtered data sequence X_filtered, the residual sequence R is calculated point by point:

[0062] R i =x i -x-filtered i (i=1,2,…N)

[0063] Perform statistical analysis on the residual sequence R and calculate its standard deviation σ. The formula is as follows:

[0064]

[0065] Based on the standard deviation σ of the residual sequence R, an outlier determination threshold T=3σ is set. If the absolute value of the residual |Ri|>T, the corresponding data point Xi is determined to be an outlier.

[0066] For the data point Xi that is judged as an outlier, the median filtering result X_filtered is used to replace it to generate the corrected data sequence X_corrected, where:

[0067]

[0068] Returns the processed data (outliers are replaced by median filter results), the index of the outlier, and the filtered data.

[0069] S3. Use the box plot model to monitor the outliers of the water level data monitored by Gate A.

[0070] Import the one-dimensional data sequence to be processed And the data point x_target to be tested, define the whisker coefficient k of the box plot, which is used to calculate the outlier detection range. The default value is 1.5.

[0071] Calculate the quartiles and interquartile range:

[0072] iqr=q3-q1

[0073] Perform outlier detection and create processed data. For data points x_i that are judged to be outliers, use the median of the data sequence median(X) to replace them and generate a corrected data sequence in:

[0074]

[0075] Returns the processed data (outliers are replaced by the median), the index of the outlier, the outlier value, and box plot statistics (Q1, Q2, Q3, lower limit, upper limit).

[0076] Draw a scatter plot of normal values ​​and outliers.

[0077] S4. Apply the isolation forest model to detect outliers in the water level data monitored by Gate A.

[0078] Receives input one-dimensional data array Where n ≥ 1, and the array element type is a real number. Define an outlier ratio parameter α to control the sensitivity of outlier detection. The default value is 0.05 (i.e., 5%). The value range of the outlier ratio parameter α is 0 < α ≤ 0.2. If it exceeds this range, it will automatically adjust to the default value.

[0079] Make sure the data is not empty and has sufficient length, reshape the data into a two-dimensional array (Isolation Forest requires two-dimensional input)

[0080] Create and train an isolation forest model to predict outliers (-1 for anomaly, 1 for normal)

[0081] Get the index and value of the outlier, create the processed data, replace the outlier with the median, and use the index to replace the outlier.

[0082] S5. Through visual evaluation of the three test results, analyze and select the test methods suitable for different environments and test their effectiveness.

[0083] Plot the original data and outliers, plot the processed data, set labels, adjust the layout, and save the image.

[0084] Draw a box plot of a single data set and mark the outliers. Compare the box plot distribution differences between the original data and the processed data.

[0085] Draw the original data curve, mark the outlier points, and draw the data curves after median filtering + 3sigama outlier detection, box plot outlier detection, and isolation forest outlier detection and outlier removal.

[0086] S6. Quantitatively compare the anomaly detection performance of the three methods. Define the evaluation metric calculation function. RMSE (root mean square error) is used to measure the overall deviation between the original data and the processed data. A smaller value indicates that the processed data is closer to the original normal data. MAE (mean absolute error) is used to reflect the average error of data processing and is insensitive to outliers. Construct an evaluation dictionary.

[0087]

[0088] Define the main function for fully automated data processing, perform anomaly detection and evaluation, filter data, perform data preprocessing, and parse the sluice station name, time series, and water level data in the table.

[0089] Perform outlier detection and visualization, traverse each gate station and location, call three detection methods, evaluate and save the results. Summarize the evaluation results by gate station and location and save the results.

[0090] The details are shown in the following table:

[0091]

[0092]

[0093] Table 1 shows the evaluation of 1020 water level information in front of sluice station A using the water level data anomaly detection processing method based on multi-method fusion. The number of outliers detected by the median filter + 3sigama model is 1, the number of outliers detected by the isolation forest model is 42, and the number of outliers detected by the box plot model is 20.

[0094] The Isolation Forest model has the highest sensitivity for detecting outliers, likely due to its tree structure's adaptability to complex distributions. The Median Filter + 3σ model is the most conservative, detecting only one outlier and potentially more suitable for scenarios requiring high-precision data correction.

[0095] Comparing the error indicators, the error of the results obtained using the median filter + 3σ model is significantly lower than that of other methods (both RMSE and MAE are the lowest), indicating that the corrected data is closest to the original normal data.

[0096] Figure 2 The comparison chart of the results obtained using the isolation forest model shows that the data fluctuations were significantly reduced after processing (the water level range dropped from 49.1m-49.4m to 49.12m-49.26m), but the number of outliers removed far exceeded that of other methods. The comparison chart of the box plot model shows that the original data had a wide distribution (49.1m to 49.5m), while the data range was narrowed after processing, and marginal outliers were removed. The comparison chart of the median filter + 3σ model shows that the original water level data had small fluctuations, and the data was smoothed after processing, with only a few outliers removed. The water level range is stable at 49.1m to 49.4m, and the outliers are likely minor noise.

[0097] Table 1 shows that the high outlier detection rate of the isolation forest model is accompanied by high error. This oversensitivity may lead to some normal data being misclassified as anomalies, introducing bias during replacement. The median filter + 3σ model is suitable for high-precision requirements and is recommended for scenarios requiring strict data quality assurance. The isolation forest model is suitable for comprehensive screening in high-risk scenarios, but manual review of the results is required. The box plot is suitable for rapid analysis and daily monitoring, balancing sensitivity and error.

[0098] According to Example 1, it can be seen that in the water level data anomaly detection process, the box plot model is simple and convenient, but there is a probability that it cannot detect the outliers of the data. The isolation forest model may misjudge more normal values ​​in the data as outliers. The median filter + 3sigama model is more suitable for water level data anomaly detection and processing.

[0099] Utilizing the processing method of the present invention, three methods, namely, box plot model, median filter + 3sigama model, and isolation forest model, are adopted to perform outlier detection processing on water level data, thus unifying the data format, detecting all abnormal points in the data, and processing dirty data in the original data well, thereby improving the data quality. This proves that the water level data detection and processing method based on multi-method fusion described in the present invention can effectively improve the quality of water level data monitored during the operation of sluice stations and pump stations, and provide reliable data support for subsequent hydrodynamic analysis and decision-making.

[0100] Compared with related technologies, the water level data anomaly detection and processing method based on multi-method fusion provided by the present invention has the following beneficial effects:

[0101] The present invention provides a water level data anomaly detection and processing method based on multi-method fusion, and uses the proposed water level data anomaly detection and processing method based on the fusion of box plot model, isolation forest model, median filter + 3sigama model to detect and clean the water level data of the selected gate station for one month. By calculating RMSE and MAE, a detection method suitable for three situations: high-precision requirements, comprehensive screening of high-risk scenarios, and rapid analysis is obtained; a curve graph of the original data and the processed data is drawn, anomaly points are marked, and the cleaning effects of the three methods are intuitively compared. The calculation results of the embodiment show that the water level data anomaly detection and processing method based on multi-method fusion proposed by the present invention can be used for water level data anomaly detection in water conservancy projects, and can effectively help engineers identify and process outliers in the water level data monitoring process, thereby improving data quality and providing reliable data support for subsequent hydrodynamic analysis and decision-making.

[0102] Second embodiment

[0103] Based on the multi-method fusion-based water level data anomaly detection and processing method provided in the first embodiment of this application, the second embodiment of this application proposes another multi-method fusion-based water level data anomaly detection and processing method. The second embodiment is merely a preferred embodiment of the first embodiment, and the implementation of the second embodiment will not affect the independent implementation of the first embodiment.

[0104] Specifically, the difference of the water level data anomaly detection and processing method based on multi-method fusion provided in the second embodiment of the present application is that the median filter + 3σ model detection in S3 includes the following steps:

[0105] Median filter denoising: Apply sliding window median filter to the input water level data data, with the window size window_size set by default, calculate the median of the data in the window point by point, and replace the original value to eliminate local noise;

[0106] Statistical feature calculation: Calculate the mean μ and standard deviation σ of the filtered data, and determine the abnormal threshold based on the 3σ principle:

[0107] Upper bound: μ + sigma_threshold × σ;

[0108] Lower bound: μ-sigma_threshold×σ;

[0109] Dynamic optimization: If the data fluctuation coefficient exceeds 0.5, the sigma_threshold will be automatically adjusted to 2.5 to avoid misjudgment of normal fluctuations;

[0110] Outlier Marking: Traverse the data points, mark the values ​​outside the range of [lower bound, upper bound] as outliers, and return the outlier index and the marked data set.

[0111] The box plot model detection in S3 includes the following steps:

[0112] Time window division: According to the time_series time series, the data is divided into multiple windows with window_hours as the unit;

[0113] For the data in each window, the lower quartile Q1, upper quartile Q3 and interquartile range IQR = Q3-Q1 are calculated.

[0114] Dynamic threshold calculation:

[0115] The basic thresholds are Q1-iqr_scale×IQR (lower bound) and Q3+iqr_scale×IQR (upper bound);

[0116] Scenario adaptation: If the data in the window contains manually reported "traffic control" semantic tags, iqr_scale is automatically adjusted to 2.0, and the threshold is relaxed to avoid misjudgment of normal fluctuations;

[0117] Outlier identification: Mark data points in each window that are below the lower bound or above the upper bound as outliers, and record the timestamp and corresponding watermark of the outlier.

[0118] The isolation forest model anomaly detection in S3 includes the following steps:

[0119] Feature vector construction:

[0120] Convert the water level data and time series time_series into a two-dimensional feature matrix;

[0121] First column: water level value (normalized to [0,1]);

[0122] Second column: time characteristics;

[0123] Model training and optimization:

[0124] Initialize the isolation forest model and set the parameters:

[0125] n_estimators: The number of trees, the default is 100. If the amount of data is less than 1000, it will be automatically reduced to 50 to improve efficiency;

[0126] Contamination: Expected anomaly ratio, the default value is 0.05, and it can be adjusted dynamically based on the historical anomaly rate.

[0127] Use grid search to optimize the max_samples parameter. When the data volume is greater than 5000, set it to auto, otherwise set it to 1 / 2 of the data volume.

[0128] Outlier prediction and marking: Call the predict() method, output 1 (normal) or -1 (abnormal), mark the data point predicted as -1 as abnormal, and return the abnormal score.

[0129] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A water level data anomaly detection and processing method based on multi-method fusion, characterized in that: The following steps are involved: S1, data collection and environment configuration; S11. Library import and environment setup: Import pandas, numpy, and matplotlib libraries for data processing and visualization, sklearn.ensemble.IsolationForest implements the isolation forest algorithm, and seaborn draws box plots; Set Chinese font display: ensure that Chinese characters in the chart are displayed normally; S12. Multi-source data collection and classification: Collect water level data from typical sluice or pump stations under different environments, including data from water diversion channel monitoring stations, on-site measured data, manually reported and uploaded data, and data automatically uploaded by automated equipment; Record data source characteristics, such as monitoring site location, measurement environment conditions, equipment number, collection frequency, and other information and store them in categories; S2. Data reading and preprocessing: data reading and structuring, data preprocessing and format unification; S3, multi-method fusion anomaly detection: median filter + 3σ model detection, box plot model detection and isolation forest model detection; S4, visualization and comparison of test results; S41. Visualization function definition: plot_water_level_comparison: plots the time-water level curve and marks the outliers of the three methods with different colors; plot_boxplot_outliers: Generates a box plot to show data distribution and outlier locations, annotating Q1, Q3, and IQR. plot_combined_results: Generates a comprehensive comparison chart showing the overlap and difference areas of the detection results of the three methods; S42. Visual content presentation: Generate multi-dimensional comparison charts and anomaly detection result comparison tables, count the number and location of outliers in different data segments for each method, and intuitively display the detection consistency and differences; S5. Method evaluation and environmental adaptation analysis: performance index calculation and environmental adaptation analysis; S6, selection of adaptation method and outlier processing; S61. Method selection and application: Based on the evaluation results, select appropriate detection methods for different environments; S62. Outlier processing: According to the anomaly type and business requirements, the detected outliers are corrected, deleted, or retained, and the data is processed in batches using the adaptation method to generate the final anomaly-processed data set; S7. Detection effect verification and method optimization: practical application verification and model optimization.

2. The water level data anomaly detection and processing method based on multi-method fusion according to claim 1 is characterized in that: Data reading and structuring in S2: define the read_excel_data function, read the specified worksheet of the Excel file, parse it into the DataFrame format, and extract key fields such as date, pump station name, and water level value; Data preprocessing and format unification: Define the preprocess_data function to process the time row, identify the pump station name, extract the water level data, integrate it into a unified DataFrame and sort it by time; Remove noise and invalid values, unify the timestamp format, use linear interpolation or polynomial interpolation for missing data, and convert it into a unified structure that includes water level value, monitoring time, data source identification, and monitoring point information.

3. The water level data anomaly detection and processing method based on multi-method fusion according to claim 1 is characterized in that: The median filter + 3σ model detection: define the detect_outliers_median_filter function, apply median filtering to the preprocessed data for denoising, calculate the mean μ and standard deviation σ, and mark outliers outside the range of [μ-3σ, μ+3σ]. Box plot model detection: define the detect_outliers_boxplot function, divide the data into segments by time window, calculate Q1, Q3, IQR, set thresholds Q1-1.5IQR and Q3+1.5IQR, and mark outliers; Isolation forest model detection: Define the detect_outliers_isolation_forest function, convert the water level data into a feature vector containing time series and water level values, train the model, and call predict() to mark outliers.

4. The water level data anomaly detection and processing method based on multi-method fusion according to claim 1 is characterized in that: Performance indicator calculation in S5: define the evaluate_outlier_detection_methods function, calculate RMSE, MAE, accuracy, recall rate, and F1 score based on manual annotation; Environmental adaptation analysis: Analyze the detection performance of each method under different water diversion channel regions, water flow conditions, and data sources, evaluate the ability to resist misjudgment of sudden fluctuations in data during the flow regulation period, and determine the optimal detection method under different environments.

5. The water level data anomaly detection and processing method based on multi-method fusion according to claim 1 is characterized in that: The practical application test in S7: define the main function, integrate the entire process, process one month of data from a typical sluice station / pump station, call various functions to complete processing, detection, visualization and evaluation, apply the results to water conservancy scheduling scenarios, and compare the support effect on scheduling decisions before and after processing; Model optimization: Collect application feedback, analyze method deficiencies, optimize model parameters or processes, regularly update the model to adapt to changes in the environment and data characteristics, and continuously improve detection capabilities.

6. The water level data anomaly detection and processing method based on multi-method fusion according to claim 1 is characterized in that: The median filter + 3σ model detection in S3 includes the following steps: Median filter denoising: Apply sliding window median filter to the input water level data data, with the window size window_size set by default, calculate the median of the data in the window point by point, and replace the original value to eliminate local noise; Statistical feature calculation: Calculate the mean μ and standard deviation σ of the filtered data, and determine the abnormal threshold based on the 3σ principle: Upper bound: μ + sigma_threshold × σ; Lower bound: μ-sigma_threshold×σ; Dynamic optimization: If the data fluctuation coefficient exceeds 0.5, the sigma_threshold will be automatically adjusted to 2.5 to avoid misjudgment of normal fluctuations; Outlier Marking: Traverse the data points, mark the values ​​outside the range of [lower bound, upper bound] as outliers, and return the outlier index and the marked data set.

7. The water level data anomaly detection and processing method based on multi-method fusion according to claim 1 is characterized in that: The box plot model detection in S3 includes the following steps: Time window division: According to the time_series time series, the data is divided into multiple windows with window_hours as the unit; For the data in each window, the lower quartile Q1, upper quartile Q3 and interquartile range IQR = Q3-Q1 are calculated. Dynamic threshold calculation: The basic thresholds are Q1-iqr_scale×IQR (lower bound) and Q3+iqr_scale×IQR (upper bound); Scenario Adaptation: If the data in the window contains manually reported "traffic control" semantic tags, iqr_scale is automatically adjusted to 2.0, relaxing the threshold to avoid misjudgment of normal fluctuations. Outlier identification: Mark data points in each window that are below the lower bound or above the upper bound as outliers, and record the timestamp and corresponding watermark of the outlier.

8. The water level data anomaly detection and processing method based on multi-method fusion according to claim 1 is characterized in that: The isolation forest model anomaly detection in S3 includes the following steps: Feature vector construction: Convert the water level data and time series time_series into a two-dimensional feature matrix; First column: water level value (normalized to [0,1]); Second column: time characteristics; Model training and optimization: Initialize the isolation forest model and set the parameters: n_estimators: The number of trees, the default is 100. If the amount of data is less than 1000, it will be automatically reduced to 50 to improve efficiency; Contamination: Expected anomaly ratio, the default value is 0.05, and it can be adjusted dynamically based on the historical anomaly rate. Use grid search to optimize the max_samples parameter. When the data volume is greater than 5000, set it to auto, otherwise set it to 1 / 2 of the data volume. Outlier prediction and marking: Call the predict() method, output 1 (normal) or -1 (abnormal), mark the data points predicted as -1 as abnormal, and return the anomaly score.

Citation Information

Cited By

  • Drainage pipe network data cleaning and intelligent repairing method fusing multiple models

    CN121327331A