Treatment method and system for monitoring data of water supply plant

By acquiring and labeling the time units of monitoring data, selecting suitable outlier identification methods and cleaning strategies, and using the ARIMA model and MLP model for data cleaning, solving the complexity and diversity of monitoring data cleaning in the water supply plant, achieving accurate and efficient data cleaning effects, and providing solid guarantees for water supply safety.

CN120198008APending Publication Date: 2025-06-24PIPE NETWORK MANAGEMENT BRANCH OF BEIJING WATERWORKS GRP CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510253927.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to adapt to the complexity and diversity of monitoring data in water supply plants, it is difficult to accurately clean two types of data, and it is difficult to adjust the data processing method according to different time units.

Method used

By acquiring and labeling the time units of the monitoring data, selecting appropriate outlier identification methods and cleaning strategies, and using the ARIMA model and MLP model for data cleaning, including null value filling and data repair.

Benefits of technology

It realizes accurate and efficient cleaning of complex and diverse monitoring data, different types and different time units, ensuring the safe data quality of water supply.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198008A_ABST
    Figure CN120198008A_ABST
Patent Text Reader

Abstract

The invention discloses a water supply plant monitoring data-oriented treatment method and system, and relates to the field of data processing, and the method comprises the steps: obtaining the monitoring data of each water affair monitoring parameter, and at least marking a time unit; selecting an optimal abnormal value discrimination method for one water affair monitoring parameter from a plurality of abnormal value discrimination methods according to an abnormal value discrimination method evaluation standard, and carrying out null value replacement on an abnormal value discriminated from the monitoring data of the water affair monitoring parameter based on the optimal abnormal value discrimination method; selecting a proper cleaning strategy according to whether the data is manually collected or not, and adjusting the cleaning parameters of the selected cleaning strategy according to the time unit of the monitoring data of the water affair monitoring parameters; and performing data cleaning including null value filling by using the cleaning strategy with the adjusted cleaning parameters, and labeling and storing a cleaning result. According to the method, the accuracy and high efficiency of data cleaning can be ensured, the method can adapt to monitoring data of complex, diversified, different types and different time units of water quality monitoring parameters, and a solid guarantee is provided for water supply safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a governance method and system for manual and online monitoring data of water supply plants. Background Art

[0002] The monitoring data of water supply plants has multi-domain characteristics and complementarity, specifically as follows:

[0003] 1. The monitoring data of water supply plants has multi-domain characteristics.

[0004] The monitoring of water plants involves different parameters in multiple fields, and the data is complex and diverse, including pH value, dissolved oxygen (DO), conductivity, turbidity, temperature, permanganate index, etc. These parameters are crucial for evaluating water quality and ensuring the effectiveness of the water treatment process. For example, changes in pH value may indicate the presence of acidic or alkaline pollutants; the dissolved oxygen level is an important indicator for measuring biological activities in water and is crucial for maintaining ecological balance; the permanganate index reflects the degree of organic pollution in water bodies.

[0005] 2. The monitoring data of water supply plants has complementarity.

[0006] The manual data and online data of water plants coexist, making the monitoring data have unique complementarity, forming an efficient data enhancement strategy. The accuracy and pertinence of manual data, combined with the real-time and extensive nature of online data, not only make up for the deficiencies of a single data source, but also improve the overall quality and applicability of the data, ensuring that more comprehensive and in-depth business insights can be obtained during the analysis process, and thus supporting more accurate and robust decision-making.

[0007] Data cleaning is the cornerstone of water supply safety. In existing technologies, in order to improve data quality, usually a single technical means (such as interpolation method, statistical method or machine learning model) is adopted. This method has the following deficiencies:

[0008] 1. It is difficult to adapt to the complexity and diversity of water service data;

[0009] 2. It is difficult to achieve precise cleaning for both manual and online data types and fails to effectively utilize the complementarity of manual and online data;

[0010] 3. It is difficult to adjust the data processing method according to different time units.

[0011] Therefore, there is an urgent need for a precise and efficient governance technology that can adapt to complex, diverse, different types, and different time units of monitoring data to provide a solid guarantee for water supply safety. Summary of the Invention

[0012] The present invention provides a governance method and system for monitoring data of water supply plants, aiming to solve the above problems existing in data cleaning in the water service industry.

[0013] A governance method for monitoring data of water supply plants provided by the present invention includes: obtaining monitoring data of each water service monitoring parameter, and at least labeling a time unit for the monitoring data of each water service monitoring parameter; according to the evaluation criteria of outlier discrimination methods, selecting the best outlier discrimination method applicable to the monitoring data of a water service monitoring parameter from several outlier discrimination methods, and replacing the outliers identified from the monitoring data of the water service monitoring parameter based on the best outlier discrimination method with null values; according to whether the monitoring data of the water service monitoring parameter is manually collected data, selecting a suitable cleaning strategy, and adjusting the cleaning parameters of the selected cleaning strategy according to the time unit of the monitoring data of the water service monitoring parameter; using the cleaning strategy with the adjusted cleaning parameters to perform data cleaning including null value filling on the monitoring data of the water service monitoring parameter, and after performing data cleaning on the monitoring data of the water service monitoring parameter, labeling and saving the cleaning result.

[0014] Preferably, the selecting the best outlier discrimination method applicable to the monitoring data of a water service monitoring parameter according to the evaluation criteria of outlier discrimination methods includes: arranging the monitoring data of the water service monitoring parameter in chronological order to obtain the time series of the water service monitoring parameter, and removing duplicates from the time series of the water service monitoring parameter through null value replacement; respectively using each outlier discrimination method to perform outlier discrimination on the time series of the water service monitoring parameter, and respectively performing null value replacement processing of outliers, data repair processing of null values, and prediction processing of time series data on the time series of the water service monitoring parameter according to the outliers identified by each outlier discrimination method, so as to evaluate the performance of each outlier discrimination method; selecting the outlier discrimination method with the optimal performance from all outlier discrimination methods, and determining the outlier discrimination method with the optimal performance as the best outlier discrimination method of the water service monitoring parameter.

[0015] Preferably, the several outlier discrimination methods at least include the moving average method, the STL decomposition method, the Z decomposition method, the exponential moving average method, and the isolation forest method.

[0016] Preferably, the selecting a suitable cleaning strategy according to whether the monitoring data of the water service monitoring parameter is manually collected data includes: judging whether the monitoring data of the water service monitoring parameter is manually collected data; if it is manually collected data, then selecting the first cleaning strategy applying the ARIMA model, otherwise, selecting the second cleaning strategy applying manually collected data and / or applying the ARIMA model.

[0017] Preferably, adjusting the cleaning parameters of the selected cleaning strategy according to the time unit of the monitoring data of the water service monitoring parameters includes: adjusting the autoregressive term coefficient p, the differencing order d, and the moving average term order q of the ARIMA model in the first cleaning strategy or the second cleaning strategy according to the time unit of the monitoring data of the water service monitoring parameters; wherein: when the time unit of the monitoring data of the water service monitoring parameters is "day", p = 5, d = 1, q = 5; when the time unit of the monitoring data of the water service monitoring parameters is "hour", p = 3, d = 1, q = 3; when the time unit of the monitoring data of the water service monitoring parameters is "5 minutes", p = 2, d = 1, q = 2.

[0018] Preferably, if the cleaning strategy is the first cleaning strategy, the cleaning of the monitoring data of the water service monitoring parameters including null value filling using the cleaning strategy with adjusted cleaning parameters includes: splitting the time series of the water service monitoring parameters to obtain multiple data segments, where the first and last values of each data segment are not null, and the last value is null; merging the data segments with null values to be filled with all the previous data segments with filled null values in a consecutive manner, and using the ARIMA model to predict and fill the null values in the merged data segments until all the null values in all the data segments are filled; or, splitting the time series of the water service monitoring parameters to obtain multiple data segments, where the first and last values of each data segment are not null, and the last value is null; determining whether to apply the MLP model according to the data volume of the time series of the water service monitoring parameters; if the MLP model is not applied, merging the data segments with null values to be filled with all the previous data segments with filled null values in a consecutive manner, and using the ARIMA model to predict and fill the null values in the merged data segments until all the null values in all the data segments are filled; if the MLP model is applied, merging the data segments with null values to be filled with all the previous data segments with filled null values in a consecutive manner, for the previous specified number of data segments, using the ARIMA model to predict and fill the null values in the merged data segments, and for the other data segments, using the MLP model to predict and fill the null values in the merged data segments until all the null values in the other data segments are filled.

[0019] Preferably, if the cleaning strategy is the second cleaning strategy, the data cleaning of the monitoring data of the water service monitoring parameters using the cleaning strategy with the adjusted cleaning parameters and including null value filling includes: querying the manually collected data corresponding to the online instrument data of the water service monitoring parameters; if no manually collected data corresponding to the online instrument data of the water service monitoring parameters is queried, using the first cleaning strategy with the adjusted cleaning parameters to fill the null values in the time series of the water service monitoring parameters; if manually collected data corresponding to the online instrument data of the water service monitoring parameters is queried, initially filling the null values by the complementary filling method of the manually collected data and the online instrument data, and after initially filling the null values, querying whether there are still unfilled null values. When it is queried that there are still unfilled null values, using the first cleaning strategy with the adjusted cleaning parameters to perform secondary null value filling on the time series of the water service monitoring parameters.

[0020] Preferably, the prediction and filling of the null values in the merged data segments using the MLP model includes: extracting the time features of each null value in the merged data segments, where the time features include year, month, date, day of the week, day of the year, quarter, number of weeks calculated using the ISO calendar system, and number of seconds since the Unix epoch; forming a time feature matrix for each null value based on the time features of each null value; and respectively inputting the time feature matrix of each null value into the trained MLP model to use the trained MLP model for null value prediction and filling the null values with the predicted values.

[0021] Preferably, the annotation and storage of the cleaning results include: annotating the cleaning results with at least the name of the best outlier discrimination method, the cleaning strategy adopted, the data name, the data unit, the time unit, and the collection location; and uploading the annotated cleaning results to the database for subsequent query.

[0022] A governance system for water supply plant detection data provided by the present invention, the system includes a memory, a processor, and a program stored on the memory and executable on the processor. When the program stored on the memory is run by the processor, the steps of the above-mentioned governance method for water supply plant detection data are implemented.

[0023] The present invention can clean the monitoring data of water quality monitoring parameters that are complex, diverse, of different types, and with different time units, and can ensure the accuracy and efficiency of data cleaning, providing a solid guarantee for water supply safety. Description of the Drawings

[0024] Figure 1 is a flowchart of a governance method for water supply plant detection data;

[0025] Figure 2 It is a detailed flowchart of a governance method for the detection data of a water supply plant;

[0026] Figure 3 It is a comparison chart of the outlier processing results of the original data and five outlier discrimination methods;

[0027] Figure 4 It is a flowchart of data cleaning using the second cleaning strategy;

[0028] Figure 5a It is a filling result chart of the test data using the recursive data increment ARIMA prediction model;

[0029] Figure 5b It is a filling result chart of the test data using the recursive data increment MLP prediction model;

[0030] Figure 5c It is a filling result chart of the test data using the recursive data increment ARIMA prediction model and the recursive data increment MLP prediction model;

[0031] Figure 6 It is a structural block diagram of a governance system for the detection data of a water supply plant. Specific implementation mode

[0032] The following will explain the embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the embodiments described below are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0033] The present invention provides a governance method and system for the monitoring data of a water supply plant. By customizing a standardized and automated cleaning process for different types of data, the accuracy and efficiency of data cleaning are achieved, adapting to complex, diverse, different types, and different time units, and being able to improve the data management and operation efficiency of the water supply plant. Among them, the cleaning process includes obtaining and unifying data specifications, automatically selecting outlier discrimination methods, null value replacement, judging and applying the first cleaning strategy based on the autoregressive integrated moving average model (ARIMA model) or the second cleaning strategy combining manual data complementary filling and the ARIMA model for data cleaning, and annotating and uploading the cleaning results to the database.

[0034] Embodiment 1

[0035] See Figure 1 A governance method for the detection data of a water supply plant may include the following steps:

[0036] Step S101: Obtain the monitoring data of each water service monitoring parameter and label at least the time unit for the monitoring data of each water service monitoring parameter.

[0037] Among them, the water service monitoring parameters include, but are not limited to, pH value, dissolved oxygen, conductivity, turbidity, temperature, chemical oxygen demand, and biochemical oxygen demand.

[0038] Among them, the monitoring data can be manually collected data (or called manually recorded data, manual data), can also be instrument online data (or simply called online data), and can also include both manually collected data and instrument online data.

[0039] Among them, the time unit can be every 5 minutes, every hour, daily, or can also be every second, every minute, weekly, monthly, and so on.

[0040] Furthermore, after obtaining the monitoring data of each water service monitoring parameter, the manually collected data and the instrument online data are standardized. The purpose of the standardization process in this embodiment is to unify the data specifications. Specifically, the text information in the monitoring data is replaced with null values.

[0041] Step S102: According to the evaluation criteria of the outlier discrimination method, select the best outlier discrimination method applicable to the monitoring data of a water service monitoring parameter from several outlier discrimination methods, and replace the outliers discriminated from the monitoring data of the water service monitoring parameter based on the best outlier discrimination method with null values.

[0042] Among them, the evaluation criteria of the outlier discrimination method is to select one outlier discrimination method with the best performance when discriminating outliers in the monitoring data of a water service monitoring parameter from several outlier discrimination methods. Taking any water service monitoring parameter as an example, arrange the monitoring data of the water service monitoring parameter in chronological order to obtain the time series (or called time series data) of the water service monitoring parameter, and through null value replacement, remove duplicates from the time series of the water service monitoring parameter. Generally, duplicate data are mostly instrument online data. Three or more consecutive identical data can be regarded as duplicate data, and the third and subsequent identical data are replaced with null values to ensure the accuracy and effectiveness of the data. Then, use each outlier discrimination method for outlier discrimination and performance evaluation to determine the best outlier discrimination method. Specifically, when implementing, use each outlier discrimination method to respectively perform outlier discrimination on the time series of the water service monitoring parameter, and according to the outliers discriminated by each outlier discrimination method, perform null value replacement processing of outliers, data repair processing of null values, and prediction processing of time series data on the time series of the water service monitoring parameter in sequence. Furthermore, perform performance evaluation on each outlier discrimination method, select the outlier discrimination method with the best performance from all outlier discrimination methods, and determine the outlier discrimination method with the best performance as the best outlier discrimination method for the water service monitoring parameter.

[0043] Among them, the several outlier identification methods at least include the moving average method, the STL decomposition method, the Z decomposition method, the exponential moving average method, and the isolation forest method.

[0044] Step S103: Select an appropriate cleaning strategy according to whether the monitoring data of the water service monitoring parameter is manually collected data, and adjust the cleaning parameters of the selected cleaning strategy according to the time unit of the monitoring data of the water service monitoring parameter.

[0045] Among them, it is judged whether the monitoring data of the water service monitoring parameter is manually collected data. If the monitoring data of the water service monitoring parameter is manually collected data, then the first cleaning strategy is selected. The first cleaning strategy applies the ARIMA model. If the monitoring data of the water service monitoring parameter is not manually collected data, that is, it is instrument online data, or even instrument online data for which corresponding manually collected data can be queried, then the second cleaning strategy is selected. The second cleaning strategy applies the manually collected data and / or applies the ARIMA model. Specifically, if corresponding manually collected data can be queried for the instrument online data, and the combination of the manually collected data and the instrument online data can fill all the null values in the time series, then the ARIMA model does not need to be applied. If corresponding manually collected data can be queried for the instrument online data, but the combination of the manually collected data and the instrument online data cannot completely fill all the null values in the time series, then the ARIMA model needs to be applied. If corresponding manually collected data cannot be queried for the instrument online data, then the ARIMA model needs to be applied.

[0046] Among them, according to the time unit of the monitoring data of the water service monitoring parameter, the autoregressive term coefficient p, the differencing order d, and the moving average term order q of the ARIMA model in the first cleaning strategy or the second cleaning strategy are adjusted. In this embodiment, according to past experience, when the time unit of the monitoring data of the water service monitoring parameter is "day", p = 5, d = 1, q = 5; when the time unit of the monitoring data of the water service monitoring parameter is "hour", p = 3, d = 1, q = 3; when the time unit of the monitoring data of the water service monitoring parameter is "5 minutes", p = 2, d = 1, q = 2. As the amount of data increases, the values of p, d, and q can be adjusted. Of course, the above settings can also be not changed.

[0047] Step S104: Use the cleaning strategy with the adjusted cleaning parameters to perform data cleaning including null value filling on the monitoring data of the water service monitoring parameter, and after performing data cleaning on the monitoring data of the water service monitoring parameter, label and save the cleaning result.

[0048] In one embodiment, if the cleaning strategy is the first cleaning strategy, the data cleaning includes: splitting the time series of the water service monitoring parameters to obtain multiple data segments, where the first and last values of each data segment are not null values and the last value is a null value; merging the data segments with null values to be filled and all the previous data segments with filled null values in a sequential manner, and using the ARIMA model to predict and fill the null values in the merged data segments until all the null values in all the data segments are filled. If the cleaning strategy is the second cleaning strategy, the data cleaning includes: querying the manually collected data corresponding to the online instrument data of the water service monitoring parameters. If no manually collected data corresponding to the online instrument data of the water service monitoring parameters is found, the first cleaning strategy with adjusted cleaning parameters is used to fill the null values in the time series of the water service monitoring parameters. If the manually collected data corresponding to the online instrument data of the water service monitoring parameters is found, the null values are initially filled by the complementary filling of the manually collected data and the online instrument data. After the initial filling of the null values, it is queried whether there are still unfilled null values. When it is found that there are still unfilled null values, the first cleaning strategy with adjusted cleaning parameters is used to perform secondary null value filling on the time series of the water service monitoring parameters.

[0049] The above embodiments do not consider the problems of long model processing time and large system CPU operation burden caused by the data volume, nor do they consider the problem of the missing seasonal variation characteristics of the data filled by the model when the number of missing values exceeds 5 consecutive time granularities. Therefore, another embodiment is provided here. Specifically, if the cleaning strategy is the first cleaning strategy, the data cleaning includes: splitting the time series of the water monitoring parameters to obtain multiple data segments, where the first and last positions of each data segment are not null values and the last position is a null value; determining whether to apply the MLP model according to the data volume of the time series of the water monitoring parameters; if the data volume is less than the preset amount, for example, 10,000 data records, the MLP model is not applied. At this time, the data segment to be filled with null values is merged with all the previous data segments that have been filled with null values in a consecutive manner, and the ARIMA model is used to predict and fill the null values in the merged data segment until all the null values in all the data segments are filled; if the data volume is greater than or equal to the preset amount, for example, 10,000 data records, the MLP model is applied. Then, the data segment to be filled with null values is merged with all the previous data segments that have been filled with null values in a consecutive manner. For the previous specified number of data segments, the ARIMA model is used to predict and fill the null values in the merged data segment, and for other data segments, the MLP model is used to predict and fill the null values in the merged data segment until all the null values in the other data segments are filled, where the data of the specified number of data segments is the data in the previous specified proportion (such as 15%, 20%, 25%, etc.) of the time series of the water monitoring parameters, and the specified number is reasonably set according to the processing situation of the ARIMA model and the actual data volume of the time series. If the cleaning strategy is the second cleaning strategy, the data cleaning includes: querying the manually collected data corresponding to the instrument online data of the water monitoring parameters. If no manually collected data corresponding to the instrument online data of the water monitoring parameters is queried, the first cleaning strategy with adjusted cleaning parameters is used to fill the null values in the time series of the water monitoring parameters. If the manually collected data corresponding to the instrument online data of the water monitoring parameters is queried, the null values are initially filled by the complementary filling method of the manually collected data and the instrument online data. After the initial filling of the null values, it is queried whether there are still unfilled null values. When it is queried that there are still unfilled null values, the first cleaning strategy with adjusted cleaning parameters is used to perform secondary null value filling on the time series of the water monitoring parameters.

[0050] Among them, using the MLP model to predict and fill in the null values in the merged data segments includes: extracting the time features of each null value in the merged data segments, then forming a time feature matrix for each null value based on the time features of each null value, and inputting the time feature matrix of each null value into the trained MLP model respectively to use the trained MLP model for null value prediction and fill in the null values with the predicted values. The time features referred to here include year, month, date, week, day of the year, quarter, number of weeks calculated using the ISO calendar system, and number of seconds since the Unix epoch.

[0051] Among them, after cleaning the monitoring data of the water service monitoring parameters, label the name of the best outlier discrimination method, the adopted cleaning strategy, data name, data unit, time unit, collection location for the cleaning result, and can further label the cleaning time, data source, outlier ratio, etc., and then upload the labeled cleaning result to the database.

[0052] By selecting the best outlier discrimination method for each water service monitoring parameter, the data governance method of the present invention is applicable to the complex and diverse monitoring data of water supply plants. By setting parameter values based on time units for some outlier discrimination methods and the ARIMA model, the data governance method of the present invention can adapt to the monitoring data at different times. In addition, during the data cleaning process, the two types of data, manual and online, can complement each other, and accurate cleaning can also be achieved for the data of the two data types, manual and online. In addition, when the data volume of the monitoring data of each water service monitoring parameter is greater than the preset amount, based on the ARIMA model, combined with the MLP model, the cleaning efficiency is improved while ensuring that the filled data maintains the seasonal change characteristics, ensuring the data accuracy.

[0053] Embodiment 2

[0054] The following combines Figures 2 to 4 、 Figures 5a to 5c to elaborate in detail on the method of the present invention.

[0055] Refer to Figure 2 A data governance method for manual and online monitoring data of water supply plants includes the following steps:

[0056] Step S201: Obtain and label the manually collected and instrument online data, and then unify the data specifications.

[0057] Among them, the labeled content covers the time unit (such as every day, every hour or every minute), data unit, and the process stage being collected (such as clear water tank).

[0058] Among them, unifying the data specifications can refer to replacing text information such as power outage and maintenance with null values.

[0059] Step S202: For the manually collected and instrument online data of each water monitoring parameter (hereinafter referred to as each type of data), automatically select the best outlier discrimination method according to the evaluation criteria.

[0060] The steps of automatically selecting the best outlier discrimination method in step S202 include:

[0061] Step S202-1: Data preprocessing.

[0062] First, sort the data in descending order according to time sequence to ensure the logicality of the analysis time sequence;

[0063] Secondly, de-duplicate the data according to the time variable to eliminate the situation where the same records may be repeatedly transmitted during the collection of online instrument data, and ensure the accuracy and effectiveness of data processing. For example, online data generally records up to three to six decimal places. The same record that appears more than three times continuously should be abnormal data and needs to be replaced with null values;

[0064] Step S202-2: Outlier discrimination methods, taking the following 5 types as examples.

[0065] 1. Moving average method

[0066] The moving average method identifies points whose difference from the rolling average exceeds four times the rolling standard deviation as outliers. At the same time, according to the time unit, adjust the window to enhance the sensitivity and smoothness of outlier detection. The value of the window determines the local range size of the time series data observed when calculating the moving average and moving standard deviation. A smaller window will make the moving average and standard deviation more sensitive to local changes, while a larger window will make them smoother to the global trend.

[0067] The specific settings are as follows:

[0068] Daily: window = 7;

[0069] Hourly: window = 24;

[0070] Every 5 minutes: window = 30.

[0071] 2. STL decomposition method

[0072] The STL decomposition method defines points whose absolute value of the residual exceeds four times the residual standard deviation as outliers through the residuals obtained by STL decomposition. And according to the time unit, adjust the period to obtain an accurate time series decomposition result, which will also improve the detection accuracy of outliers.

[0073] The specific settings are as follows:

[0074] Daily: period = 10;

[0075] Hourly: period = 24;

[0076] Every 5 minutes: period = 120.

[0077] 3. Z - score method

[0078] The Z - score method marks the points where the absolute value of the Z - score exceeds 4 as outliers.

[0079] 4. Exponential moving average method

[0080] The exponential moving average method regards the points that deviate from the exponential moving average by more than four times the standard deviation of the time series as outliers. The "time span" of the exponentially weighted moving average directly determines the relative importance of new data points and old data points in the weighted average, thus affecting the sensitivity of outlier detection.

[0081] The specific settings are as follows:

[0082] Daily: span = 10;

[0083] Hourly: span = 24;

[0084] Every 5 minutes: span = 120.

[0085] 5. Isolation forest method: Set the expected outlier proportion to 1%, and identify the points marked as - 1 as outliers according to the prediction of the isolation forest model.

[0086] Step S202 - 3: Outlier handling and data repair.

[0087] For each screening method, replace the identified outliers with null values, and then use a linear regression model to fill these missing values.

[0088] Step S202 - 4: Performance evaluation.

[0089] Use the ARIMA model to predict the time series after filling, and evaluate the performance of each outlier screening method by calculating the mean squared error (MSE).

[0090] Step S202 - 5: Selection of the best method.

[0091] By comparing the MSE values corresponding to the five methods, select the screening method with the smallest MSE value as the optimal choice.

[0092] Taking the original data of a water monitoring parameter from January 1, 2017 to January 31, 2017 as an example, using these original data as test data, see Figure 3, the abscissa represents time, the ordinate represents test data, the gray line represents the original data, and the black dots represent outliers. The five outlier screening methods pointed out in the above step S202-2 are respectively used to screen the outliers of the test data. By comparing Figure 3 the test data in Figure 3 with the results of different outlier screening methods, it can be found that there are significant differences in the outlier detection of the test data of this water monitoring parameter among these outlier screening methods. Among them, the Isolation Forest method detects the largest number of outliers, which may indicate that this method has a high sensitivity to the outliers in the current test data. In contrast, the Moving Average method detects the smallest number of outliers, which reflects that it may not be sensitive enough in the outlier detection of the current test data, but it does not mean that this method is completely unusable. At the same time, the number of outliers detected by the Z-Score method, the Exponential Moving Average (EMA) method, and the STL method is between the Isolation Forest method and the Moving Average method, indicating that their sensitivity to the outliers in the current test data is at a moderate level. To further evaluate the performance of these outlier processing methods, the present invention uses the data processed by linear interpolation, predicts them respectively using the ARIMA model, and compares their mean squared errors. The comparison results of the mean squared errors show that the Isolation Forest method is recognized as the best method in the current test data of this water monitoring parameter.

[0093] Step S203: Use the best outlier screening method for null value replacement.

[0094] Step S204: Determine whether the requirements of the first cleaning strategy can be met.

[0095] The specific content of the above step S204 is to determine whether the data is artificial data. If it is, the requirements of the first cleaning strategy are met; if not, the second cleaning strategy is performed.

[0096] Step S205: For the data that meets the requirements of the first cleaning strategy, use the first cleaning strategy for data cleaning.

[0097] Among them, the first cleaning strategy is a recursive data increment ARIMIA prediction model constructed based on the autoregressive integrated moving average model (ARIMA), and its key variables (p, d, q) are correspondingly modified according to the time unit marked in step S201.

[0098] The specific content of the above step S205 includes the following steps:

[0099] Step S205-1: Data segmentation.

[0100] Check the time series data, find the null values, and split the data into multiple segments accordingly. Each segment may have one or more consecutive null values that only appear at the end.

[0101] Step S205-2: Establish the basic prediction model.

[0102] Apply the recursive data increment ARIMA prediction model to the first data segment to obtain a basic prediction result and use it to fill the null values in the first data segment.

[0103] The specific process of the recursive data increment ARIMIA prediction model is as follows:

[0104] Step S205-2-1: Establish the ARIMA model.

[0105] Use the valid data part of the data segment to establish an ARIMA model:

[0106]

[0107] Where: Y t is the value of the time series at time t; p is the order of the autoregressive term; φ1, φ2,..., φ p are the autoregressive coefficients; ∈ t is the white noise error term; d is the order of differencing; q is the order of the moving average term; θ1, θ2,..., θ q are the moving average coefficients; μ is the mean of the time series; represents the first-order differencing operation.

[0108] Then, according to the time unit of the input data, modify p, d, q accordingly. The parameters are as follows:

[0109] Daily: p, d, q = 5, 1, 5;

[0110] Hourly: p, d, q = 3, 1, 3;

[0111] Every 5 minutes: p, d, q = 2, 1, 2;

[0112] Of course, it is also possible not to change the settings and directly use p, d, q = 1, 1, 1.

[0113] Step S205-2-2: Predict and fill the missing values.

[0114] Use the current ARIMA model to predict the missing values in the data segment and fill the predicted values to the original missing positions.

[0115] Step S205-2-3: Recursively optimize the model.

[0116] When the recursive data increment ARIMIA prediction model fills in each time, it only fills in one null value and needs to loop multiple times until all null values in the filled segment are filled. If there are still missing values in the data segment in step S205-1, the ARIMA model is reconstructed using the already filled data segment, and the prediction filling is performed again until all missing values in the data segment are completely eliminated.

[0117] Step S205-3: Recursively fill in the missing values.

[0118] Combine the filling result of step S205-2 with the next data segment, and use the recursive data increment ARIMA prediction model again to fill in the null values (or missing values).

[0119] Step S205-4: Complete the filling of all missing values.

[0120] Repeat step S205-3 until all null values (or missing values) are filled.

[0121] Step S206: For the data that does not meet the requirements of the first cleaning strategy, use the second cleaning strategy to clean the data.

[0122] Among them, the second cleaning strategy combines manual data complementary filling and the recursive data increment ARIMIA prediction model, and its key variables (p, d, q) are correspondingly modified according to the time unit marked in step S201.

[0123] See Figure 4 , the said step S206 includes the following steps:

[0124] Step S206-1: Verification of manual data records.

[0125] Check whether the online data has corresponding manual data records; if there is corresponding manual data, execute the process of step S206-2; if not, then turn to execute the operation of step S2066-4.

[0126] Step S206-2: Data time alignment and fusion.

[0127] Align the time of the manual data and the online data to ensure that their time units are consistent, and complement each other according to the time sequence to fill in the blank areas in the data.

[0128] Step S206-3: Secondary check for missing values.

[0129] Review the data set again to confirm whether there are still missing values. If there are still vacancies, continue to execute the process of step S206-4; if there are no missing values, you can directly enter the subsequent link of step S207.

[0130] Step S206-4: Implement the first cleaning strategy.

[0131] Adopt the first cleaning strategy to fill in the missing values.

[0132] Step S207: Label the cleaning results and upload them to the database for subsequent query.

[0133] An example of the labeled content is as follows: Outlier identification method name + cleaning strategy + data time unit + data name + collection location + data unit.

[0134] It should be noted that in the above process, the data volume is not considered. In practical applications, taking the online instrument data as an example, when processing the online instrument data with more than 10,000 records, the original recursive data incremental ARIMA prediction model has the problem of too long processing time, which not only increases the operation burden of the CPU but also may cause the computer system to restart. In addition, when filling in the missing values for more than 5 consecutive time granularities, the recursive data incremental ARIMA prediction model only uses the average value for filling, which ignores the seasonal pattern of the data and makes the filled data lose its original seasonal variation characteristics. To solve the above problems, it is necessary to optimize the foregoing cleaning strategy, that is, adopt a composite prediction method that combines the recursive data incremental ARIMA prediction model and the recursive data incremental MLP prediction model to replace the single prediction method that only uses the recursive data incremental ARIMA prediction model in the above process. Taking the data volume of the online instrument data exceeding the preset amount of 10,000 records as an example, step S206 specifically includes the following steps:

[0135] Step S206-1: Manual data record verification.

[0136] Check whether the online data has corresponding manual data records; if there are corresponding manual data, execute the process of step S206-2; if not, then turn to execute the operation of step S2066-4.

[0137] Step S206-2: Data time alignment and fusion.

[0138] Align the time of the manual data and the online data to ensure that their time units are consistent, and supplement each other according to the time sequence to fill in the blank areas in the data.

[0139] Step S206-3: Secondary check for missing values.

[0140] Review the data set again to confirm whether there are still missing values. If there are still vacancies, continue to execute the process of step S206-4; if there are no missing values, you can directly enter the subsequent link of step S207.

[0141] Step S206-4: Data segmentation.

[0142] Check the time series data, find the null values, and split the data into multiple segments accordingly. Each segment may have one or more consecutive null values that only appear at the end.

[0143] Step S206-5: Establish the basic prediction data.

[0144] Apply the recursive data increment ARIMA prediction model to the specified number of previous data segments (i.e., about the first 20% of the time series data) to obtain a basic prediction result with a large amount of data.

[0145] Step S206-6: Fill in the missing values with a neural network model.

[0146] Combine the basic prediction result of step S206-5 with the next data segment, and use the recursive data increment MLP prediction model to fill in the missing values in the current segment.

[0147] That is to say, first use the ARIMA model to fill the first 20% of the data, then splice this part of the filled data with the remaining unprocessed data, and then use the MLP model to further fill the entire spliced data, and finally complete the filling.

[0148] Step S206-7: Fill in the missing values in a loop.

[0149] Combine the prediction result of step S206-6 with the next data segment, and continue to use the recursive data increment MLP prediction model to fill in the missing values in the current segment.

[0150] Step S206-8: Complete filling all the missing values.

[0151] Repeat step S206-7 until all the missing values are filled.

[0152] Among them, the specific process of the recursive data increment MLP prediction model in step S206-6 is as follows:

[0153] Step S206-6-1: Data preprocessing.

[0154] Convert time into available features; extract the year (Year) from the time index; extract the month (Month) from the time index; extract the day (Day) from the time index; extract the day of the week (Weekday) from the time index, where Monday is 0 and Sunday is 6; extract the day of the year (DayOfYear) from the time index; extract the quarter (Quarter) from the time index; extract the week number (WeekNumber) from the time index, calculated using the ISO calendar system; convert the time index to the number of seconds since the Unix epoch (January 1, 1970).

[0155] Step S206-6-2: Delete the rows containing missing values for subsequent model training.

[0156] Check whether each row contains missing values through the dropna(inplace=True) function in the dataframe. If a row contains missing values, delete that row. This can provide a dataset without missing values for subsequent model training and ensure the normal progress of model training.

[0157] Step S206-6-3: Feature engineering.

[0158] Prepare the feature matrix X and the target vector y, where x only contains all the features in Step S206-6-1, and y is the valid data of the data segment.

[0159] Divide the dataset. The size of the training set is 80% of the original data, and the size of the test set is 20% of the original data.

[0160] Step S206-6-4: Feature standardization.

[0161] Perform standardization processing on the features.

[0162] Step S206-6-5: Model training and parameter tuning.

[0163] Define a multi-layer perceptron regression model (MLPRegressor);

[0164] Build a three-layer MLP model with 1 input layer, 1 output layer, and 1 hidden layer;

[0165] When running for the first time (initial=True), the program will start the GridSearchCV process. This process involves comprehensively traversing all hyperparameter combinations specified in param_grid and evaluating the performance of each hyperparameter combination through 5-fold cross-validation. Through this process, it aims to identify the hyperparameter combination that can achieve the best model performance. Among them, param_grid contains the following variables:

[0166] 'hidden_layer_sizes':[(10,),(20,),(50,)]: Defines the different numbers of neurons in the hidden layer. Here, three cases are considered: one hidden layer contains 10, 20, and 50 neurons respectively;

[0167] 'activation':['relu']: Specifies the activation function used in the hidden layer. Here, only the ReLU (Rectified Linear Unit) activation function is considered;

[0168] 'solver':['adam']: Specifies the optimization algorithm used to train the model. Here, only the Adam (Adaptive Moment Estimation) optimizer is considered;

[0169] 'alpha':[0.0001,0.001,0.01]: Defines the strength of L2 regularization (weight decay). Here, three different regularization coefficients are considered;

[0170] 'max_iter':[500,1000,1500]: Defines the maximum number of iterations (epochs) during the training process. Here, three different numbers of iterations are considered;

[0171] Train the model using the best combination of hyperparameters.

[0172] Step S206-6-6: Model saving and loading.

[0173] If it is the first run, save the best model; if not, load the saved model.

[0174] Step S206-6-7: Missing value imputation.

[0175] Use the trained model to predict missing values;

[0176] Step S206-6-8: Model evaluation.

[0177] Use the mean squared error (MSE) to evaluate the performance of the trained model on the test set.

[0178] Similarly, if the amount of manually collected data exceeds a preset amount, such as 10,000 records, steps S205-2 to S205-4 can also be replaced by steps S206-5 to S206-8 to improve efficiency and data prediction accuracy.

[0179] It can be seen that for the case of massive data, the present invention first uses a recursive data increment ARIMA prediction model to preliminarily fill a certain amount of missing data, and then takes this filling result as input and substitutes it into the recursive data increment MLP prediction model to complete the accurate prediction and filling of the remaining missing data.

[0180] Taking the original data of a water monitoring parameter from January 1, 2017 to January 31, 2017 (data missing from January 15, 2017 to January 17, 2017) as an example, using these original data as test data, the effect diagrams of data prediction and filling using the ARIMA model, MLP model, and the combination of ARIMA and MLP are shown in Figures 5a to 5c , where in each effect diagram, the abscissa represents time, the ordinate represents the test data, the gray line represents the original data, and the black dots represent the filled data. From Figure 5a it can be seen that the recursive data increment ARIMA model shows better results when dealing with small sample data. This model can relatively accurately predict and fill a single missing value. However, in the face of long continuous missing values in the time series, such as Figure 5a the long continuous missing values from January 15, 2017 to January 17, 2017 in, the filled values of the ARIMA model tend to be fixed, that is, it performs poorly in filling continuous missing values and has low efficiency. From Figure 5a and Figure 5b it can be seen that in the case of insufficient data sample size, in the time period before January 5, 2017, the prediction and filling results of the MLP model have a relatively large deviation compared with the ARIMA model and tend to exceed the original data range. From January 15, 2017 to January 17, 2017, at the same long continuous missing value points, the data filled by the MLP model can better capture the change trend of the original data compared with the ARIMA model. That is, compared with the ARIMA model, the MLP model shows more excellent prediction and filling performance when dealing with large-scale data sets, especially in the case of long continuous missing values. In addition, the operation time of the MLP model is relatively short and it is more efficient when dealing with large-scale data sets. By integrating the advantages of the ARIMA model in predicting and filling small sample data and the advantages of the MLP model in dealing with long continuous missing values, the adaptability of the model to the sample quantity can be significantly improved, so that the combined use of the ARIMA + MLP model can not only handle small sample data, but also maintain high efficiency on large-scale data sets, and can provide relatively accurate filling results under different data fluctuation conditions. See Figure 5c, compared with using ARIMA or MLP alone, the ARIMA+MLP filling method performs better in most cases, can better capture the change trend of the original data, and can maintain good consistency in both the areas with large data fluctuations and the relatively stable areas. It can be seen that the present invention is applicable to the governance of massive and complex data.

[0181] Embodiment III

[0182] Referring to FIG. 5, the present invention further provides a governance system 1 for the detection data of a water supply plant. The system 1 includes a memory 100, a processor 200, and a program stored on the memory 100 and operable on the processor 200. When the program stored on the memory 100 is run by the processor 200, it implements the steps of the governance method for the detection data of the water supply plant in the above-mentioned Embodiment I or the above-mentioned Embodiment II.

[0183] The system of the present invention can ensure the accuracy and efficiency of data cleaning and is suitable for the monitoring data with different time units, different data volumes, and different water service monitoring parameters.

[0184] Although the present invention has been described in detail above, the present invention is not limited thereto, and those skilled in the art of the present technology can make various modifications according to the principle of the present invention. Therefore, all modifications made according to the principle of the present invention should be understood to fall within the protection scope of the present invention.

Claims

1. A method for managing monitoring data of a water supply plant, characterized in that: The method comprises: Obtain monitoring data for each water service monitoring parameter, and mark the monitoring data for each water service monitoring parameter with at least a time unit; According to the evaluation criteria of the outlier identification method, an optimal outlier identification method applicable to the monitoring data of the water affairs monitoring parameter is selected from a plurality of outlier identification methods for a water affairs monitoring parameter, and an outlier identified from the monitoring data of the water affairs monitoring parameter based on the optimal outlier identification method is replaced with a null value; Selecting a suitable cleaning strategy according to whether the monitoring data of the water service monitoring parameter is manually collected data, and adjusting the cleaning parameters of the selected cleaning strategy according to the time unit of the monitoring data of the water service monitoring parameter; The cleaning strategy with adjusted cleaning parameters is used to perform data cleaning including null value filling on the monitoring data of the water affairs monitoring parameters, and after the data cleaning is performed on the monitoring data of the water affairs monitoring parameters, the cleaning results are marked and saved.

2. The method according to claim 1, characterized in that The method of selecting the best outlier identification method applicable to the monitoring data of any water affairs monitoring parameter according to the outlier identification method evaluation criteria includes: Arrange the monitoring data of the water affairs monitoring parameters in chronological order to obtain the time series of the water affairs monitoring parameters, and remove duplication of the time series of the water affairs monitoring parameters by replacing null values; Using each outlier identification method to identify outliers in the time series of the water affairs monitoring parameters, and according to the outliers identified by each outlier identification method, performing null value replacement processing, null value data repair processing, and time series data prediction processing on the time series of the water affairs monitoring parameters in sequence, so as to evaluate the performance of each outlier identification method; The outlier identification method with the best performance is selected from all outlier identification methods, and the outlier identification method with the best performance is determined as the best outlier identification method for the water affairs monitoring parameter.

3. The method according to claim 1 or 2, characterized in that: The several outlier identification methods include at least a moving average method, an STL decomposition method, a Z decomposition method, an exponential moving average method, and an isolation forest method.

4. The method according to claim 1, characterized in that: The selecting of a suitable cleaning strategy according to whether the monitoring data of the water service monitoring parameter is manually collected data includes: Determining whether the monitoring data of the water affairs monitoring parameter is manually collected data; If the data is collected manually, the first cleaning strategy of applying the ARIMA model is selected, otherwise, the second cleaning strategy of applying the manually collected data and / or applying the ARIMA model is selected.

5. The method according to claim 4, characterized in that The adjusting the cleaning parameters of the selected cleaning strategy according to the time unit of the monitoring data of the water service monitoring parameter comprises: According to the time unit of the monitoring data of the water service monitoring parameter, the autoregressive term coefficient p, the difference order d, and the moving average term order q of the ARIMA model in the first cleaning strategy or the second cleaning strategy are adjusted; wherein: The time unit of the monitoring data of the water affairs monitoring parameters is "day", p=5, d=1, q=5; When the time unit of the monitoring data of the water affairs monitoring parameter is "hour", p=3, d=1, q=3; When the time unit of the monitoring data of the water affairs monitoring parameter is "5 minutes", p=2, d=1, q=2.

6. The method according to claim 4 or 5, characterized in that: If the cleaning strategy is the first cleaning strategy, then the cleaning strategy with adjusted cleaning parameters is used to clean the monitoring data of the water service monitoring parameters including filling in null values, including: The time series of the water affairs monitoring parameter is segmented to obtain a plurality of data segments, wherein the first bit of each data segment is not a null value and the last bit is a null value; the data segment to be filled with null values ​​is merged with all previous data segments that have been filled with null values ​​in a continuous manner, and the null values ​​in the merged data segment are predicted and filled by using the ARIMA model until the null values ​​in all data segments are filled; Alternatively, data segmentation is performed on the time series of the water affairs monitoring parameters to obtain multiple data segments, wherein the first digit of each data segment is not a null value and the last digit is a null value; whether to apply the MLP model is determined based on the data volume of the time series of the water affairs monitoring parameters; if the MLP model is not applied, the data segments to be filled with null values ​​are merged with all previous data segments that have been filled with null values ​​in a consecutive manner, and the ARIMA model is used to predict and fill the null values ​​in the merged data segments until the null values ​​in all data segments are filled; if the MLP model is applied, the data segments to be filled with null values ​​are merged with all previous data segments that have been filled with null values ​​in a consecutive manner, and for a specified number of previous data segments, the ARIMA model is used to predict and fill the null values ​​in the merged data segments, and for other data segments, the MLP model is used to predict and fill the null values ​​in the merged data segments until the null values ​​in other data segments are filled.

7. The method according to claim 6, characterized in that If the cleaning strategy is the second cleaning strategy, then the cleaning strategy with adjusted cleaning parameters is used to perform data cleaning including null value filling on the monitoring data of the water service monitoring parameters, including: Querying manually collected data corresponding to the instrument online data of the water service monitoring parameter; If no manually collected data corresponding to the online instrument data of the water affairs monitoring parameter is found, then the first cleaning strategy with adjusted cleaning parameters is used to fill in the empty values ​​of the time series of the water affairs monitoring parameter; If manually collected data corresponding to the online data of the instrument of the water monitoring parameter is queried, the null values ​​are initially filled by complementarily filling the manually collected data with the online data of the instrument, and after the null values ​​are initially filled, it is queried whether there are any unfilled null values. When it is queried that there are still unfilled null values, the first cleaning strategy with adjusted cleaning parameters is used to perform secondary null value filling on the time series of the water monitoring parameter.

8. The method according to claim 6, characterized in that The method of using the MLP model to predict and fill in the empty values ​​in the merged data segments includes: Extracting time features of each null value in the merged data segment, the time features comprising year, month, date, week, day of the year, quarter, week number calculated using the ISO calendar system, and seconds since the Unix epoch; Based on the time feature of each null value, a time feature matrix of each null value is formed; The time feature matrix of each null value is respectively input into the trained MLP model to perform null value prediction using the trained MLP model, and to perform null value filling using the predicted value.

9. The method according to claim 1, characterized in that: The marking and saving of the cleaning results comprises: The cleaning results should at least be labeled with the name of the best outlier identification method, the cleaning strategy used, the data name, data unit, time unit, and collection location; Upload the marked cleaning results to the database for subsequent query.

10. A management system for water supply plant detection data, characterized in that: The system includes a memory, a processor, and a program stored in the memory and executable on the processor. When the program stored in the memory is executed by the processor, the program implements the steps of the method for managing water supply plant detection data as described in any one of claims 1 to 9.

Citation Information

Cited By

  • Data cleaning method and device, electronic equipment, storage medium and program product

    CN121255788A