A Diagnostic and Imputation Method for Abnormal Hydrological Data Based on the RF-Adaboost Model
Through the diagnosis and interpolation method based on the RF-Adaboost model, the problem of abnormal water condition data in open channel water diversion projects is solved, real-time and accurate prediction and interpolation of water condition data is achieved, and the data reliability and engineering scheduling accuracy are improved.
Patent Information
- Application Number
- CN202210116677.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-07
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-02-07
AI Technical Summary
In open channel water diversion projects, water condition data is abnormal due to equipment failure and human factors. Direct deletion or simple modification will lead to information loss and inaccurate data, affecting the gate hydraulic calculation and scheduling control.
The diagnostic and interpolation method based on the RF-Adaboost model is used to analyze the water condition data through median filtering, draw a 3-Sigma graph to identify abnormal data, and build and train the Adaboost model based on RF improvement to conduct real-time prediction and outlier interpolation of the water condition data.
It effectively improves the real-time diagnosis and prediction accuracy of water condition monitoring data, enhances the reliability and integrity of data, can objectively reflect changes in water condition, and guides engineering scheduling.
Smart Images

Figure CN114493023B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water regime monitoring, and particularly to a real-time anomaly diagnosis and interpolation method for flow monitoring data. Background Art
[0002] With the development of intelligent water transfer, a large number of automatic water regime monitoring devices have been installed along the built open-channel water transfer projects to monitor water regime information such as water level, flow rate, and gate opening. Due to various random device failures and regular manual maintenance and other interference factors, the water regime data may become abnormal. When facing abnormal water regime data, directly deleting it often causes a large amount of information loss, destroying the integrity, continuity, and consistency of the water regime data. Simple modification will seriously reduce the reliability and accuracy of the data set, and may have a greater impact on the hydraulic calculation and scheduling control of the gate in severe cases. Therefore, on the premise of ensuring the integrity, continuity, and consistency of the water regime data, timely identifying missing and abnormal data and performing prediction and interpolation have important application value and scientific significance for improving the reliability of the data, objectively reflecting the changes in the water regime, and effectively guiding the project scheduling.
[0003] Using a single machine learning method such as cubic spline (Spline) interpolation, random forest (RF) interpolation, etc. to perform interpolation prediction on abnormal data often contains a large subjective one-sidedness, and the interpolation prediction effect is very different from the measured value, unable to reflect the real water regime situation, and will be greatly limited when applied to the real-time interpolation of monitoring data. Therefore, effectively identifying abnormal data and using reasonable data for real-time interpolation prediction are the key problems to be solved in the water regime monitoring of open-channel water transfer projects.
[0004] Therefore, there is an urgent need to find a method for flow prediction and diagnosis to solve the above technical problems. Summary of the Invention
[0005] The purpose of the present invention is to provide a diagnosis and interpolation method based on the RF-Adaboost model for abnormal water regime data, so as to solve the foregoing problems existing in the prior art.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0007] A diagnosis and interpolation method based on the RF-Adaboost model for abnormal water regime data, comprising the following steps:
[0008] S1, perform median filtering analysis on the obtained water regime monitoring data to obtain the trend term and residual term of the time series;
[0009] S2, perform real-time diagnosis on abnormal data using the 3-Sigma chart method;
[0010] S3. Construct and train an Adaboost model improved based on RF. The input of the model is the time of the water regime data with noise values, missing values, and outliers removed, and the output is the flow rate of the water regime data.
[0011] S4. Apply the RF - Adaboost model to perform real - time prediction and imputation of outliers on the water regime data, and output the prediction variables.
[0012] Preferably, the prediction variables refer to real - time water regime monitoring data, including the water level and flow rate of the sluice gate.
[0013] Preferably, step S1 specifically includes:
[0014] S11. Define a long window with a length L = 3 for the obtained water regime data. Suppose at a certain moment, the signal samples within the window are x(i - 1), x(i), x(i + 1), where x(i) is the signal sample value at the center of the window. After arranging these 3 signal samples in ascending order of the sample values, the sample value at i is defined as the output value of the median filter, thereby eliminating isolated noise points, which is the trend term of the time series.
[0015] S12. Calculate the difference between the trend term of the time series and the measured value, which is the residual term.
[0016] Preferably, step S2 specifically includes:
[0017] The specific standard for real - time diagnosis is: data less than μ - βσ or greater than μ + βσ is abnormal data, where μ is the mean, σ is the standard deviation, and β is a variable parameter.
[0018] Preferably, step S3 specifically includes:
[0019] Given the training data set: (x1, y1),......, (x N , y N ), where y i ∈{1, - 1} is used to represent the class label of the training sample;
[0020] S31. Initialize the sample set weight as
[0021]
[0022] In the formula: D1(i) represents the initial weight distribution of the training sample set;
[0023] w 1i Each training sample is initially assigned the same weight;
[0024] N is the number of samples;
[0025] S32. Perform multiple rounds of iteration, where \(m = 1, 2, \cdots, N\) represents which round of iteration it is.
[0026] e) Select a random forest (RF) as the base classifier and learn from the training dataset with the data weight distribution \(D\). m Calculate the weak classifier \(G\). m (x);
[0027]
[0028] The above equation indicates that the weak classifier \(G\). m (x) classifies the sample data through a threshold. All data on one side of the threshold will be classified into class - 1, while data on the other side will be classified into class + 1.
[0029] f) Calculate the classification error rate of the weak classifier \(G\). m (x) on the training dataset.
[0030]
[0031] In the formula: \(e\). m Is the error rate.
[0032] \(P(G\). m (x i ) \(\neq y\). i ) is the probability that \(G\). m (x i ) \(\neq y\). i That is, the probability of misclassification.
[0033] \(I(G\). m (x i ) \(\neq y\). i ) means that if \(G\). m (x i ) \(= y\). i Then \(I = 1\), otherwise \(I = 0\).
[0034] \(w\). mi Is the weight assigned to the \(i\)-th training sample in the \(m\)-th round of iteration.
[0035] From this, it can be seen the relationship between the data weight distribution \(D\). m And the classification error rate of the weak classifier \(G\). m (x).
[0036] g) Calculate the coefficient of the weak classifier \(G\). m (x):
[0037]
[0038] \(\alpha\). m Represents the weak classifier \(G\). m(x) Importance in the final classifier. When α m ≥0, and α m increases as e m decreases;
[0039] At this time, the classifier is: f m (x) = α m G m (x)
[0040] h) Update the weight distribution of the training data for the next iteration:
[0041] D m+1 = (w m+1,1 , w m+1,2 ... w m+1,i ..., w m+1,N ) (4)
[0042]
[0043] Here, Z m is the normalization factor, making D m+1 a probability distribution:
[0044]
[0045] S33. Combine each weak classifier according to the weak classifier weight α m , that is
[0046]
[0047] Through the action of the sign function, a strong classifier is obtained as:
[0048]
[0049] Preferably, the RF - Adaboost model established in step S3 is:
[0050] Use random forest (RF) as the weak classifier to classify the sample data set of Adaboost;
[0051] After step S1 is completed, use the water regime data after median filtering to remove isolated noise points, and the measured water regime data after removing outliers after step S2 to train the constructed RF - Adaboost model;
[0052] The training data is the 2 - hour measured water regime data with a time span of one year;
[0053] In each round of training, a new weak classifier is obtained through classification by Random Forest (RF). That is, by changing the weights of the samples, especially the previously misclassified samples will get greater weights until the error rate is lower than the specified value or the preset maximum number of iterations is reached.
[0054] Preferably, the water regime monitoring data obtained in step S1 is updated to 2-hour time series data and input into the trained RF-Adaboost model in step S3. The output predicted variable values are the real-time prediction values and the correction values of the outliers.
[0055] The beneficial effects of the present invention are:
[0056] The present invention discloses a method for diagnosis and interpolation of abnormal water regime data based on the RF-Adaboost model, which includes obtaining water regime monitoring data, performing median filtering to remove obvious noise points, analyzing the residual terms; drawing a 3-Sigma graph, and real-time identifying and diagnosing abnormal data based on the 3-Sigma graph; constructing and training an Adaboost model improved based on Random Forest (RF); applying the RF-Adaboost model to perform real-time prediction on water regime monitoring data and interpolate abnormal data. Using this method can effectively improve the real-time diagnosis, prediction and interpolation of flow monitoring data, thereby improving the reliability of the data, objectively reflecting the changes in water regime, and effectively guiding project scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 is a schematic flow chart of the method for diagnosis and interpolation of abnormal water regime data based on the RF-Adaboost model in Embodiment 1;
[0058] Figure 2 is the result of identifying and diagnosing outliers in the flow data of the check gate by the 3-Sigma graph in Embodiment 1;
[0059] Figure 3 is the result after predicting and interpolating the measured flow data of the check gate using the RF-Adaboost model in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0061] Embodiment 1
[0062] This embodiment adopts a specific example, taking the middle route of the South-to-North Water Diversion Project as an example, to provide a method for diagnosis and interpolation of abnormal water regime data based on the RF-Adaboost model, which specifically includes the following steps:
[0063] Step 1, perform median filtering analysis on the obtained water regime monitoring data to obtain the trend term and residual term of the time series.
[0064] Affected by various factors such as the hydraulic characteristics of the river channel, the operation period of the project, and the labor intensity of personnel in the open channel water diversion project, the water regime monitoring frequency of the open channel water diversion project is 2 hours, which can objectively reflect the water regime changes and effectively guide the project dispatching. Select the flow data of the Hutuo River Check Gate from January 1, 2018 to December 31, 2018 as the research object.
[0065] Adopt the median filtering method to extract the trend term and residual term of the flow data time series. The frequency distribution diagram of the residual term obtained after the flow data is median filtered shows obvious normal distribution characteristics.
[0066] Step 2, use the 3-Sigma chart method to perform real-time diagnosis on abnormal data.
[0067] Draw a 3-Sigma chart to depict the discrete distribution of the data. To meet the abnormal data diagnosis needs of different data sets, introduce the parameter β into the 3-Sigma method, and use the data less than μ - βσ or greater than μ + βσ as the judgment standard for abnormal data, where μ is the mean value, σ is the standard deviation, and β = 1.62. The abnormal value results of the check gate flow data are as Figure 2 .
[0068] Step 3, construct and train a model based on RF-Adaboost:
[0069] Given the training data set: (x1, y1),......, (x N , y N ), where y i ∈{1, -1} is used to represent the class label of the training sample.
[0070] (1) Initialize the sample set weight as
[0071]
[0072] In the formula: D1(i) is the initial weight distribution of the training sample set;
[0073] w 1i is the same weight initially assigned to each training sample;
[0074] N is the number of samples;
[0075] (2) Perform multiple rounds of iteration, and use m = 1, 2,..., N to represent which round of iteration
[0076] ① Select the random forest (RF) as the basic classifier, and learn from the training data set with weights D m to calculate the weak classifier G m (x);
[0077]
[0078] The above formula indicates that the weak classifier G m (x) classifies the sample data through a threshold. All data on one side of the threshold will be assigned to class - 1, while data on the other side will be assigned to class + 1;
[0079] ② Calculate the classification error rate of G m (x) on the training data set
[0080]
[0081] In the formula: e m is the error rate;
[0082] P(G m (x i ) ≠ y i ) is the probability that G m (x i ) ≠ y i , that is, the probability of misclassification;
[0083] I(G m (x i ) ≠ y i ): That is, if G m (x i ) = y i , then I = 1, otherwise I = 0;
[0084] w mi is the weight assigned to the i - th training sample in the m - th round of iteration;
[0085] From this, it can be seen the relationship between the data weight distribution D m and the classification error rate of the basic classifier G m (x).
[0086] ③ Calculate the coefficient of the weak classifier G m (x):
[0087]
[0088] α m represents the importance of G m (x) in the final classifier. When When a m is 30, and α m increases as e m decreases, the smaller the classification error rate, the greater the role of the basic classifier in the final classifier.
[0089] At this time, the classifier is: f m (x) = α m G m (x)
[0090] ④ Update the weight distribution of the training data for the next iteration.
[0091] D m+1 = (w m+1,1 , w m+1,2 ... w m+1,i ..., w m+1,N ) (4)
[0092]
[0093] Here, Z m is the normalization factor, making D m+1 become a probability distribution:
[0094]
[0095] It can be seen from this that the weights of the samples misclassified by the weak classifier G m (x) are enlarged, while the weights of the samples correctly classified are reduced.
[0096] (3) Combine each weak classifier according to the weak classifier weight α m , that is
[0097]
[0098] Through the action of the sign function, a strong classifier is obtained as:
[0099]
[0100] The input of the model is the time series of normal data after removing the noise values by median filtering and the outliers diagnosed by the 3-Sigma diagram, and the output is the flow value of the normal data. A new weak classifier will be added in each round of training, that is, the weights of the samples are changed through the decision tree. In particular, the previously misclassified samples will get greater weights until a sufficiently low error rate is reached or the specified maximum number of iterations is reached.
[0101] Step 4, Apply the RF-Adaboost model to predict and interpolate the noise values and outliers of the flow monitoring.
[0102] Use the 2-hour time series data of the water regime monitoring data obtained in step 1 as the input of the RF-Adaboost model, and apply it to the RF-Adaboost model trained in step S3. The output prediction variable values are the real-time prediction values and the corrected values of the outliers, such as Figure 3 .
[0103] By adopting the above technical solutions disclosed in the present invention, the following beneficial effects are obtained:
[0104] The present invention discloses a diagnosis and interpolation method for abnormal water regime data based on the RF-Adaboost model, which relates to the technical field of water regime monitoring. The method includes the following steps: obtaining real-time water regime data, performing time series analysis on the water regime data by using median filtering, drawing a 3-Sigma chart, and performing real-time identification and diagnosis of abnormal data based on the 3-Sigma chart; constructing and training an Adaboost model improved based on the random forest (RF), applying the RF-Adaboost model to perform real-time prediction on the water regime monitoring data, and interpolating the abnormal data. By adopting this method, the real-time prediction and monitoring of the water regime monitoring data can be effectively improved, the abnormal data can be diagnosed and interpolated in a timely manner, so as to improve the reliability of the data, objectively reflect the change of the water regime, and effectively guide the project scheduling.
[0105] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A diagnostic and interpolation method for abnormal water regime data based on the RF - Adaboost model, characterized in that, It includes the following steps: S1. Perform median filtering analysis on the obtained water regime monitoring data to obtain the trend term and residual term of the time series; S2. Use the 3-Sigma chart method to perform real-time diagnosis on abnormal data; S3. Construct and train an Adaboost model improved based on RF. The input of the model is the moment of water regime data after removing noise values, missing values and abnormal values, and the output is the flow rate of the water regime data; S4. Apply the RF-Adaboost model to perform real-time prediction and interpolation of abnormal values on the water regime data, and output the prediction variable; The prediction variable refers to real-time water regime monitoring data, including the water level and flow rate of the regulating sluice; Step S1 specifically includes: S11. Define a long window with a length L = 3 for the obtained water regime data; Suppose at a certain moment, the signal samples within the window are x(i - 1), x(i), x(i + 1), where x(i) is the signal sample value located at the center of the window; After arranging these 3 signal samples in ascending order of the sample value, the sample value at i is defined as the output value of the median filtering, so as to eliminate isolated noise points, which is the trend term of the time series; S12. Calculate the difference between the trend term of the time series and the measured value, which is the residual term; Step S2 specifically includes: The specific standard for real-time diagnosis is: data less than μ - βσ or greater than μ + βσ is abnormal data, where μ is the mean value, σ is the standard deviation, and β is a variable parameter; Step S3 specifically includes: Given training data set: (x1, y1),......, (x N , y N ), where y i ∈ {1, -1} is used to represent the class label of the training sample; S31. Initialize the sample set weight as In the formula: D1(i) represents the initial weight distribution of the training sample set; w 1i Each training sample is initially assigned the same weight; N is the number of samples; S32. Perform multiple rounds of iteration, and use m = 1, 2,..., N to represent which round of iteration; a) Select RF as the basic classifier, and learn from the training data set with the data weight distribution D m to calculate the weak classifier G m (x); The above formula represents the weak classifier G m (x) classifies the sample data by a threshold, and all data on one side of the threshold will be assigned to class -1, while data on the other side will be assigned to class +1; b) Calculate the weak classifier G m (x) on the training data set where: e m is the error rate; P(G m (x i )≠y i ) is the probability of G m (x i )≠y i , that is, the probability of misclassification; I(G m (x i )≠y i ):that is, if G m (x i )=y i then I = 1, otherwise I = 0; w mi is the weight assigned to the i-th training sample in the m-th iteration; It can be seen therefrom that the data weight distribution D m and the weak classifier G m (x) in terms of the relationship of the classification error rate; c) Calculate the coefficient of the weak classifier G m (x): α m indicates the importance of the weak classifier G m (x) in the final classifier. When , α m ≥0, and α m increases as e m decreases; At this time, the classifier is: f m (x) = α m G m (x) d) Update the weight distribution of the training data for the next round of iteration: D m+1 = (w m+1,1 , w m+1,2 ... w m+1,i ..., w m+1,N ) (4) Here, Z m is a normalization factor such that D m+1 becomes a probability distribution: S33, according to the weak classifier weight α m Combine each weak classifier, that is Through the action of the sign function, a strong classifier is obtained as:
2. The diagnostic and interpolation method for abnormal water regime data based on the RF - Adaboost model according to claim 1, characterized in that, The RF-Adaboost model established in step S3 is: Use RF as the basic classifier to classify the sample data set of Adaboost; After step S1 is completed, use the water regime data after median filtering to remove isolated noise points, and the measured water regime data after removing abnormal values after step S2 to train the constructed RF-Adaboost model; The training data is 2-hour measured water regime data with a time span of one year; In each round of training, a new weak classifier will be obtained through classification by RF, that is, by changing the weights of the samples, the previously misclassified samples will get greater weights until the error rate is lower than the specified value or the preset maximum number of iterations is reached.
3. The diagnostic and interpolation method for abnormal water regime data based on the RF - Adaboost model according to claim 1, characterized in that, Update the obtained water regime monitoring data in step S1 to 2-hour time series data, input it into the Adaboost model improved based on RF trained in step S3, and the output prediction variable value is the real-time prediction value and the corrected value of the abnormal value.
Citation Information
Patent Citations
Transformer fault diagnosis method based on improved multi-class AdaBoost
CN108229581A
Real-time anomaly diagnosis and interpolation method for water regime monitoring data
CN111307123A