Dam safety monitoring data error identification, classification and correction method

By combining a variety of advanced algorithms (LOF, CEEMDAN, CUMSUM, SSA and RF), automated error identification, classification and correction of dam safety monitoring data is achieved, and the problems of low processing efficiency and inaccurate error identification are solved, improving data processing capabilities and analysis accuracy.

CN120011854APending Publication Date: 2025-05-16CHINA THREE GORGES CORPORATION +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510062054.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Among the existing dam safety monitoring technology, there are problems such as low efficiency in monitoring data processing, inaccurate error identification, and incomplete data processing.

Method used

Local outlier factor (LOF) algorithm, complete empirical modal decomposition (CEEMDAN) method of adaptive noise, accumulation sum (CUMSUM) algorithm, singular spectrum analysis (SSA) algorithm and random forest (RF) algorithm are used to realize automated error identification, classification and correction of monitoring data.

Benefits of technology

It significantly improves the data processing capability and analysis accuracy of the dam safety monitoring system, providing more scientific and reliable technical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011854A_ABST
    Figure CN120011854A_ABST
Patent Text Reader

Abstract

The invention discloses a dam safety monitoring data error identification, classification and correction method. The method comprises the following steps: (1) selecting a data sequence and collecting influence factors; (2) carrying out gross error identification; (3) judging a fluctuation error; (4) detecting abrupt change points in the monitoring data by adopting a cumulative sum CUMSUM algorithm; (5) processing fluctuation errors; (6) complementing missing values; according to the invention, more scientific and reliable technical support is provided for dam safety management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water conservancy engineering, and specifically to a method for identifying, classifying and correcting errors in dam safety monitoring data. Background Art

[0002] As an important water conservancy project facility, the safety of dams is directly related to the safety of life and property of people in the downstream area and the protection of the ecological environment. Therefore, dam safety monitoring is a vital task, which mainly uses various sensors and monitoring equipment to monitor the displacement, seepage, uplift pressure and other data of the dam in real time, and evaluates the structural safety of the dam based on the monitoring data.

[0003] Gross errors and noise often exist in monitoring data, and these interference factors will affect the accuracy of data analysis. Commonly used methods for identifying gross errors include outlier detection and statistical analysis, but the accuracy and robustness of existing methods still need to be improved for large-scale data and complex environments. Secondly, there are various types of errors in monitoring data, including systematic errors, random errors, and gross errors. Existing error processing methods often focus on gross errors, and there is little research on methods for comprehensively processing multiple errors. Summary of the invention

[0004] Purpose of the invention: The purpose of the present invention is to provide a method for error identification, classification and correction of dam safety monitoring data, which solves the problems of low monitoring data processing efficiency, inaccurate error identification and incomplete data missing processing existing in the prior art by introducing a variety of advanced algorithms and technologies, including local outlier factor (LOF) algorithm, complete empirical mode decomposition of adaptive noise (CEEMDAN) method, cumulative sum (CUMSUM) algorithm, singular spectrum analysis (SSA) algorithm and random forest (RF) algorithm.

[0005] Technical solution: A method for identifying, classifying and correcting dam safety monitoring data errors according to the present invention comprises the following steps:

[0006] (1) Select data series and collect impact factors;

[0007] (2) Identify gross errors;

[0008] (3) Determine volatility error;

[0009] (4) Using the cumulative CUMSUM algorithm to detect mutation points in monitoring data;

[0010] (5) Dealing with volatility errors;

[0011] (6)Fill in missing values.

[0012] Furthermore, step (1) is specifically as follows: taking a preset time length as a frequency interval, continuously selecting monitoring data of n preset time lengths to form a data sequence; and collecting a corresponding influencing factor sequence; wherein the data sequence includes displacement, seepage or uplift pressure data, and its influencing factors include water level, temperature and time factor.

[0013] Furthermore, step (2) is specifically as follows: in the process of automatically identifying continuous points of dam monitoring data, the local outlier factor LOF algorithm is first used to identify gross errors.

[0014] Furthermore, step (3) is specifically as follows: using the complete empirical mode decomposition (CEEMDAN) algorithm based on adaptive noise to decompose the gross error identification data; wherein the CEEMDAN method introduces a noise signal; performing a volatility analysis on the residual term after decomposition to determine whether there is a fluctuation error.

[0015] Furthermore, step (4) is as follows: the CUMSUM algorithm compares each data point with the target value, calculates the deviation, and then accumulates these deviations; if the cumulative deviation exceeds the set threshold, it indicates that the data has changed significantly; wherein the target value is the mean.

[0016] Furthermore, step (5) is specifically as follows: for the volatility error, the singular spectrum analysis SSA algorithm is used to further extract and reconstruct the high-frequency IMF obtained after CEEMDAN decomposition, and the remaining high-frequency components are eliminated and accumulated to obtain the dam monitoring data without the volatility error.

[0017] Furthermore, step (6) is specifically as follows: for the missing values ​​after gross errors are eliminated, fitting is performed based on the random forest RF method using various environmental quantity data as input and the data to be processed as output, and finally the missing values ​​of the monitoring data are filled in according to the environmental quantity data.

[0018] A dam safety monitoring data error identification, classification and correction system according to the present invention comprises:

[0019] Acquisition module: used to select data sequences and collect influencing factors;

[0020] Gross error module: used for gross error identification;

[0021] Discrimination module: used to distinguish volatility errors;

[0022] Detection module: used to detect mutation points in monitoring data using the cumulative CUMSUM algorithm;

[0023] Processing module: used to process volatility errors;

[0024] Processing module: used to fill missing values.

[0025] An electronic device described in the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, it implements any one of the methods for identifying, classifying, and correcting errors in dam safety monitoring data.

[0026] A storage medium described in the present invention stores a computer program, and when the computer program is executed by a processor, it implements a dam safety monitoring data error identification, classification and correction method according to any one of the items.

[0027] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: by introducing a variety of advanced algorithms and technologies, including the local outlier factor (LOF) algorithm, the complete empirical mode decomposition of adaptive noise (CEEMDAN) method, the cumulative sum (CUMSUM) algorithm, the singular spectrum analysis (SSA) algorithm and the random forest (RF) algorithm, the automatic error identification, classification and correction of monitoring data can be achieved. Through the method of the present invention, the data processing capability and analysis accuracy of the dam safety monitoring system can be significantly improved, providing more scientific and reliable technical support for dam safety management. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a technical flow chart of the present invention;

[0029] Figure 2 It is a schematic diagram of the principle of fluctuation error identification and correction in the present invention. DETAILED DESCRIPTION

[0030] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0031] like Figure 1 As shown, an embodiment of the present invention provides a method for identifying, classifying and correcting dam safety monitoring data errors, comprising the following steps:

[0032] S1: Continuously select monitoring data of n preset time intervals to form a data sequence, and collect the corresponding influencing factor sequence. The monitoring data sequence can be displacement, seepage or uplift pressure data, and its influencing factors include water level, temperature and time factor. Taking the dam displacement monitoring data as an example, it mainly consists of three parts: water pressure component, temperature component and time component, namely

[0033] δ=δ H +δ T +δ θ

[0034] Where: δ is the displacement fitting value; δ H , δ T and δ θThey are water pressure component, temperature component and time aging component respectively.

[0035] S11: The water pressure component is the deformation of the dam under the action of reservoir water load. From engineering mechanics and dam construction theory, it can be known that the water pressure component of the displacement of the gravity dam δ H Mainly related to the water depth H, H 2 and H 3 about, that is

[0036]

[0037] Where: H is the water depth in front of the dam; α i is the regression coefficient of water pressure factor, i=1~3.

[0038] S12: The temperature component is the displacement caused by the temperature change of the dam concrete and bedrock. Depending on the layout of the thermometers, the temperature component can be calculated in the following two ways:

[0039] (a) When a sufficient number of thermometers are placed in the dam body and bedrock, it can be expressed as

[0040] or

[0041] Where: m1 is the number of thermometers; T i is the measured temperature; m2 is the equivalent temperature count; is the equivalent average temperature measured by the thermometer at the i-th layer; β i is the temperature gradient of the thermometer at the i-th layer; b i , b 1i , b 2i is a statistical parameter.

[0042] (b) When the dam has been in operation for many years and the heat of concrete hydration has been completely dissipated or only air temperature observation data are available, the temperature component can be expressed as

[0043]

[0044] Where: t is the number of days from the monitoring day to the initial monitoring day; b 1i , b 2i is the temperature factor regression coefficient.

[0045] S13: The time-dependent component can be considered to be mainly caused by the creep of concrete and bedrock and the plastic deformation of cracks and joints under pressure. The time-dependent component of a normally operating dam has a law of rapid changes in the early stage and gradually stabilization in the later stage. The time-dependent component is often expressed as a combination of linear function and logarithmic function, which can be written as

[0046] δ θ =c1θ+c2lnθ

[0047] Wherein: c1, c2 are regression coefficients of time-effect factor; θ = t / 100, t has the same meaning as above.

[0048] S2: In the process of automatically identifying continuous points of deformation monitoring data, the LOF algorithm is first used to identify gross errors. The main concepts involved in the LOF algorithm are the k-distance of data objects, k-distance neighborhood, reachable distance of data objects, reachable density and local outlier factor. The relevant concepts are defined as follows:

[0049] k distance d of object p k (p): Let k be a positive integer, and the k-distance of data object p is denoted by d k (p). In the data set D, the distance between two data objects p and o is denoted as d(p,o). k (p)=d(p,o), must satisfy:

[0050] (c) There are at least k points o'∈C / {p} in the data set D that do not include p, satisfying d(p,o')≤d(p,o).

[0051] (d) There are at most k-1 points o'∈C / {p} in the dataset D that do not contain p, satisfying d(p,o') <d(p,o)。

[0052] The k-th distance neighborhood of object p: The k-th distance neighborhood of data object p is the set of all data objects o whose distance from p is less than or equal to k, that is:

[0053] N k (p)={o∈Dd(p,o)≤d k (p)}

[0054] Reachable distance: The reachable distance between data objects p and o is recorded as:

[0055] reach-dist k (p,o)=max{d k (p),d(p,o)}

[0056] Local reachability density: The local reachability density of a data object p represents the inverse of the average reachable distance from point p to point p in the kth neighborhood, that is:

[0057]

[0058] Local outlier factor: The local outlier factor of a data object p represents the neighborhood N of point p k The average of the ratios of the local reachable density of point (p) to the local reachable density of point p, that is:

[0059]

[0060] If the LOF value is closer to 1, it means that the density of p is closer to that of its neighborhood objects, and it is more likely to belong to the same cluster as the neighborhood; if the LOF value is greater than 1, it means that the density of p is less than the density of its neighborhood points, and it is more likely to be an outlier.

[0061] S3: After removing the gross errors in the monitoring data, the complete empirical mode decomposition with adaptive noise (CEEMDAN) method is used to identify data fluctuation errors. CEEMDAN is an improved empirical mode decomposition (EMD) method that enhances the robustness of EMD by introducing noise signals, effectively alleviating the problems of modal aliasing and uneven distribution of modal gaps that may occur when processing non-stationary signals. The specific process of CEEMDAN to identify fluctuation system errors is as follows:

[0062] S31: Introduce multiple sets of different white noise n into the original signal x(t) i (t), generate multiple noisy signal samples x i (t) = x(t) + n i (t).

[0063] S32: For each noisy signal sample x i (t) Perform empirical mode decomposition and extract the first intrinsic mode function IMF, denoted as IMF i1 (t).

[0064] S33: Calculate the average of the first IMF of all noisy signal samples to obtain the first IMF, which is recorded as

[0065]

[0066] S34: Subtract the first IMF from the original signal x(t) to obtain a residual signal r1(t)=x(t)-IMF1(t).

[0067] S35: Introduce white noise into the residual signal r1(t) again to generate multiple groups of noisy residual signal samples r 1i (t) = r1(t) + n i (t). Perform EMD decomposition on these noisy residual signal samples and extract their first IMF, recorded as IMF1(t). Average the first IMFs of all noisy residual signals to obtain the second IMF, recorded as IMF2(t). Subtract the second IMF from the residual signal to obtain a new residual signal r2(t)=r1(t)-IMF2(t).

[0068] S36: Repeat the above steps and continue to perform CEEMDAN decomposition on the new residual signal until all predetermined IMFs are decomposed or the residual signal becomes stable.

[0069] S37: Perform volatility analysis on the final residual term. Calculate the statistical characteristics of the residual term, such as standard deviation and variance, to evaluate its volatility. If the volatility of the residual term is large, it means that there is a significant volatility error in the signal.

[0070] According to the preset volatility threshold, it is determined whether the volatility of the residual term exceeds the normal range. If the volatility of the residual term exceeds the threshold, it is determined that there is a volatility error in the signal.

[0071] S4: Use the Cumulative Sum (CUMSUM) algorithm to detect mutation points in error data. The CUMSUM algorithm detects change points by calculating the deviation of the cumulative sum of data points. The basic idea is to compare each data point with the target value (usually the mean), calculate the deviation, and then accumulate these deviations. If the cumulative deviation exceeds the set threshold, it indicates that the data has changed significantly. The specific steps are as follows:

[0072] S41: Calculate target value: Calculate the mean of the monitoring data as the target value.

[0073] S42: Calculate deviation: For each data point, calculate its deviation from the target value.

[0074] S43: Cumulative deviation: Accumulate the deviation of each data point to form a cumulative sum series.

[0075] S44: Setting a threshold: setting a cumulative sum threshold according to actual needs.

[0076] S45: Detect change point: When the cumulative sum exceeds the set threshold, the point is marked as a jump point, indicating that the data has changed significantly at this point.

[0077] S5: Correct the fluctuation error in the monitoring data. For the decomposition results in S3, the continuous mean square error (CMSE) strategy is used to evaluate the obtained IMF signal. The CMSE strategy is essentially a technique for evaluating the square accumulation of the amplitude of the corresponding IMF component at each point in the time domain, which can be regarded as a characterization of the energy density of the IMF component. This strategy can accurately quantify the energy distribution and concentration of the IMF component in the time domain. When the error of the kth signal shows a significant change, this point is regarded as the boundary between the signal and the noise. The calculation method of CMSE is as follows:

[0078]

[0079] In the formula, represents the partially reconstructed signal, n is the signal length, and N is the number of IMFs.

[0080] The above method is used to calculate the CMSE between adjacent IMFs and plot them into a line graph. The inflection point where the slope of the line changes the most is the dividing point between the high-frequency and low-frequency IMFs. All high-frequency IMFs are superimposed to obtain the overall high-frequency component as the input of the singular spectrum analysis (SSA), and the trend characteristics of the original signal contained therein are further analyzed to reconstruct the effective components. After completing this step, the overall high-frequency component extracted by SSA is integrated with all low-frequency IMFs obtained by CEEMDAN, and finally the data with volatility errors removed after deep feature re-extraction and reconstruction is obtained. The entire CEEMDAN-SSA denoising process is as follows: Figure 2 shown.

[0081] S6: For the missing values ​​after the gross error is eliminated, the random forest (RF) method is used to complete the data. The environmental quantity data is used as input and the data to be processed is used as output for fitting. Finally, the missing values ​​of the monitoring data are completed based on the environmental quantity data. The random forest algorithm is based on statistical learning theory and uses the self-help resampling technology to extract multiple sample sets from the training samples. The extracted sample sets are used to build decision tree models respectively, and then several decision trees are clustered together to obtain the final result through majority voting or averaging. It has the following characteristics:

[0082] S61: Ensemble learning: Random forest can effectively reduce the bias and variance of a single model by integrating the prediction results of multiple decision tree models. Each decision tree has a certain degree of randomness in the generation process, so it can capture different patterns of data.

[0083] S62: Self-service resampling: Self-service resampling generates multiple different sub-datasets, each of which is used to train a decision tree model. This method increases the diversity of the model, allowing each tree to provide a different perspective.

[0084] S63: Random feature selection: In the process of generating a decision tree, each time a node is selected for partitioning, a portion of features is randomly selected for optimal partitioning. This random feature selection mechanism further enhances the diversity of the model and reduces the overfitting problem caused by feature correlation.

[0085] S64: Model prediction: Random forest can effectively improve the stability and accuracy of prediction by combining the prediction results of multiple decision trees. The final prediction value is obtained by averaging the prediction results of all decision trees.

[0086] S65: Use the predicted value of RF to fill in the missing data to obtain a complete monitoring sequence.

Claims

1. A method for identifying, classifying and correcting errors in dam safety monitoring data, characterized in that: The following steps are involved: (1) Select data series and collect impact factors; (2) Identify gross errors; (3) Determine volatility error; (4) Using the CUMSUM algorithm to detect mutation points in monitoring data; (5) Dealing with volatility errors; (6)Fill in missing values.

2. A method for identifying, classifying and correcting dam safety monitoring data errors according to claim 1, characterized in that: Step (1) is as follows: taking a preset time length as a frequency interval, continuously selecting monitoring data of n preset time lengths to form a data sequence; and collecting a corresponding influencing factor sequence; wherein the data sequence includes displacement, seepage or uplift pressure data, and its influencing factors include water level, temperature and aging factor.

3. A method for identifying, classifying and correcting dam safety monitoring data errors according to claim 1, characterized in that: Step (2) is as follows: In the process of automatically identifying continuous points of dam monitoring data, the local outlier factor LOF algorithm is first used to identify gross errors.

4. A method for identifying, classifying and correcting dam safety monitoring data errors according to claim 1, characterized in that: Step (3) is as follows: the gross error identification data is decomposed using the complete empirical mode decomposition (CEEMDAN) algorithm based on adaptive noise; wherein the CEEMDAN method introduces noise signals; and the residual term after decomposition is subjected to volatility analysis to determine whether there is a fluctuation error.

5. A method for identifying, classifying and correcting dam safety monitoring data errors according to claim 1, characterized in that: Step (4) is as follows: The CUMSUM algorithm compares each data point with the target value, calculates the deviation, and then accumulates these deviations; if the cumulative deviation exceeds the set threshold, it indicates that the data has changed significantly; where the target value is the mean.

6. A method for identifying, classifying and correcting dam safety monitoring data errors according to claim 1, characterized in that: Step (5) is as follows: for the volatility error, the singular spectrum analysis SSA algorithm is used to further extract and reconstruct the high-frequency IMF obtained after CEEMDAN decomposition, and the remaining high-frequency components are eliminated and accumulated to obtain the dam monitoring data without volatility error.

7. A method for identifying, classifying and correcting dam safety monitoring data errors according to claim 1, characterized in that: Step (6) is as follows: for the missing values ​​after gross errors are eliminated, the random forest RF method is used to fit the environmental quantity data as input and the data to be processed as output, and finally the missing values ​​of the monitoring data are filled in according to the environmental quantity data.

8. A dam safety monitoring data error identification, classification and correction system, characterized in that: include: Acquisition module: used to select data sequences and collect influencing factors; Gross error module: used for gross error identification; Discrimination module: used to distinguish volatility errors; Detection module: used to detect mutation points in monitoring data using the cumulative CUMSUM algorithm; Processing module: used to process volatility errors; Processing module: used to fill missing values.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, a method for identifying, classifying and correcting errors in dam safety monitoring data according to any one of claims 1 to 7 is implemented.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a method for identifying, classifying and correcting errors in dam safety monitoring data according to any one of claims 1 to 7 is implemented.