Jump point cleaning method and system for bridge monitoring data

By smoothing the bridge monitoring data and dynamic threshold calculation of the difference sequence, the problem that the 3σ criterion in the prior art cannot effectively handle non-normal distributed bridge monitoring data, and efficient point-hop cleaning and data analysis are achieved.

CN120216870APending Publication Date: 2025-06-27广州珠江黄埔大桥建设有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510263028.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has limitations when processing point jumps in bridge monitoring data. The 3σ criterion is based on the normal distribution assumption and cannot effectively adapt to non-normal distribution bridge monitoring data, and is not reliable enough when the number of measurements is small.

Method used

By smoothing the original data, high-frequency noise interference is suppressed, the influence of random fluctuations on the 3σ criterion is reduced, and the probability of misjudgment is reduced. At the same time, the threshold is dynamically calculated based on the mean and standard deviation of the difference sequence to adapt to the non-stationary characteristics of the bridge monitoring data.

Benefits of technology

It realizes effective point-jump cleaning of bridge monitoring data, improves the balance of detection accuracy and efficiency, and is suitable for real-time or quasi-real-time monitoring systems, avoiding the problems of over-cleaning or under-cleaning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216870A_ABST
    Figure CN120216870A_ABST
Patent Text Reader

Abstract

The invention discloses a jump point cleaning method and system for bridge monitoring data. The method comprises the following steps: obtaining health monitoring to-be-cleaned original data; smoothing the original data to obtain preprocessed data; calculating a mean value and a standard deviation of difference values of the original data and the preprocessed data; and judging and recording the position of the outlier by adopting a 3 sigma criterion, and replacing data at the same position as the position of the outlier in the original data with a null value. According to the method, high-frequency noise interference is suppressed through smooth processing, the influence of random fluctuation in original data on the 3 sigma criterion is reduced, the misjudgment probability is reduced, and meanwhile, the overall trend of effective signals is reserved; the threshold value is dynamically calculated based on the mean value and the standard deviation of the difference value sequence, the method can adapt to the non-stationary characteristic of bridge monitoring data, and the problem of over-cleaning or under-cleaning caused by a fixed threshold value is avoided; the whole process of'smoothing-difference calculation-replacement 'can be completed only through single data traversal, the calculation complexity is low, and the method is suitable for a real-time or quasi-real-time monitoring system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of bridge health monitoring, and specifically relates to a method and system for cleaning jump points in bridge monitoring data. Background Art

[0002] Bridge health monitoring is a technology that continuously and automatically measures, records, and manages set parameters of a bridge through technologies such as sensors, signal acquisition, signal processing, networks, and computers. It can obtain real-time data on the bridge environment, actions, structural responses, and structural changes, and perform data analysis and applications. Since monitoring data may be affected by various noises and interferences, data cleaning is required to eliminate incorrect data to ensure the accuracy and reliability of data analysis. Data jump points are a common problem, which refers to sudden and discontinuous changes in data, which may be caused by sensor failures, data transmission errors, or other abnormal conditions. The existence of jump points will seriously affect the results of data analysis, so a special cleaning method is needed to handle these jump points.

[0003] The traditional method is generally the 3σ criterion, that is, the three-standard-deviation rule, which is a commonly used outlier detection method in statistics. It is based on the characteristics of the normal distribution, that is, in the case of a normal distribution, the data is concentrated near the mean, and about 99.7% of the data falls within three standard deviations of the mean. According to the 3σ criterion, any data point that falls outside three standard deviations of the mean is considered an outlier or a deviant point.

[0004] However, this method has certain limitations:

[0005] Firstly, this method is the most basic and simple method, and its cleaning rate for eliminating jump point data is not high enough. Secondly, the 3σ criterion is based on the assumption that the data follows a normal distribution, but actual bridge monitoring data (such as dynamic strain, temperature gradient, etc.) often shows asymmetric, multi-modal, or heavy-tailed distributions; for example: the dynamic response data of a bridge under vehicle loads may show a skewed distribution, resulting in inaccurate threshold calculation. In addition, the 3σ criterion is applicable to the case where the number of measurements is sufficiently large. When the number of measurements is small, it is not reliable to use this criterion to eliminate gross errors. Summary of the Invention

[0006] In view of the problems existing in the prior art, the present application proposes a method and system for cleaning jump points in bridge monitoring data. By means of smoothing processing, high-frequency noise interference is suppressed, the influence of random fluctuations in the original data on the 3σ criterion is reduced, the probability of misjudgment is decreased, and at the same time, the overall trend of the effective signal is retained. Based on the mean and standard deviation of the difference sequence, the threshold is dynamically calculated, which can adapt to the non-stationary characteristics of bridge monitoring data and avoid over-cleaning or under-cleaning problems caused by a fixed threshold. The entire process of "smoothing - difference calculation - replacement" can be completed with only a single pass through the data, and the computational complexity is low, making it suitable for real-time or quasi-real-time monitoring systems.

[0007] In a first aspect, a method for cleaning jump points in bridge monitoring data includes the following steps.

[0008] S1: Obtain the original data to be cleaned for health monitoring;

[0009] S2: Perform smoothing processing on the original data to obtain preprocessed data;

[0010] S3: Calculate the mean and standard deviation of the difference between the original data and the preprocessed data;

[0011] S4: Use the 3σ criterion to judge and record the positions of the outliers, and replace the data at the same positions as the outliers in the original data with null values, thus completing the data cleaning.

[0012] In an implementation manner of the first aspect, step S1 includes the following contents:

[0013] S101: Obtain the data to be cleaned for health monitoring, denoted as dataset X. Assume that dataset X has n data points, denoted as X = (x1, x2,... x i ,..., x n ), where i ∈ {1, 2,..., n};

[0014] S102: Perform smoothing processing on dataset X to obtain the smoothed dataset S, denoted as S = (s1, s2,... s i ,..., s n ), where i ∈ {1, 2,..., n}.

[0015] In an implementation manner of the first aspect, the smoothing processing in step S2 adopts the moving average method, and the formula is as follows:

[0016]

[0017] where k is the window radius, s i is the smoothed value of the i-th data point, and x i is the original value of the i-th data point.

[0018] In an implementation of the first aspect, the smoothing process in step S2 adopts the single exponential smoothing method, and the formula is as follows:

[0019] s i = αx i +(1 - α)s i-1

[0020] where s i is the smoothed value of the i-th data point, x i is the original value of the i-th data point, and α is the smoothing factor.

[0021] The value range of the smoothing factor α is 0 < α < 1.

[0022] In an implementation of the first aspect, step S3 includes the following:

[0023] S301: Subtract the original data from the preprocessed data:

[0024] d i = s i - x i

[0025] where s i is the smoothed value of the i-th data point, x i is the original value of the i-th data point;

[0026] S302: Denote the data set composed of each difference as D = (d1, d2,..., d i ,..., d n ), and calculate the mean and standard deviation of the data set D:

[0027]

[0028] In an implementation of the first aspect, step S4 includes the following:

[0029] S401: Record the position information P:

[0030] P = (p1, p2,..., p i ,..., p n )

[0031] where p i represents the i-th position information, and the value is as follows:

[0032]

[0033] S402: The data after cleaning is:

[0034]

[0035] where nan represents a null value, and y i represents the data value after the i-th cleaning;

[0036] S403: Denote the cleaned dataset as Y = (y1, y2,..., y i ,..., y n ).

[0037] In a second aspect, a system for performing the jump point cleaning method includes:

[0038] An acquisition unit for obtaining the original data to be cleaned for health monitoring;

[0039] A preprocessing unit for smoothing the original data to obtain preprocessed data;

[0040] A mean and standard deviation unit for calculating the difference between the original data and the preprocessed data, and then calculating the corresponding mean and standard deviation;

[0041] A data cleaning unit that uses the 3σ criterion to judge and record the positions of outliers, and replaces the data at the same positions as the outliers in the original data with null values, thus completing the data cleaning.

[0042] The traditional jump point cleaning algorithm based on the 3σ criterion mainly calculates the mean u and standard deviation v of the data column, and determines that the numerical values not within the interval [u - 3v, u + 3v] are all regarded as outliers and will be cleaned as jump point values. Based on the traditional 3σ criterion and data smoothing processing, this application proposes an improved data cleaning algorithm. First, the data column needs to be smoothed, and then the difference between the smoothed data and the original data is calculated. Based on the difference, the 3σ criterion is used to clean the outliers; specifically, this application has the following beneficial effects:

[0043] Balance between detection accuracy and efficiency: By smoothing processing (such as moving average or median filtering) to suppress high-frequency noise interference, reduce the influence of random fluctuations in the original data on the 3σ criterion, reduce the probability of misjudgment, and at the same time retain the overall trend of the effective signal. The combination of the two avoids the sensitivity defect of directly applying the 3σ criterion to non-Gaussian distributed data, and simplifies the complexity of outlier identification through preprocessing.

[0044] Dynamic adaptability and robustness: Dynamically calculate the threshold based on the mean and standard deviation of the difference sequence, which can adapt to the non-stationary characteristics of bridge monitoring data (such as baseline drift caused by temperature and load changes), and avoid the problems of over-cleaning or under-cleaning caused by fixed thresholds.

[0045] Process lightweight and engineering practicability: The entire process of "smoothing - difference calculation - replacement" can be completed with only a single data traversal, with low computational complexity, and is applicable to real-time or quasi-real-time monitoring systems. In addition, the null value replacement strategy preserves the integrity of the original data timestamps, facilitating subsequent missing value imputation and in-depth analysis. Description of the Drawings

[0046] Figure 1 This is a flowchart of the data cleaning method for an embodiment of the present application. Detailed Embodiments

[0047] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings.

[0048] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0049] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0050] In the first aspect, the present application provides a method for cleaning jump points in bridge monitoring data, including the following steps:

[0051] S1: Obtain the data to be cleaned for health monitoring, that is, the original data;

[0052] S2: Smooth the original data to obtain preprocessed data;

[0053] S3: Calculate the mean and standard deviation of the difference between the original data and the preprocessed data;

[0054] S4: Use the 3σ criterion to record the positions of the outliers, and replace the data at the same positions as the outliers in the original data with null values, that is, complete the data cleaning.

[0055] Specifically, step S1 includes the following content:

[0056] S101: Obtain the data to be cleaned for health monitoring, denoted as dataset X. Assume that dataset X has n data points, denoted as X = (x1, x2,... x i ,..., x n ), where i ∈ {1, 2,..., n};

[0057] S102: Smooth the dataset X to obtain the smoothed dataset S, denoted as S = (s1, s2,... s i ,..., s n ), where i ∈ {1, 2,..., n}.

[0058] Optionally, in step S2, the algorithms used for the smoothing process include but are not limited to the moving average algorithm, median filtering algorithm, exponential smoothing algorithm, Savitzky - Golay, and LOESS. Conventionally, for the application scenario of this application, the moving average method or the first - order exponential smoothing method can be used.

[0059] The formula for the moving average method is as follows:

[0060]

[0061] where k is the window radius, s i is the smoothed value of the i - th data point, and x i is the original value of the i - th data point.

[0062] The formula for the first - order exponential smoothing method is as follows:

[0063] s i = αx i + (1 - α)s i-1

[0064] where s i is the smoothed value of the i - th data point, x i is the original value of the i - th data point, α is the smoothing factor, and the value range of α is 0 < α < 1.

[0065] Furthermore, step S3 includes the following content:

[0066] S301: Subtract the original data from the pre - processed data:

[0067] d i = s i - x i

[0068] where s i is the smoothed value of the i - th data point, and x i is the original value of the i - th data point;

[0069] S302: Denote the dataset composed of each difference as D = (d1, d2,..., d i ,..., d n ), and calculate the mean and standard deviation of the dataset D:

[0070]

[0071]

[0072] Further, step S4 includes the following:

[0073] S401: Record the position information P:

[0074] P = (p1, p2,..., p i ,..., p n )

[0075] where pi i represents the i-th position information, and the value is as follows:

[0076]

[0077] S402: The data after cleaning is:

[0078]

[0079] where nan represents a null value, and yi i represents the i-th data value after cleaning;

[0080] S403: Then the cleaned data set is Y = (y1, y2,..., y i ,..., y n ).

[0081] In a second aspect, the present application provides a system for executing the above jump point cleaning method, including:

[0082] An acquisition unit, configured to obtain the data to be cleaned for health monitoring, that is, the raw data;

[0083] A preprocessing unit, configured to perform smoothing processing on the raw data to obtain preprocessed data;

[0084] A mean and standard deviation unit, configured to calculate the difference between the raw data and the preprocessed data, and further calculate the corresponding mean and standard deviation;

[0085] A data cleaning unit, which uses the 3σ criterion to judge and record the positions of outliers, and replaces the data at the same positions as the outliers in the raw data with null values, that is, completes the data cleaning.

[0086] It should be understood that the division of each processing unit in the above system is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. In addition, the processing units in the system can be implemented in the form of a processor calling software. For example, the system includes a processor, the processor is connected to a memory, instructions are stored in the memory, and the processor calls the instructions stored in the memory to implement any of the above methods or the functions of each processing unit in the system. The processor is a general-purpose processor, such as a central processing unit or a microprocessor, and the memory is a memory inside or outside the system.

[0087] The above embodiments are only used to illustrate the technical concept and features of the present application. The purpose is to enable those skilled in the art to understand the content of the present application and implement it accordingly, and it cannot be used to limit the protection scope of the present application. For those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for cleaning jump points of bridge monitoring data, characterized in that: The following steps are included. S1: Collect the raw data to be cleaned from health monitoring; S2: Smoothing the original data to obtain preprocessed data; S3: Calculate the mean and standard deviation of the difference between the original data and the preprocessed data; S4: Use the 3σ criterion to determine and record the position of the outlier, and replace the data at the same position as the outlier in the original data with a null value, thus completing the data cleaning.

2. The jump point cleaning method according to claim 1, characterized in that: Step S1 includes the following contents: S101: Obtain the health monitoring data to be cleaned, denoted as data set X. Assume that data set X has n data points, denoted as X=(x1, x2, ... x i ,...,x n ), where i∈{1,2,...,n}; S102: Smoothing the data set X to obtain a smoothed data set S, where S = (s1, s2, ...s i ,...,s n ), where i∈{1,2,...,n}.

3. The jump point cleaning method according to claim 2, characterized in that: In step S2, the smoothing process adopts the moving average method, and the formula is as follows: Where k is the window radius, s i is the smoothed value of the ith data point, x i is the original value of the ith data point.

4. The jump point cleaning method according to claim 2, characterized in that: In step S2, the smoothing process adopts a first exponential smoothing method, and the formula is as follows: s i =αx i +(1-α)s i-1 Among them, s i is the smoothed value of the ith data point, x i is the original value of the ith data point and α is the smoothing factor.

5. The jump point cleaning method according to claim 4, characterized in that: The value range of the smoothing factor α is 0<α<1.

6. The jump point cleaning method according to claim 1, characterized in that: Step S3 includes the following contents: S301: Subtract the original data from the pre-processed data: d i =s i -x i Among them, s i is the smoothed value of the ith data point, x i is the original value of the ith data point; S302: Record the data set composed of each difference as D=(d1, d2, ..., d i ,...,d n ), calculate the mean and standard deviation of data set D:

7. The jump point cleaning method according to claim 1, characterized in that: Step S4 includes the following contents: S401: Recording location information P: P=(p1,p2,...,p i ,...,p n ) Among them, p i Represents the i-th position information, and its values ​​are as follows: S402: The cleaned data is: Among them, nan represents a null value, y i Represents the data value after the i-th cleaning; S403: The cleaned data set is recorded as Y=(y1, y2, ..., y i ,...,y n ).

8. A system for executing the jump point cleaning method according to any one of claims 1 to 7, characterized in that: include: A collection unit, used to collect raw data to be cleaned from health monitoring; A preprocessing unit, used for smoothing the original data to obtain preprocessed data; The mean and standard deviation unit is used to calculate the difference between the original data and the preprocessed data, and then calculate the corresponding mean and standard deviation; The data cleaning unit uses the 3σ criterion to determine and record the position of the outlier, and replaces the data at the same position as the outlier in the original data with a null value, thus completing the data cleaning.