Burr and step detection and data anomaly identification method and bridge health monitoring method
By using glitch abnormality detection method in data abnormality recognition, the trend distribution is used to calculate the reference value and dynamic threshold, and glitch abnormalities in massive data are identified, the problem of low accuracy in the existing technology is solved and detection accuracy and flexibility are improved.
Patent Information
- Application Number
- CN202510296812.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has low accuracy in data abnormal identification, and it is prone to misidentification or misidentification, especially when massive data is processed.
A glitch abnormality detection method is proposed. By dividing the time series data set into multiple data intervals with equal numerical widths, the fluctuations at each time point are calculated, and the continuous time points with fluctuations greater than the glitch threshold are identified as glitches. At the same time, the trend distribution is used to calculate the benchmark value, remove the influence of sudden factors, and improve the detection accuracy.
The glitch detection accuracy is improved, the false alarm rate is reduced, and the method flexibility is enhanced through dynamic threshold and trend distribution analysis.
Smart Images

Figure CN120217235A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of data anomaly recognition and bridge safety monitoring, and in particular to a method for detecting burrs and steps, recognizing data anomalies, and a method for bridge health monitoring. Background Art
[0002] In recent years, with the rapid expansion of the amount of data collected by monitoring systems, and at the same time due to reasons such as environmental noise, abnormal sensor devices, abnormal network transmission devices, abnormal acquisition and parsing devices, and software system anomalies, the data sets to be analyzed often contain a lot of noisy, incomplete, or even inconsistent data; before establishing a time series model, it is necessary to preprocess the dynamic monitoring data, check the basic statistical characteristics of these sample monitoring data, and classify the abnormal data that does not conform to the actual rules to ensure the confidence and reliability of the established time series model and meet certain accuracy requirements.
[0003] According to the current common data anomaly manifestations, they are divided into four categories: missing, constant value, burr, and step anomaly. The missing anomaly is manifested as the difference between the theoretical data volume and the actual data volume. The original data curve is shown in Figure 3 ; the constant value anomaly is manifested as the data within a continuous time period remaining a constant value and no longer changing. The original data curve is shown in Figure 4 ; the burr anomaly is manifested as a sudden rise or fall in the original data curve at certain moments, and then returns to the normal change range. The original data curve is shown in Figure 5 . The step anomaly is manifested as the original data curve changing from one level range to another level range at a certain moment and maintaining a certain time period. The original data curve is shown in Figure 6 .
[0004] The current data processing methods are basically based on the mean of the data sequence or the difference between adjacent data for anomaly judgment. Facing a large amount of data, the judgment accuracy cannot be guaranteed; it is very easy to have the situation of misrecognition or missed recognition. Summary of the Invention
[0005] In order to overcome the defect of low accuracy in data anomaly recognition in the above-mentioned prior art, the present invention proposes a method for detecting burr anomalies, which considers the trend distribution, focuses on long-term features, and improves the burr detection accuracy.
[0006] A method for detecting burr anomalies proposed by the present invention first divides the time series data set into multiple data intervals with equal numerical widths; and takes the mean value of the data interval containing the largest number of numerical values as the reference value;
[0007] Calculate the difference between the numerical value at each time point of the time series data and the reference value as the fluctuation at that time point, and statistically identify the continuous time points with fluctuations greater than the burr threshold as burrs.
[0008] Preferably, fluctuations greater than the burr threshold are denoted as bumps; when the number of time points included in consecutive bumps is greater than the set point threshold, the consecutive bumps are identified as burrs.
[0009] Preferably, the number of data intervals is a set value and is proportional to the number of numerical values of the time series data.
[0010] Preferably, the number of data intervals is set as: k = 1 + 3.322lgN, where k is the number of data intervals and N is the number of numerical values of the time series data.
[0011] A step anomaly detection method proposed by the present invention first obtains a time series data set and performs window division, and calculates the reference value of each window data; traverses adjacent two window data, if the absolute value of the difference between the reference values of adjacent two window data is greater than the set step threshold, it is determined that there is a step anomaly in the adjacent two window data;
[0012] The calculation method of the reference value of the window data is: divide the window data into multiple data intervals with equal numerical widths; let the mean value of the data interval containing the largest number of numerical values be the reference value of the window data.
[0013] Preferably, if the absolute value of the difference between the reference values of adjacent two window data is greater than the set step threshold, it is determined that there is a step anomaly in the adjacent two window data, and the start time of the step is located at the center point of the front window data, and the end time of the step is located at the center point of the back window data.
[0014] A data anomaly recognition method proposed by the present invention first samples the input data to obtain a time series data set; performs sampling frequency detection on the time series data set, and when the sampling frequency is normal, then performs missing anomaly detection, constant value anomaly detection, burr anomaly detection and step anomaly detection respectively;
[0015] The burr anomaly detection method is: first divide the time series data set into multiple data intervals with equal numerical widths; let the mean value of the data interval containing the largest number of numerical values be the reference value;
[0016] Calculate the difference between the numerical value at each time point of the time series data and the reference value as the fluctuation at that time point, and statistically identify the consecutive time points with fluctuations greater than the burr threshold as burrs.
[0017] Preferably, the step anomaly detection method is: first obtain a time series data set and perform window division, calculate the reference value of each window data; traverse adjacent two window data, if the absolute value of the difference between the reference values of adjacent two window data is greater than the set step threshold, it is determined that there is a step anomaly in the adjacent two window data.
[0018] A bridge health monitoring method proposed by the present invention sets multiple strain monitoring points on the bridge to collect bridge strain data, and uses the data anomaly recognition method to detect anomalies in the bridge strain data; for the bridge strain data determined to be abnormal, mark the strain monitoring points and display the data anomalies.
[0019] A data anomaly recognition system proposed by the present invention includes a memory and a processor. A computer program is stored in the memory, and the processor is connected to the memory. The processor is used to execute the computer program to implement the data anomaly recognition method.
[0020] The advantages of the present invention are as follows:
[0021] (1) The spike anomaly detection method proposed by the present invention calculates the reference value based on the trend distribution, and then uses the difference between the sampled value and the reference value to judge the spike. Considering the trend distribution, the influence of sudden factors is well removed, and only the influence of long-term factors is considered for data analysis, achieving a better spike detection effect and improving the spike detection accuracy.
[0022] (2) The present invention uses continuous bump determination to avoid single-point noise interference and reduce the false alarm rate. Moreover, the present invention uses a dynamic threshold, and the number of data intervals is proportional to the data volume, adapting to different-scale data sets, avoiding human division deviation, and improving the flexibility of the method.
[0023] (3) The step anomaly detection method proposed by the present invention first calculates the reference value based on the trend distribution on the window data to ensure that the reference value reflects the long-term trend of the window data and avoids the influence of short-term fluctuations, thereby improving the accuracy of step analysis. In the present invention, the step anomaly is determined by the difference between the reference values of adjacent windows, and the start and end times are located by the center points of the front and rear windows, improving the time positioning accuracy of sudden events.
[0024] (4) The data anomaly recognition method proposed by the present invention integrates four types of anomaly detections: missing, constant value, spike, and step, comprehensively covering common data problems in bridge monitoring, and realizing multi-dimensional anomaly coverage. The present invention pre-detects the sampling frequency and preprocesses the data to ensure the quality of the input data and reduce the risk of subsequent misjudgment.
[0025] (5) The present invention provides an intelligent method and system for identifying abnormal data in bridge monitoring, which can accurately and efficiently identify outliers in monitoring data through an automated analysis model, improving the accuracy and real-time performance of bridge monitoring. By analyzing common abnormalities such as data loss, delay, interference, and jump in the field of bridge monitoring, an abnormal identification model based on frequency distribution and statistical analysis methods are adopted to complete functions such as data preprocessing, feature extraction, abnormal identification, and result discrimination. The invention can respond to various abnormal situations in bridge monitoring data in real time, has strong adaptability, and can evaluate the stability of various devices through the identification results of abnormal data.
[0026] (6) The present invention combines innovative reference value calculation, dynamic threshold strategy, and multi-abnormality collaborative detection mechanism to achieve high-precision and high-efficiency abnormal identification in the bridge monitoring scenario, taking into account real-time performance and reliability, and providing intelligent technical support for infrastructure health management. Brief Description of the Drawings
[0027] Figure 1 It is a flowchart of the data abnormality identification method;
[0028] Figure 2 It is a flowchart of the burr abnormality detection method;
[0029] Figure 3 It is a flowchart of the step abnormality detection method;
[0030] Figure 4 It is a data analysis diagram of the embodiment;
[0031] Figure 5 It is a display of the effect of detecting burrs by the moving average method;
[0032] Figure 6 It is a display of the effect of detecting burrs by the method of the present invention;
[0033] Figure 7 It is a display of the abnormal data identification of the Sx14 sensor detection data in the embodiment. Detailed Embodiment
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0035] Refer to Figure 1, the data anomaly recognition method proposed in this embodiment first samples the input data to obtain a time series data set X; performs sampling frequency detection on the time series data set X, and when the sampling frequency is normal, performs missing anomaly detection, constant value anomaly detection, glitch anomaly detection, and step anomaly detection respectively; and displays the distribution of abnormal data through an abnormal data Gantt chart.
[0036] Specifically, the method for performing sampling frequency detection on the time series data set X is as follows: calculate the theoretical sampling time of each sampling point by combining the time length of the input data and the set sampling frequency; compare the actual sampling time of each sampling value in the time series data set X with the corresponding theoretical sampling time. If the actual sampling time of all sampling values is consistent with the corresponding theoretical sampling time, it is determined that the sampling frequency of the time series data set X is normal.
[0037] In specific implementation, the theoretical sampling times can be sorted according to time sequence. Combining the sampling sequence of the sampling values, compare the actual sampling time of the sampling values with the theoretical sampling times in the corresponding order. If they are consistent, then compare the next sampling value; once the situation where the actual sampling time is inconsistent with the theoretical sampling time appears, end the comparison and determine that the sampling frequency of the time series data set X is abnormal.
[0038] Alternatively, a set can be established to store the theoretical sampling times, and then compare the actual sampling times of each sampling value in the time series data set X with the theoretical sampling times in the set. If a consistent theoretical sampling time is found, delete the theoretical sampling time from the set and search for the next sampling value; if no consistent theoretical sampling time is found, end the comparison and determine that the sampling frequency is abnormal.
[0039] Alternatively, calculate the interval between adjacent sampling times and determine whether all intervals conform to the sampling frequency. If so, it means the sampling frequency is normal; otherwise, it means the sampling frequency is abnormal. The formula is as follows:
[0040] Sampling is performed over a specified time period to form a time information sequence T = {t1, t2,....., t n ,.., t N}, where t1, t2, t n ,, t N respectively represent the sampling times of the 1st, 2nd, nth, and Nth sampling values, and N is the total number of values in the time series data set X;
[0041] Let the i-th sampling interval DT i = t i+1 - t i
[0042] t i+1 represents the sampling time of the (i + 1)-th sampling value, and t iRepresents the sampling time of the i-th sampled value, where 1 ≤ i ≤ N;
[0043] If DT i = 1 / F s , it indicates that there is an abnormal sampling interval, that is, the sampling frequency is abnormal; otherwise, it means the sampling interval is normal; if all sampling intervals are normal, the data passes the sampling frequency detection.
[0044] The method for detecting missing anomalies is as follows: First, calculate the theoretical number of samples by combining the time length of the input data and the set sampling frequency, and compare the actual number of samples with the theoretical number of samples. If the actual number of samples is greater than or equal to the theoretical number of samples, it is determined that the data has no missing values; otherwise, it is determined that the data is missing.
[0045] The method for detecting constant value anomalies is as follows: Count the number of cases where the difference between adjacent sampled values in the time series dataset X in consecutive time is 0. If there are more than N1 consecutive cases where the difference between adjacent sampled values is 0, it is determined that a constant value appears in this consecutive time; N1 is the set threshold.
[0046] The method for detecting spike anomalies proposed in this embodiment divides the time series data into multiple data intervals with equal numerical widths. The numerical width of the data interval is the difference between the maximum value and the minimum value on the data interval;
[0047] Let the mean value of the data interval containing the largest number of numerical values be used as the reference value;
[0048] Calculate the difference between the numerical value at each time point of the time series data and the reference value as the fluctuation at that time point, and count the consecutive time points where the fluctuation is greater than the spike threshold as spikes.
[0049] As Figure 2 shown, this method specifically includes the following steps:
[0050] S1. Obtain the time series dataset X and divide it into consecutive data intervals with equal widths;
[0051] In specific implementation, the width of the data interval can be set to w;
[0052]
[0053] X = {x1, x2, …, x n , …, x N}
[0054] where maxX is the maximum value of the time series dataset X, minX is the minimum value of the time series dataset X; x n is the numerical value at the n-th time point in the time series dataset X, where 1 ≤ n ≤ N, and N is the number of numerical values in the time series dataset X; that is, x1, x2, x NThey are the values at the 1st, 2nd, and Nth time points in the time series dataset X respectively;
[0055] k is the number of data intervals, and specifically, k can take the value k = 1 + 3.322 log 10 (N).
[0056] Then the value range of the data interval i is [minX + (i - 1)w, minX + iw], where 1 ≤ i ≤ k.
[0057] S2. Calculate the frequency of each data interval. The frequency is the number of values contained in the data interval.
[0058] S3. Calculate the average value of the data interval with the largest frequency as the reference value D x ;
[0059] S4. Calculate the fluctuation curve of the time series dataset X. Let the fluctuation of the value x n be denoted as dif n ;
[0060] dif n = |x n - D x |
[0061] S5. Judge whether dif n is greater than the spike threshold;
[0062] If it is less than the threshold, it is defined as normal;
[0063] If it is greater than the threshold, it is defined as a spike point and step S6 is executed;
[0064] S6. Statistically identify the spike points that are identified as spike points in continuous time and whose continuous time length is greater than or equal to the set point threshold as glitch anomalies; Take the occurrence time of the first spike point in continuous time as the glitch occurrence time, and the time span of continuous spike points as the duration of the glitch.
[0065] Referring to Figure 3 , the step anomaly detection method proposed in this embodiment includes the following steps:
[0066] St1. Obtain the time series dataset X and perform window partitioning. The window step size is the set value a, where a < N;
[0067] In this way, let M = floor(N / a), where floor represents rounding down, and N is the number of values in X; If N is an integer multiple of a, the number of windows is M; Otherwise, the number of windows is M + 1; The mth window is denoted as X(m), where 1 ≤ m ≤ M;
[0068] X(m) = {x (m-1)a+1 , x (m-1)a+2 , …, xma}
[0069] Among them, x (m-1)a+1 , x (m-1)a+2 , x ma are the (m - 1)a + 1, (m - 1)a + 2, and ma-th values in X respectively;
[0070] When Ma < N, the last window data is X(M + 1) = {x Ma+1 , x Ma+2 , …, x N}], x Ma+1 , x Ma+2 , x N are the Ma + 1, Ma + 2, and N-th values in X respectively.
[0071] St2. Calculate the reference values of each window data;
[0072] The calculation method of the reference value of the window data is as follows:
[0073] First, divide the window data into equally wide data intervals. The width of the data interval is the difference between the maximum value and the minimum value. The width of the data interval can be set directly or calculated according to the following method;
[0074] That is:
[0075] k(h) = 1 + 3.322log 10 (N h ).
[0076] Among them, N h is the number of values of the h-th window data, k(h) is the number of data intervals divided by the h-th window data, and w(h) is the width of the data interval divided by the h-th window data.
[0077] St3. Determine whether the absolute value of the difference between the reference values of two adjacent window data is greater than the set step threshold;
[0078] No, it is determined that there is no step in the current window;
[0079] Yes, it is determined that there is a step, and the starting time of the step is at the center point of the previous window data, and the duration of the step is from the center point of the previous window data to the center point of the subsequent window data.
[0080] That is, when |DX(h) - DX(h - 1)| > c, it indicates that there is a step anomaly, and the step extends from t(h - 1) to t(h);
[0081] Among them, DX(h) is the reference value of the h-th window data; t(h) is the center point of the h-th window data, that is, if Nh is an even number, t(h) is the midpoint of the sampling time of the two middle sampling values on the window. If N h is an odd number, t(h) is the sampling time of the middle sampling value on the window; X(h - 1) is the reference value of the h - 1th window data; t(h - 1) is the center point of the h - 1th window data.
[0082] Referring to Figure 4 , the above burr anomaly detection method is verified by combining the strain data collected by sensors from 0:00 to 1:00 on the 4th of a certain month for a certain bridge.
[0083] In this embodiment, the data length is 1 hour and the sampling interval is 10 minutes.
[0084] In this embodiment, burrs are found by two methods respectively:
[0085] The first method, that is, the burr anomaly detection method provided by the present invention, the difference between each calculated value and the reference value is shown as the probability distribution trend value in the figure;
[0086] The second method, for each value, the mean value within 10 minutes centered on itself is taken as the reference value, and the difference between each value and the reference value is shown as the moving average trend value in the figure.
[0087] The original data curve of a certain sensor is as Figure 4 shown. Referring to Figures 4 to 6 , it can be seen that the probability distribution trend value obviously shows a smoother change trend, and its change is not affected by the protruding peaks and valleys; while the moving average trend value is affected by the fluctuations and the change is not smooth, showing a stepped shape locally.
[0088] Referring to Figure 5 , when using the moving average value trend method to judge burrs, only 2 burrs can be identified. This is because the moving average method is greatly affected by the curve fluctuations, which affects the burr threshold judgment.
[0089] Referring to Figure 6 , when using the burr anomaly detection method proposed by the present invention, 5 burrs can be identified, 3 more burrs than the moving average method.
[0090] In this embodiment, the analysis of more sensor data is shown in the following table.
[0091] Table 1: Bridge Strain Statistics
[0092]
[0093] In this embodiment, the sensor data at Sx14 is also extracted for data anomaly identification, and the results are shown as Figure 7 shown.
[0094] As can be seen from Table 1, in any case, the burr anomaly detection method proposed by the present invention (abbreviated as the probability distribution method) has achieved a significant improvement in accuracy. This is because the trend value should not consider the influence of sudden factors such as vehicle load, but only consider the influence of long-term factors such as temperature and structural creep. Therefore, the change of the trend value should be gentle. The present invention adopts the trend distribution method to effectively eliminate the influence of sudden factors and only consider the influence of long-term factors, so good results can be obtained.
[0095] Certainly, for those skilled in the art, the present invention is not limited to the details of the above exemplary embodiments, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0096] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
[0097] The technologies, shapes, and structures not described in detail in the present invention are all well-known technologies.
Claims
1. A burr anomaly detection method, characterized in that: First, divide the time series data set into multiple data intervals with equal value widths; let the mean of the data interval containing the largest number of values be the benchmark value; The difference between the value at each time point of the time series data and the reference value is calculated as the fluctuation at that time point, and the continuous time points whose statistical fluctuation is greater than the glitch threshold are identified as glitches.
2. The burr abnormality detection method according to claim 1, characterized in that: The fluctuations greater than the glitch threshold are recorded as burrs; when the number of time points contained in the continuous burrs is greater than the set point number threshold, the continuous burrs are identified as burrs.
3. The burr abnormality detection method according to claim 1 or 2, characterized in that: The number of data intervals is a set value and is proportional to the number of values in the time series data.
4. The burr abnormality detection method according to claim 3, characterized in that: The number of data intervals is set to: k=1+3.322lgN, where k is the number of data intervals and N is the number of values of the time series data.
5. A step anomaly detection method, characterized in that: First, the time series data set is obtained and divided into windows, and the benchmark value of each window data is calculated; the data of two adjacent windows are traversed, and if the absolute value of the difference between the benchmark values of the two adjacent window data is greater than the set step threshold, it is determined that the two adjacent window data have step anomalies; The calculation method of the reference value of the window data is as follows: divide the window data into multiple data intervals with equal value widths; and take the mean value of the data interval containing the largest number of values as the reference value of the window data.
6. The step abnormality detection method according to claim 5, characterized in that: If the absolute value of the difference between the baseline values of two adjacent window data is greater than the set step threshold, it is determined that there is a step anomaly in the two adjacent window data, and the starting time of the step is located at the center point of the previous window data, and the ending time of the step is located at the center point of the subsequent window data.
7. A method for identifying data anomalies, characterized in that: First, the input data is sampled to obtain a time series data set; the sampling frequency of the time series data set is detected, and when the sampling frequency is normal, missing anomaly detection, constant value anomaly detection, burr anomaly detection and step anomaly detection are performed respectively; The glitch anomaly detection method is as follows: first, the time series data set is divided into multiple data intervals with equal value widths; the mean value of the data interval containing the largest number of values is taken as the reference value; The difference between the value at each time point of the time series data and the reference value is calculated as the fluctuation at that time point, and the continuous time points whose statistical fluctuation is greater than the glitch threshold are identified as glitches.
8. The data anomaly identification method according to claim 7, characterized in that: The step anomaly detection method is as follows: first, obtain the time series data set and divide it into windows, and calculate the baseline value of each window data; traverse the data of two adjacent windows, and if the absolute value of the difference between the baseline values of the two adjacent window data is greater than the set step threshold, it is judged that the two adjacent window data have a step anomaly.
9. A bridge health monitoring method, characterized in that: A plurality of strain monitoring points are arranged on the bridge to collect bridge strain data, and the data anomaly identification method as described in claim 7 or 8 is used to perform anomaly detection on the bridge strain data; the strain monitoring points are marked for the bridge strain data judged to be abnormal and the data anomaly is displayed.
10. A data anomaly identification system, characterized in that: It includes a memory and a processor, the memory stores a computer program, the processor is connected to the memory, and the processor is used to execute the computer program to implement the data anomaly identification method as described in claim 7 or 8.
Citation Information
Cited By
Alarm suppression method and device, computer equipment, readable storage medium and program product
CN121333880A