A data distortion recognition method based on association density clustering and its application

The gray B-type correlation density clustering method is used to identify data distortion caused by hardware failures, which solves the problem of frequent sensor failures, improves the accuracy and reliability of the monitoring system, and is suitable for health monitoring of ancient wooden structures.

CN116662847BActive Publication Date: 2025-07-04BEIJING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310279317.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-21
Publication Date
2025-07-04
Estimated Expiration
2043-03-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify data distortion caused by hardware failures, which affects the accuracy and reliability of structural safety assessments. Especially in the health monitoring of long-lived structures such as ancient building wood structures, sensor failures occur frequently.

Method used

A method based on gray B-type correlation density clustering is adopted to measure the correlation between the overall displacement difference, the overall first-order slope difference and the overall second-order slope difference of the strain measurement points, a classification model is constructed to identify abnormalities in real-time monitoring data.

Benefits of technology

It improves the reliability and robustness of data distortion recognition, can promptly handle data distortion caused by hardware failures, ensures the accuracy and reliability of the monitoring system, and reduces structural abnormalities and misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116662847B_ABST
    Figure CN116662847B_ABST
Patent Text Reader

Abstract

The present invention discloses a data distortion recognition method based on correlation density clustering, including: for the historical time series of n strain measurement points of a strain sensor, wherein the correlation degree between each measurement point sequence is measured from the perspectives of the overall displacement difference representing the similarity of the n strain measurement points based on grey type B correlation, the overall first-order slope difference representing the similarity of development trends, and the overall second-order slope difference representing the similarity of development speeds, and a classification model is constructed, and the real-time monitoring data is recognized based on the classification model. The present invention can complete the recognition of data distortion caused by hardware failures, utilize the correlation of multi-source data to realize the fault diagnosis of multi-measurement-point devices, and improve the reliability and robustness of clustering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a data distortion recognition method based on association density clustering and its application. Background Art

[0002] When evaluating the structural safety state based on monitoring data, real and reliable measured data is the prerequisite for exploring the actual state of the structure. Therefore, the data obtained by monitoring sensors cannot be simply assumed to be accurate. For most health monitoring systems, the service life of sensors and other devices is only a few decades or even more than a decade. Compared with the lifespan of hundreds of years of ancient wooden structures, failures caused by various reasons such as external environment, aging, and electromagnetic interference during their service life are not uncommon. Therefore, realizing real-time identification of data distortion caused by equipment failures can not only timely handle the failures, ensure the reliability of the monitoring system, but also provide reliable and accurate data for structural safety assessment, and reduce the misjudgment of structural anomalies in the monitoring system.

[0003] Data distortion caused by system hardware failures is different from the anomalies caused by structural damage. Therefore, ensuring correct distinction is the primary prerequisite. Data distortion caused by equipment anomalies is often single or a few measurement points. Under the influence of environmental temperature, there is a significant correlation between each strain measurement point in structural strain. Considering that the data information of a single sensor is single and has strong one-sidedness, the structural strain has relatively stable time series characteristics is utilized at the same time. Summary of the Invention

[0004] The purpose of the present invention is to provide a data distortion recognition method based on association density clustering to solve the technical problems in the prior art.

[0005] To solve the above technical problems, the present invention specifically provides the following technical solutions:

[0006] In the first aspect of the present invention, a data distortion recognition method based on association density clustering is provided, including:

[0007] For the historical time series X = {x1(t), x2(t),..., x n (t)} of n strain measurement points of strain sensors, where 0 ≤ t ≤ m, the correlation degree between each measurement point sequence is measured from the perspectives of the overall displacement difference representing the similarity of n strain measurement points, the overall first-order slope difference representing the similarity of development trends, and the overall second-order slope difference representing the similarity of development speeds based on grey type B association, and a classification model is constructed, and real-time monitoring data is recognized based on the classification model.

[0008] Preferably, it further includes the standardization processing of the historical time series:

[0009] Divide the historical time series of n measurement points over a certain period into stepwise clustering windows. The length w determines the time period for anomaly localization. The shorter the window, the stronger the immediacy of the determination.

[0010] Use z-score to standardize each sequence in the window to eliminate the dimension or balance the measurement scales of each strain sensor.

[0011] Suppose the time series of a certain window after standardization is shown in Equation (6).

[0012]

[0013] where n is the measurement point number and m is the time number.

[0014] Preferably, it also includes smoothing and denoising the historical time series after the standardization process.

[0015] Select the Savitzky-Golay method for smoothing and denoising, and combine convolution and polynomial regression to achieve smoothing filtering. Different smoothing effects are achieved by adjusting the sliding window and the fitting order, which is expressed as:

[0016]

[0017] where 2l + 1 is the sliding window length, x k represents the center of the sliding window, and h i is the smoothing coefficient, which is obtained by fitting a polynomial using the least squares method.

[0018] Among them, the smaller the polynomial fitting order and the longer the window length, the more significant the smoothing effect.

[0019] Preferably, constructing the classification model includes:

[0020] Perform differential processing on the overall displacement difference, overall first-order slope difference, and overall second-order slope difference of the window sequence after smoothing and denoising by columns respectively, and merge them with matrix X to form a new matrix. At the same time, add the weight coefficients to the new matrix to obtain the clustering feature matrix Y.

[0021] Perform clustering analysis. In the clustering calculation process, determine the Manhattan distance as the calculation method for the distance between sequences.

[0022] Determine the parameters (∈, MinPts) according to the actual engineering needs to obtain the final clustering result and construct the classification model.

[0023] Preferably, the new matrix includes:

[0024] Overall displacement difference matrix:

[0025]

[0026] First-order difference matrix:

[0027] Second-order difference matrix:

[0028] Clustering feature moment

[0029] where dif ij = x i(j+1) - x ij , w1, w2, w3 are weights. If more dependence on proximity features in clustering analysis, then increase w1 and decrease w2 and w3; if more dependence on similarity features, then increase w2 and w3 and decrease w1.

[0030] Preferably, the Manhattan distance is as shown in Equation (11):

[0031]

[0032] Preferably, identifying real-time monitoring data based on the classification model includes:

[0033] For the real-time collected monitoring data, perform clustering analysis step by step with the same window length w;

[0034] When the clustering results of all measurement points are consistent with the clustering model, it is determined that the data for this period is normal;

[0035] When the clustering results are inconsistent and outliers appear, it is regarded as distorted;

[0036] When the clustering results are inconsistent and there are no outliers, further perform data distortion identification.

[0037] On the other hand, the present invention provides an application of a data distortion identification method, which is applied to the health monitoring of wood structures.

[0038] Preferably, the measurement points include strain measurement points of homogeneous components and strain measurement points of heterogeneous components.

[0039] The present invention has the following beneficial effects compared with the prior art:

[0040] The present invention completes the identification of data distortion caused by hardware failures based on the method of grey type-B correlation density clustering. By using the correlation of multi-source data, the fault diagnosis of multi-measurement-point devices can be realized. The improved grey type-B correlation degree realizes the determination and calculation of characteristic attributes, improves the reliability and robustness of clustering, and better identifies distorted data. Description of the Drawings

[0041] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only exemplary. For those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.

[0042] Figure 1 Layout diagram of the strain sensor of the present invention;

[0043] Figure 2 Physical diagram of the layout of the strain and temperature / humidity sensors of the present invention;

[0044] Figure 3 Schematic diagram of clustering mode 1 of strain measurement points of similar components of the present invention;

[0045] Figure 4 Schematic diagram of clustering mode 2 of strain measurement points of similar components of the present invention;

[0046] Figure 5 Schematic diagram of clustering mode 3 of strain measurement points of similar components of the present invention.

[0047] Figure 6 Schematic diagram of clustering mode 1 of strain measurement points of dissimilar components of the present invention;

[0048] Figure 7 Schematic diagram of clustering mode 2 of strain measurement points of dissimilar components of the present invention;

[0049] Figure 8 Schematic diagram of clustering mode 3 of strain measurement points of dissimilar components of the present invention;

[0050] Figure 9 Flowchart of data distortion identification based on density clustering of the present invention;

[0051] Figure 10 Schematic diagram of C2-4 outlier identification of the present invention;

[0052] Figure 11 Schematic diagram of C2-4 constant offset identification of the present invention;

[0053] Figure 12 Schematic diagram of C2-3 data jamming identification of the present invention;

[0054] Figure 13 Overall flowchart of the method of the present invention. Detailed implementation

[0055] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] As Figure 1 shown, the present invention provides a data distortion recognition method based on correlation density clustering. For the historical time series X = {x1(t), x2(t),..., x n (t)} of n strain measurement points of a strain sensor, where 0 ≤ t ≤ m, the correlation degree between each measurement point sequence is measured from the perspectives of the overall displacement difference representing the similarity of n strain measurement points, the overall first-order slope difference representing the similarity of the development trend, and the overall second-order slope difference representing the similarity of the development speed based on grey type-B correlation, and a classification model is constructed, and real-time monitoring data is recognized based on the classification model.

[0057] Among them, theoretically, similarity and resemblance can be expressed as:

[0058] Resemblance:

[0059] Similarity:

[0060]

[0061] Where is the overall displacement difference, is the overall first-order slope difference, representing the similarity of the development trend; is the overall second-order slope difference, representing the similarity of the development speed. Furthermore, the correlation degree is defined as:

[0062]

[0063] Utilize the core idea in the original theory of measuring the correlation degree between sequences from the two perspectives of similarity and resemblance.

[0064] When calculating the correlation degree, according to the characteristics of the monitoring data, balance the data scales of the two measurement indicators of similarity and resemblance, and at the same time consider the actual engineering application situation. Referring to the weights of these two measurement indicators, multiply the overall displacement difference, the overall first-order slope difference, and the overall second-order slope difference by weight coefficients respectively to obtain the final correlation feature as shown in Equation (5):

[0065]

[0066] If more dependence is placed on proximity features in cluster analysis, then increase the weight w1 and decrease the weights w2 and w3. If more dependence is placed on similarity features, then increase the weights w2 and w3 and decrease the weight w1.

[0067] The normalization process for the historical time series is as follows:

[0068] For the historical time series of n measurement points over a certain period, divide the step-by-step clustering window. The length w determines the time segment for anomaly localization. The shorter the window, the stronger the immediacy of the determination.

[0069] Use z-score to normalize each sequence in the window to eliminate the dimension or balance the measurement scales of each strain sensor.

[0070] Let the time series of a certain window after normalization be shown in Equation (6).

[0071]

[0072] where n is the measurement point number and m is the time number.

[0073] The smoothing and denoising of the historical time series after normalization is as follows:

[0074] Select the Savitzky-Golay method for smoothing and denoising. Combine convolution and polynomial regression to achieve smoothing filtering. Different smoothing effects are achieved by adjusting the sliding window and the fitting order, expressed as:

[0075]

[0076] where 2l + 1 is the sliding window length, x k represents the center of the sliding window, and h i is the smoothing coefficient, obtained by fitting a polynomial using the least squares method;

[0077] Among them, the smaller the polynomial fitting order and the longer the window length, the more significant the smoothing effect.

[0078] Constructing the classification model includes:

[0079] Perform differential processing on the overall displacement difference, overall first-order slope difference, and overall second-order slope difference of the window sequence after smoothing and denoising by column respectively, and merge them with the matrix X into a new matrix. At the same time, add the weight coefficients to the new matrix to obtain the clustering feature matrix Y;

[0080] The new matrix includes:

[0081] Overall displacement difference matrix:

[0082]

[0083] First-order difference matrix:

[0084] Second-order difference matrix:

[0085] Clustering feature moment

[0086] where dif ij = x i(j+1) - x ij , w1, w2, w3 are weights. If more dependence on proximity features in clustering analysis, then increase w1 and decrease w2 and w3; if more dependence on similarity features, then increase w2 and w3 and decrease w1.

[0087] Perform clustering analysis. In the process of clustering calculation, determine the Manhattan distance as the calculation method for the distance between sequences;

[0088] The Manhattan distance is as shown in Equation (11):

[0089]

[0090] Determine the parameters (∈, MinPts) according to the actual engineering needs, obtain the final clustering result, and construct a classification model. Minpts is the number of clusters generated by clustering, and ∈ refers to the weight parameters of w1, w2, w3.

[0091] Identify real-time monitoring data based on the classification model as follows:

[0092] For the real-time collected monitoring data, perform clustering analysis step by step with the same window length w;

[0093] When the clustering results of all measurement points are consistent with the clustering model, it is determined that the data for this period is normal;

[0094] When the clustering results are inconsistent and outliers appear, it is regarded as distorted;

[0095] When the clustering results are inconsistent and there are no outliers, further data distortion identification is performed.

[0096] The present invention also provides an application of a data distortion identification method, which is applied to the health monitoring of wood structures. The measurement points include strain measurement points of homogeneous components and strain measurement points of heterogeneous components.

[0097] The following provides specific embodiments for illustration:

[0098] As Figure 1 and Figure 2 shown, taking a wood-structured ancient building with a structural health monitoring system arranged as an object, the relevant diagram of the strain sensor is as attachedFigure 1 , attached Figure 2 As shown, the sampling time interval is 10 minutes, the distributed clustering window length is 24 hours, and there are 144 sampling points in each window.

[0099] Perform normalization and smoothing denoising preprocessing according to the above method. Based on the characteristics of the collected data, determine the smoothing order to be 3 and the moving window length to be 25 to ensure that the noise can be filtered while minimizing the impact of the smoothing process on the distorted data.

[0100] When constructing a new feature attribute matrix after smoothing and denoising, considering the strain data characteristics of this application and the requirements for similarity features and resemblance features, select the weight coefficients as w1 = 0.1, w2 = 0.45, and w3 = 0.45 respectively.

[0101] Cluster analysis of strain measurement points of similar components:

[0102] First, take 12 strain measurement points of the through columns as an example and use historical strain data for training. The parameters (ε, MinPts) in the clustering process are (5.4, 1), which basically meet the requirements for obtaining a stable clustering result and have good robustness. The final clustering results show three patterns:

[0103] Clustering pattern one: After clustering analysis, all 12 measurement points in this pattern are classified into one category, and there are two presentation forms. One is that its time series diagram shows a form similar to a sine curve, and the other is that there is no fixed form, as attached Figure 3 .

[0104] Clustering pattern two: In this pattern, among the 12 measurement points, except for three measurement points C1-2, C1-3, and C1-4 that are easily classified as outliers, the other measurement points are classified into one category. The presentation forms of the measurement points classified into one category also show two states. One is that the time series diagram is also approximately in the form of a sine curve, and the other is that there is no fixed form, as attached Figure 4 .

[0105] Clustering pattern three: In this pattern, the time series diagrams of each measurement point are disorderly, and the clustering situation has no obvious regularity. Most measurement points cannot be classified, and there is no consistent clustering result, as attached Figure 5 .

[0106] Cluster analysis of strain measurement points of different components:

[0107] Based on the above general column strain analysis, the strain of the frame beam is added, and cluster analysis is carried out on a total of 24 measurement points. The clustering results presented under the two conditions of relatively high and low environmental humidity are basically the same as the results of the simple general column strain. Except for the three column strain measurement points C1-2, C1-3, and C1-4, B1-4, B3-2, and B3-4 are the measurement points in the beam strain that are easily classified as outliers. It can be seen from this that the strain clustering results of different types of components have consistent regularities, and it also reflects that the clustering analysis method based on the grey theory has good robustness, which can provide a reliable clustering model for the subsequent real-time identification of data distortion, such as attached Figure 6 , Figure 7 , Figure 8 .

[0108] Verification of the distortion data recognition method:

[0109] Construct a clustering model: Establish a classification model with 9 measurement points such as C1-1, C2-1, C2-2, and C2-3. Then the working conditions during training can be adjusted to 2 cases, corresponding to constructing two clustering models, namely "clustered into one category" and "irregular". At this time, the proportion of the corresponding mode one during training is shown in Table 2, and the total proportion reaches 86%. This shows that the clustering method can simultaneously handle the data distortion situation of not less than 86% in the time series of these 9 measurement points. Therefore, when performing clustering analysis on the real-time multi-measurement point time series, if it is the same as model one, it means the data is normal; if outliers appear, the corresponding measurement points can be determined as data distortion; if different classification situations appear, further determination of structural anomalies will be carried out for subsequent processing. The specific identification process is as attached Figure 9 shown.

[0110] Table 2 Establish a clustering model

[0111]

[0112] Using the 9 general column strain measurement points in June 2020 to conduct simulation distortion data recognition verification, add outlier, constant offset, and data stuck three forms of distortion simulation data to the strain sequences of some measurement points on the 4th, 6th, 8th, 14th, and 15th. This method can identify the corresponding data distortion measurement points, as shown in attached Figure 10 , 11 , Figure 12.

[0113] To further explore the sensitivity of this method to the recognition of three types of data distortion, the abnormal data is simulated at different levels.

[0114] (1) Outlier

[0115] By analyzing the historical strain time series, on a yearly basis, the average strain fluctuation range at each measuring point is less than 400 (με). Therefore, the single-point outlier anomalies are divided into three grades, as shown in Table 3. For each grade, 10 sets of outlier distortion data are applied respectively, and the anomaly value applied in each simulation is a random value within the grade range. After analysis, Table 3 lists the probability of the method identifying the measuring points containing outliers under each grade. The results show that the method can effectively identify absolute outliers above 300 (με), but cannot identify outliers below 200 (με).

[0116] Table 3 Comparison of outlier identification sensitivity

[0117] Table 3 Comparison of outlier identification sensitivity

[0118]

[0119] (2) Data stuck

[0120] For the continuous data distortion type of data stuck, it is necessary to explore the sensitivity of the method to the time span of continuous distorted data. At a sampling frequency with an average of 10 minutes, after calculation, the method can effectively identify data stuck for more than 40 consecutive sampling points, that is, it is applicable to situations with a time span of more than 6.7 hours. Since the method performs clustering analysis under the condition that the window length is 24 hours, this type of data distortion can be identified in a timely manner.

[0121] (3) Constant offset

[0122] Constant offset also belongs to the type of continuous data distortion, and it contains two variables: the continuous distortion time span and the offset. To control variables, referring to the results of data stuck, under the condition that the continuous time span is 7 hours, the sensitivity of the method to the offset is explored. The offset is divided into three grades, and the range of each grade is shown in Table 4. For each grade, 10 sets of outlier distortion data are applied respectively, and the offset applied in each simulation is a random value within the grade range. The results show that the method can effectively identify this type of distorted data when the absolute value of the offset is greater than 50 (με).

[0123] Table 4 Comparison of outlier identification sensitivity

[0124] Table 4 Comparison of outlier identification sensitivity

[0125]

[0126] Through the sensitivity analysis of three types of anomalies, it can be concluded that due to the limitations of the smoothing and noise reduction method, the sensitivity to the identification of outliers is lower than that of other types of data distortion identification, while the sensitivity to the identification of continuous data distortion types such as constant offset and data jamming is better.

[0127] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.

Claims

1. A data distortion recognition method based on association density clustering, characterized in that Including: For the historical time series of n strain measurement points of the strain sensor , where Based on the grey B-type correlation, the correlation degree between each measurement point sequence is measured from the perspectives of the overall displacement difference representing the similarity of n strain measurement points, the overall first-order slope difference representing the similarity of development trends, and the overall second-order slope difference representing the similarity of development speeds, and a classification model is constructed, and the real-time monitoring data is identified based on the classification model, where the strain measurement points include column strain measurement points and beam strain measurement points; Building a classification model includes: The window sequences after smooth denoising are respectively subjected to difference processing of the overall displacement difference, the overall first-order slope difference, and the overall second-order slope difference by column and merged with the matrix X to form a new matrix. At the same time, the weight coefficients are added to the new matrix to obtain the clustering feature matrix Y ; Performing clustering analysis, and determining the Manhattan distance as the calculation method for the distance between sequences during the clustering calculation process; Determine parameters according to the actual engineering needs , obtain the final clustering result, and construct a classification model; The new matrix includes: Overall displacement difference matrix: ; First-order difference matrix: ; (8) Second-order difference matrix: ; (9) Clustering feature matrix ; (10) Among them, , is the weight. If more dependence is placed on similarity features in cluster analysis, then increase , and decrease and ; if more dependence is placed on resemblance features, then increase and , and decrease .

2. The data distortion recognition method based on association density clustering according to claim 1, wherein It also includes the normalization processing of the historical time series: Dividing the historical time series of n measuring points in a certain period into step-by-step clustering windows. The length w determines the time period for anomaly location. The shorter the window, the stronger the immediacy of the determination; Using z-score to perform normalization processing on each sequence in the window to eliminate the dimension or balance the measurement scales of each strain sensor; Suppose the time series of a certain window after standardization is shown in Equation (6), (6) Where n is the measuring point serial number and m is the time serial number.

3. The data distortion recognition method based on correlation density clustering according to claim 2, characterized in that It also includes smoothing and denoising the normalized historical time series: The Savitzky-Golay method is selected for smoothing and denoising. Combining convolution and polynomial regression to achieve smoothing filtering, different smoothing effects are achieved by adjusting the sliding window and the fitting order, which is expressed as: (7) Among them, 2 l +1 is the sliding window length, represents the center of the sliding window, is the smoothing coefficient, which is obtained by fitting a polynomial with the least squares method; Among them, the smaller the polynomial fitting order and the longer the window length, the more significant the smoothing effect.

4. A data distortion recognition method based on association density clustering according to claim 1, characterized in that The Manhattan distance is as shown in Equation (11): (11).

5. The data distortion recognition method based on association density clustering according to claim 4, wherein Identifying real-time monitoring data based on the classification model includes: Performing clustering analysis step by step on the real-time collected monitoring data with the same window length w; When the clustering results of all measuring points are consistent with the clustering model, it is determined that the data during this period is normal; When the clustering results are inconsistent and outliers appear, it is regarded as distorted; When the clustering results are inconsistent and there are no outliers, further data distortion identification is performed.

6. An application of the data distortion recognition method according to any one of claims 1-5, characterized in that, Applied to the health monitoring of wooden structures.

7. The application according to claim 6, wherein The measuring points include strain measuring points of homogeneous components and strain measuring points of heterogeneous components, and the strain measuring points include column strain measuring points and beam strain measuring points.

Citation Information

Patent Citations

  • Structure health monitoring method based on distributed strain dynamic test

    CN101221104A

  • Principal element degree of association sensor fault detection method and apparatus based on density clustering

    CN105894027A