Structural health monitoring multi-type data anomaly diagnosis and processing method based on Bernoulli sparse Bayesian model

By combining the Bernoulli sparse Bayesian model with the moving median method and sparse Bayesian learning, the problem of identifying and processing multiple types of data anomalies in existing technologies is solved, achieving high-precision and robust monitoring data processing, which is suitable for structural health monitoring systems.

CN121579923APending Publication Date: 2026-02-27CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511781657.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-29
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing structural health monitoring methods struggle to efficiently and accurately identify and process various types of data anomalies when dealing with monitoring data under complex operating conditions, leading to incorrect assessments of structural condition and even masking potential safety hazards.

Method used

We employ a Bernoulli sparse Bayesian model-based approach, using the moving median method for initial screening and sparse Bayesian learning (SBL) for refined modeling, combined with two regression analysis strategies to identify and process various types of outlier data.

Benefits of technology

It achieves high-precision and robust multi-type anomaly identification and processing, ensuring the reliability of monitoring data, applicable to different dynamic signals, adapting to complex data fluctuations, and improving the accuracy and security of monitoring data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121579923A_ABST
    Figure CN121579923A_ABST
Patent Text Reader

Abstract

The invention provides a multi-type data exception diagnosis and processing method based on a Bernoulli sparse Bayesian model, aiming at the problems that multi-type data exceptions coexist in structure health monitoring, and a traditional method is difficult in unified identification and poor in robustness. According to the method, the probability that each data point is normal or abnormal is deduced through a Bernoulli sparse Bayesian model, and abnormal identification and classification of three types of data including an outlier, a deviation value and a drift value are achieved. Two regression analysis strategies (including a two-stage rejection reconstruction method and a data correction method) suitable for different abnormal conditions are further provided, and data exception processing is carried out; wherein the two-stage rejection reconstruction method is specially used for processing data with dominant outliers, and the data correction method is suitable for data with long-section abnormities such as deviation values and drift values. Experimental results show that the method provided by the invention is a reliable multi-type data anomaly diagnosis and processing method, and data processed by the method and subjected to regression analysis can be regarded as reliable real monitoring data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of Structural Health Monitoring (SHM) data processing technology, specifically involving a method for diagnosing and processing multiple types of data anomalies in structural health monitoring based on a Bernoulli sparse Bayesian model. The purpose is to identify and classify multiple types of data anomalies in the monitoring data and carry out data anomaly processing to obtain real and reliable monitoring data. Background Technology

[0002] With the rapid development of global infrastructure construction, critical engineering structures such as bridges, dams, and large buildings inevitably face severe challenges during their service life, including material aging, environmental erosion, extreme loads, and cumulative damage from operational loads. These factors lead to a gradual degradation of structural performance, posing increasingly severe tests to their safe operation and maintenance.

[0003] Structural Health Monitoring (SHM), a key technology capable of real-time, online assessment of structural health and early warning of potential risks, has become a research hotspot and cutting-edge development in fields such as civil engineering and aerospace. SHM systems continuously collect physical data reflecting the structural state by deploying sensor networks at critical structural locations. Based on this data, they perform in-depth analysis and intelligent assessment, thereby achieving early diagnosis of structural performance degradation, damage identification, and life-cycle safety assurance.

[0004] In the entire SHM system, high-quality monitoring data is the cornerstone of all subsequent analysis. However, in practical engineering applications, the data acquisition process is highly susceptible to interference from complex operating conditions. The causes of data anomalies are varied, primarily including: sensor malfunctions, aging, zero-point drift, or calibration errors; data acquisition equipment malfunctions, unstable power supply, signal transmission interruptions, or packet loss; and complex factors such as strong environmental noise, electromagnetic interference, and extreme weather.

[0005] These factors inevitably lead to the inclusion of various outliers in the collected monitoring data. These outliers manifest in various forms, but can be mainly summarized as: outliers with short-term, sharp fluctuations; deviations where the overall data suddenly increases or decreases abruptly at a certain point in time; and drifts where the data slowly but continuously deviates from the normal trajectory over time.

[0006] The presence of data anomalies can severely distort the true response characteristics of a structure. If these anomalies are not processed and used directly in subsequent analysis, they will lead to incorrect assessments of the structural condition and may even mask real safety hazards, resulting in catastrophic consequences. Therefore, efficient and accurate automated preprocessing of SHM monitoring data, especially the precise identification and appropriate handling of various types of outliers, is the primary step in ensuring the reliability of the entire monitoring system and a key scientific and technological problem that urgently needs to be solved in the field of SHM.

[0007] Currently, various outlier detection techniques have been applied in the field of structural health monitoring. However, these methods still have certain limitations when dealing with complex monitoring data in actual engineering projects. Specifically, statistical model-based methods, represented by the 3σ rule and box plots, are based on the assumption that normal data follows a specific probability distribution. However, these methods are highly dependent on the data distribution, while actual monitoring data often fails to strictly meet the ideal distribution assumption. Furthermore, the statistics used to identify outliers are easily affected by outliers themselves, potentially leading to masking or flooding effects, thus affecting the accuracy of the judgment. In addition, these methods are usually more effective for obvious outliers, but have limited ability to identify gradual anomalies such as offsets and drifts. Another common type of method is detection techniques based on proximity or distance, including the k-nearest neighbor algorithm and the local anomaly factor algorithm. The core idea of ​​these methods is to define outliers as isolated points that deviate significantly from the dense area of ​​normal data. However, their computational complexity is high, making it difficult to achieve efficient processing when dealing with massive amounts of monitoring data. At the same time, the detection effect is highly sensitive to parameter settings and lacks adaptability in scenarios with dynamically changing data. In addition, detection methods based on clustering and time series modeling have also been applied. The former also faces the problems of sensitivity to parameter selection and insufficient automation; while the latter, although considering the temporal characteristics of the data, has a relatively complex modeling process and limited ability to capture the non-stationary and nonlinear features commonly found in structural health monitoring data.

[0008] In summary, SHM monitoring data is characterized by its massive volume, non-stationarity, nonlinearity, strong temporal correlation, and diverse anomaly types. Existing traditional detection methods generally suffer from poor robustness, sensitivity to parameter settings, difficulty in simultaneously and effectively handling multiple mixed anomaly types, and computational efficiency that cannot meet the real-time requirements of massive data. Therefore, developing a new method that can efficiently, accurately, and automatically identify and process multiple types of outliers in SHM monitoring data has significant engineering and application value. Summary of the Invention

[0009] Purpose of the Invention: To address this challenge, this study proposes a method for diagnosing and processing multiple types of data anomalies in structural health monitoring based on a Bernoulli sparse Bayesian model. This method transforms anomaly detection into probabilistic inference of the hidden states of data. It achieves high-precision identification and classification of multiple types of anomalies through preliminary screening using the moving median method and refined modeling using Sparse Bayesian Learning (SBL). Regression analysis is performed on the identified and classified data using two data anomaly processing strategies applicable to different situations. The resulting regression model can be considered reliable and accurate monitoring data.

[0010] To achieve the above technical objectives, this invention discloses a method for diagnosing and processing multiple types of data anomalies in structural health monitoring based on a Bernoulli sparse Bayesian model. The specific steps are as follows:

[0011] An SHM system is installed on the engineering structure. The SHM system includes multiple monitoring sensors of different types installed at different locations on the engineering structure, as well as a computer for recording various monitoring sensor data.

[0012] Add simulated quantities to the monitoring data under healthy conditions to test the model's accuracy in identifying unconventional errors, and to verify the model's accuracy and practicality;

[0013] First, the monitoring data is initially screened using a model. A moving median filter is applied to the monitoring data to generate a local baseline. This baseline is used to robustly fit the local nonlinear trend of the monitoring data. Then, the residual between the monitoring data and this local baseline is calculated and applied to a strict statistical threshold based on the absolute deviation of the median. Finally, the initial screening of outliers is completed, providing a data sequence with relatively few data anomalies for subsequent fine modeling.

[0014] The data sequence after preliminary screening is used to infer whether the data points in the original monitoring data are abnormal, and the data anomalies are classified according to different residual patterns of the abnormal data. They can be divided into three types of data anomalies: outliers, deviations, and drifts.

[0015] After anomaly identification, the theoretical model proposes and compares two strategies for handling anomalies in regression analysis data: The first is a two-stage elimination and reconstruction method, which is suitable for anomaly monitoring data dominated by outliers. First, a preliminary Sparse Bayesian Learning (SBL) model is trained on the original data containing anomalies to identify the anomalies. Second, all samples identified as anomalies are completely removed from the dataset. Finally, a more accurate SBL model is trained on the dataset with some anomalies removed to effectively isolate the negative impact of data anomalies on the accuracy of the final model. The second is a data correction method, which is suitable for monitoring data with long-term anomalies such as bias and drift. Through the good data filling characteristics of the Sparse Bayesian algorithm, filled values ​​that conform to the nonlinear trend of the data can be generated to avoid problems caused by the discontinuity of the data sequence, thereby improving the accuracy of the regression model.

[0016] Data processed through regression analysis can be considered reliable and accurate monitoring data.

[0017] Furthermore, the Bernoulli sparse Bayesian model employs outlier identification and analysis based on the sparse Bayesian learning theory model and the Bernoulli likelihood model. The basic form of the sparse Bayesian learning theory model is as follows:

[0018] (1)

[0019] in It is a basis function vector, usually using a Gaussian kernel function. This maps the input data to a high-dimensional feature space. It is the corresponding weight vector. It is noise that follows a zero-mean Gaussian distribution, i.e. .

[0020] Therefore, the likelihood function of the objective value is:

[0021] (2)

[0022] in yes The design matrix, its elements .

[0023] The core of SBL lies in the weighting A hierarchical prior is introduced. Specifically, for each weight... Assume an independent, zero-mean Gaussian prior distribution:

[0024] (3)

[0025] in It is a control weight The hyperparameters of variance. By analyzing the hyperparameters... and noise accuracy By placing these hyperparameters under a Gamma distribution prior and using the type II maximum likelihood method for iterative optimization, SBL can automatically optimize most of the unrelated basis functions. Pushing it toward infinity, thus increasing its weight. The posterior distribution is concentrated at zero, thus achieving sparsity of the model.

[0026] After training is complete, for a new input point SBL can not only provide the predicted mean It can also provide the prediction variance. This constitutes a complete prediction probability distribution:

[0027] (4)

[0028] (5)

[0029] (6)

[0030] and They are weights The posterior mean and covariance. Prediction variance. This constitutes the confidence interval of the predicted value, which is the key basis for subsequent anomaly probability identification.

[0031] The basic form of the Bernoulli likelihood model is:

[0032] Assuming each data point There is a binary hidden variable behind it. ,in This indicates that the value is an outlier. This indicates that it is a normal value.

[0033] The latent variable follows a Bernoulli distribution:

[0034] (7)

[0035] in It is the prior probability that a data point is an anomaly.

[0036] This invention does not directly estimate Instead, it uses the prediction results of the SBL model to... The state is inferred posteriorly. Specifically, the SBL model is used for each input. The predicted mean and the predicted standard deviation It is possible to build a Confidence interval If the observed value If a value falls outside this range, it is considered an anomaly. This decision-making process can be formalized as the analysis of latent variables. Judgment:

[0037] (8)

[0038] This Bernoulli classification method, based on sparse Bayesian learning of probability output, is more objective and robust than the traditional fixed threshold method, and can adaptively respond to dynamic changes in data.

[0039] Furthermore, the overall approach of this method is to first perform preliminary screening and then achieve accurate identification through probabilistic inference, aiming to achieve high-precision and robust identification of outliers.

[0040] The first stage of this method employs the moving median method to construct a locally adaptive baseline. This method exhibits good robustness and can effectively track the local central tendency of the data without being affected by extreme outliers. By calculating the deviation between the original data points and the moving median baseline, the most significant outliers can be identified and temporarily removed. The main goal of this stage is to provide a data sequence with relatively few data anomalies for subsequent refined modeling.

[0041] After the initial screening, the next step is to make a precise probability inference as to whether the data is abnormal.

[0042] This invention first defines the problem from a probabilistic perspective. Its core assumption is that for each data point in a time series... Behind each of these lies an implicit state variable that cannot be directly observed. .

[0043] when When the time is specified, it indicates that the point is a normal point.

[0044] when When the value is 0, it indicates that the point is an outlier.

[0045] This binary state variable can be considered as being sampled from a Bernoulli distribution:

[0046] (9)

[0047] in This is the prior probability that a point is an anomaly. The goal of the entire anomaly detection process is to use the observed data... To infer the most likely underlying state .

[0048] The model treats the data processed in the first stage as observed samples sampled from a normal distribution. Training the SBL regression model on this batch of data, which has relatively few anomalies, essentially constructs a conditional probability model for the normal state. Leveraging SBL's powerful non-linear fitting capabilities, the model can learn the complex trends hidden behind the data. After training, the final baseline generated by the model can be considered the expected value of the normal probability distribution. .

[0049] Once an accurate normal model is available, state inference can be performed on each point in the original contamination data. For a point to be measured... Calculate the residual between it and the corresponding normal expected value. :

[0050] (10)

[0051] This residual can be considered a key statistic. If a point truly belongs to the normal state ( If the residual is large, then its residual should be small, consistent with the noise distribution predicted by the model. Conversely, if the residual is significantly large, then the probability that this point was generated by a normal model is extremely low, and therefore there is sufficient reason to believe that its state is abnormal. ).

[0052] Therefore, the threshold judgment method used here can be regarded as an efficient approximation of Bayesian decision-making, which constructs a threshold predicted by the SBL model. The confidence interval is used as a criterion to determine whether the state of a point is normal or abnormal, thus completing the inference of the implicit Bernoulli state of the data point;

[0053] After completing the Bernoulli state deduction (determined to be abnormal), After that, the present invention further applies to all those marked as The points are further subdivided, and by analyzing the residual stability of continuous outlier segments, they are classified into three types of data anomalies: outliers, deviations, and drifts.

[0054] Furthermore, using the identified data anomalies, regression analysis of the data is performed through a Bernoulli sparse Bayesian model:

[0055] Based on the aforementioned high-precision anomaly identification methods, this study constructs and compares two anomaly handling strategies for regression analysis data: the two-stage elimination and reconstruction method and the data correction method. The two-stage elimination and reconstruction method is a strategy that effectively isolates the negative impact of outliers and pursues the highest predictive accuracy of the final model. This method first trains a preliminary SBL model on the original monitoring data containing anomalies. Then, based on the identification framework, it locates as many outlier samples as possible and completely removes them from the dataset, forming a training subset with relatively few data anomalies. On this training subset, a more accurate SBL model is trained. Finally, this refined model is used to predict all original time points, thereby reconstructing a smooth and continuous high-precision data curve. This method is suitable for anomaly monitoring data dominated by outliers. The data correction method ensures data continuity through the data imputation function of SBL, exhibiting excellent robustness. This method avoids numerical oscillations caused by discontinuities in the data sequence by replacing rather than deleting outliers. This method ensures the continuity and integrity of the training data; therefore, even without fine-tuning of hyperparameters, the model still exhibits high robustness, and its regression results have clear physical meaning. The data correction method provides a well-conditioned dataset for model optimization, enabling the optimization process to converge successfully to an effective global optimum, even though it is rapid. This ensures the reliability of the final model. This method is suitable for monitoring data with long-term data anomalies such as deviation values ​​and drift values.

[0056] A computer device includes: a processor, a memory, and a network interface;

[0057] A computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor, comprising the above-described method for diagnosing and processing multiple types of data anomalies in structural health monitoring based on a Bernoulli sparse Bayesian model.

[0058] Beneficial effects: (1) This method adopts a preliminary screening and then achieves accurate identification through probability inference. By combining the local adaptive preliminary screening of the moving median method and the accurate modeling of SBL, it can efficiently and accurately identify a variety of mixed anomalies, including outliers, deviations and drifts. At the same time, it has strong adaptability to noise and complex data fluctuations. (2) This invention proposes two regression analysis data anomaly processing strategies for different states, and the processed data can be considered as reliable real monitoring data. The first method is the two-stage elimination and reconstruction method, which is suitable for abnormal monitoring data dominated by outliers. First, a preliminary SBL model is trained on the original data containing outliers to identify the outliers. Second, all samples that are judged to be outliers are completely removed from the dataset. Finally, a more accurate SBL model is trained on the dataset that has removed some data outliers to effectively isolate the negative impact of data outliers on the accuracy of the final model. The second method is the data correction method, which is suitable for monitoring data with long-span data outliers such as deviation values ​​and drift values. Through the good data filling characteristics of the sparse Bayes algorithm, filling values ​​that conform to the nonlinear trend of the data can be generated to avoid problems caused by the discontinuity of the data sequence, thereby improving the accuracy of the regression model. (3) This method does not rely on specific structural physics or finite element models. It has been verified on various types of monitoring data such as strain, acceleration and displacement, proving its wide applicability and high reliability for different dynamic signals. It is a reliable and universal data preprocessing tool. Attached Figure Description

[0059] Figure 1 This is a schematic diagram of a method for diagnosing and processing multiple types of data anomalies in structural health monitoring based on a Bernoulli sparse Bayesian model, according to the present invention.

[0060] Figure 2 This is a schematic diagram of the sensor arrangement on the Tsing Ma Bridge in Hong Kong, as described in an embodiment of the present invention.

[0061] Figure 3 This invention relates to the identification of outliers in strain monitoring data in this embodiment.

[0062] Figure 4 This invention relates to the identification of outliers in acceleration monitoring data in embodiments of the invention.

[0063] Figure 5 This invention relates to the identification of outliers in displacement monitoring data in embodiments of the invention.

[0064] Figure 6 For regression analysis of monitoring data dominated by outliers using the data correction method based on the Bernoulli sparse Bayesian model;

[0065] Figure 7Regression analysis of monitoring data dominated by outliers using a two-stage elimination and reconstruction method based on the Bernoulli sparse Bayesian model;

[0066] Figure 8 This invention provides an analysis of the model instability of the two-stage elimination and reconstruction method based on the Bernoulli sparse Bayesian model when dealing with anomalies in long data segments.

[0067] Figure 9 This invention provides an analysis of model instability when using a data correction method based on a Bernoulli sparse Bayesian model to handle anomalies in long data segments. Detailed Implementation

[0068] The embodiments of the present invention will be further described below with reference to the accompanying drawings:

[0069] like Figure 1 As shown, this invention proposes a method for diagnosing and processing multiple types of data anomalies in structural health monitoring based on a Bernoulli sparse Bayesian model. This method has several beneficial effects. First, it can achieve high-precision and robust anomaly identification. By first conducting preliminary screening and then achieving accurate identification through probabilistic inference, it combines locally adaptive preliminary screening using the moving median method with accurate modeling using Sparse Bayesian learning (SBL). This method can efficiently and accurately identify various mixed anomalies and has strong adaptability to complex data fluctuations. Second, this invention provides two regression modeling strategies applicable to different situations of anomaly monitoring data: a two-stage elimination and reconstruction method and a data correction method. The first is the two-stage elimination and reconstruction method, which is suitable for anomaly monitoring data dominated by outliers; the second is the data correction method, which is suitable for monitoring data with long-term data anomalies such as deviation values ​​and drift values. Third, this invention demonstrates excellent model stability and performance. When dealing with monitoring data dominated by outliers, the two-stage elimination and reconstruction method demonstrates higher accuracy. When handling long stretches of continuous data anomalies, the data correction method exhibits better robustness, providing a well-formed dataset for model optimization, enabling rapid convergence to an effective solution, and maintaining a moderate weight distribution, thus ensuring model stability. Finally, both strategies possess broad versatility and applicability, having been validated on various monitoring data such as strain, acceleration, and displacement, proving their reliability and high accuracy for different dynamic signals.

[0070] The Bernoulli Sparse Bayes model employs a combination of Sparse Bayesian learning theory and Bernoulli likelihood model for outlier identification and analysis. The overall process follows a principle of initial screening followed by probabilistic inference. First, the moving median method is used for preliminary data screening, providing a data sequence with relatively few anomalies for subsequent refined SBL modeling. After the initial screening, the precise diagnosis phase begins, which combines SBL and the Bernoulli model.

[0071] SBL was first used to construct probabilistic models of normal states. Its basic theoretical model is as follows:

[0072]

[0073] in It is a basis function vector, usually using a Gaussian kernel function. This maps the input data to a high-dimensional feature space. It is the corresponding weight vector. It is noise that follows a zero-mean Gaussian distribution, i.e. .

[0074] The core of SBL theory lies in the weighting A hierarchical prior is introduced. Specifically, for each weight... Assume an independent, zero-mean Gaussian prior distribution:

[0075]

[0076] in It is a control weight The hyperparameters of variance. By analyzing the hyperparameters... and noise accuracy By placing these hyperparameters under a Gamma distribution prior and using the type II maximum likelihood method for iterative optimization, SBL can automatically optimize most of the unrelated basis functions. Pushing it toward infinity, thus increasing its weight. The posterior distribution is concentrated at zero, thus achieving sparsity of the model.

[0077] After training is complete, for a new input point SBL can not only provide the predicted mean It can also provide the prediction variance. This constitutes a complete prediction probability distribution:

[0078]

[0079]

[0080]

[0081] and They are weights The posterior mean and covariance. Prediction variance. This constitutes the confidence interval of the predicted value, which is the key basis for subsequent anomaly probability identification.

[0082] The basic form of the Bernoulli likelihood model is:

[0083] Assuming each data point There is a binary hidden variable behind it. ,in This indicates that the value is an outlier. This indicates that it is a normal value.

[0084] The latent variable follows a Bernoulli distribution:

[0085]

[0086] in It is the prior probability that a data point is an anomaly.

[0087] This invention does not directly estimate Instead, it uses the prediction results of the SBL model to... The state is inferred posteriorly. Specifically, the SBL model is used for each input. The predicted mean and the predicted standard deviation It is possible to build a Confidence interval If the observed value If a value falls outside this range, it is considered an anomaly. This decision-making process can be formalized as the analysis of latent variables. Judgment:

[0088]

[0089] This Bernoulli classification method, based on sparse Bayesian learning of probability output, is more objective and robust than the traditional fixed threshold method, and can adaptively respond to dynamic changes in data.

[0090] Once an accurate normal model is available, state inference can be performed on each point in the original contamination data. For a point to be measured... Calculate the residual between it and the corresponding normal expected value. :

[0091]

[0092] The residual can be considered a key statistic. If a point belongs to a normal state, its residual should be small, consistent with the noise distribution predicted by the model. Conversely, if the residual is significantly large, the probability that the point was generated by a normal model is extremely low, thus providing sufficient reason to believe that its state is abnormal.

[0093] Subsequently, using the identified data anomalies of various types, regression analysis of the data is performed through a Bernoulli Sparse Bayes model. Based on the above high-precision anomaly identification method, this invention constructs the following two advanced data anomaly processing strategies: a two-stage elimination and reconstruction method and a data correction method. The two-stage elimination and reconstruction method is suitable for anomaly monitoring data where only outliers dominate. First, a preliminary SBL model is trained on the original data containing anomalies to identify the anomalies. Second, all samples identified as anomalies are completely removed from the dataset. Finally, a more accurate SBL model is trained on the dataset where some data anomalies have been removed to effectively isolate the negative impact of data anomalies on the accuracy of the final model. The data correction method is suitable for monitoring data with long-span data anomalies such as bias and drift. Through the good data imputation characteristics of the sparse Bayes algorithm, imputation values ​​that conform to the nonlinear trend of the data can be generated to avoid problems caused by the discontinuity of the data sequence, thereby improving the accuracy of the regression model.

[0094] Finally, the data processed through regression analysis can be considered reliable and accurate monitoring data.

[0095] The following example, using the Tsing Ma Bridge in Hong Kong, illustrates the present invention's method for diagnosing and processing multiple types of data anomalies in structural health monitoring based on a Bernoulli sparse Bayesian model:

[0096] The Tsing Ma Bridge in Hong Kong is renowned for its unique double-deck structure. The upper deck carries a six-lane highway, while the lower deck includes a railway and two backup lanes to cope with extreme weather conditions such as typhoons. The bridge has a main span of 1,377 meters, a main tower reaching 206 meters, and utilizes over 22,000 kilometers of high-strength steel cables. To ensure long-term safety, the bridge is equipped with a large-scale structural health monitoring system, employing hundreds of sensors to monitor key data such as traffic loads, wind speed, and bridge stress in real time.

[0097] This time, the following was adopted: Figure 2 The data collected by some sensors in the transverse frame at mileage marker 23623 is used as the research object.

[0098] Figures 3-5 The process of identifying outliers in three types of monitoring data—strain, acceleration, and displacement—was demonstrated.

[0099] Figure 3(a) illustrates the data preparation phase of this experiment. The gray scatter points in the figure represent the original monitoring signal used as a reference, which exhibits a certain nonlinear fluctuation trend. To simulate complex situations that may occur in the real world, three typical data anomalies were artificially injected into the pure signal through programming: high-amplitude outliers were injected at multiple discrete time points; a constant negative offset was injected within a continuous interval; and a linearly increasing drift was injected within another subsequent interval. After the above processing, the detection data containing mixed anomalies, as shown by the blue scatter points in the figure, was formed. This data will serve as the input for all subsequent algorithm steps.

[0100] Figure 3 (b) The process of the first stage of the framework's local adaptive preliminary screening is presented in detail. The black scatter points in the figure represent contaminated data to be detected. The algorithm first processes the contaminated data using a moving median filter with a preset window size, generating a baseline that adaptively fits the local nonlinear trend of the data. Subsequently, by calculating the residual between the original data points and this local baseline, and applying a strict statistical threshold based on the absolute deviation of the median, the outliers marked with red crosses in the figure, representing the highest confidence levels, are identified. According to the program's execution log, this preliminary screening stage identified a total of 30 obvious outliers.

[0101] Figure 3 (c) describes the data filtering steps after the initial screening. The goal of this step is to generate a training sample with relatively few data anomalies for the subsequent model to learn normal patterns. In the figure, the outliers identified in the previous stage are considered missing data points. The algorithm uses linear interpolation to fill these missing positions with valid data points on both sides of each outlier, ultimately generating the smooth green data curve shown in the figure. This curve eliminates most of the anomaly interference in the original data, preserving the true trend of the data relatively completely, thus laying a solid foundation for the accurate training of the subsequent model.

[0102] Figure 3 (d) shows the final diagnostic results of the entire anomaly detection framework. In the figure, the final baseline learned and generated by the core model accurately reconstructs the normal dynamic pattern of the data; this curve is unaffected by any outliers. Based on this baseline, the framework performs final anomaly detection and classification on the original contaminated data. The diagnostic log shows that 28 anomalies were precisely detected in this stage and classified according to their characteristics. As shown in the figure, outliers are marked with orange crosses, stable deviations are marked with green squares, and progressively changing drifts are marked with yellow diamonds. The final diagnostic report confirms that the model successfully identified 5 outliers, 11 deviations, and 12 drift points, intuitively verifying the high accuracy and reliability of this invention in complex anomaly detection tasks.

[0103] Figure 3 (e) presents the performance evaluation results of the algorithm under another set of parameter configurations, aiming to analyze the impact of changes in model sensitivity on recognition accuracy. The three metrics are precision, recall, and F1 score. Precision represents the proportion of simulated data anomalies identified out of all data anomalies; recall represents the proportion of simulated anomalies identified; and the F1 score is a combined score of the two. In this experiment, the precision for outlier detection dropped sharply to 0.12, while the recall remained at 1.00. This result indicates that although the model can capture all true outliers, the cost is a large number of false positives, leading to a corresponding drop in its F1 score to a low level of 0.21. Meanwhile, the model maintains perfect robustness in recognizing biased anomalies, with all metrics at 1.00. The performance in detecting drifting data anomalies is consistent with the previous experiment, with a precision of 1.00 and a recall of 0.75. However, due to the severe deterioration in outlier precision, the model's overall precision and F1 score dropped significantly to 0.53 and 0.66, respectively. This highlights that in anomaly detection tasks, excessively pursuing the recall rate of a certain type of anomaly may have a significant negative effect on the overall reliability and practicality of the model. Finding a balance among various indicators is crucial.

[0104] In order to analyze in depth Figure 3 The underlying reason for the sharp drop in accuracy of outliers in (e) Figure 3 (f) The classification error of data anomalies was quantified using a confusion matrix. The matrix data shows that the core problem lies in the large-scale misclassification of normal samples: a total of 23 normal samples were incorrectly classified as outliers by the model. This number is significantly higher than in the previous experiment and is the direct cause of the outlier precision rate of only 11.5%. This reflects that under the current model parameter settings, the detection threshold has become too strict for normal data fluctuations, thus misclassifying a large amount of noise as anomaly signals. Furthermore, the matrix also shows that the number of missed detections of true drift samples has increased from 3 to 4. In summary, this confusion matrix not only explains the drastic changes in macroscopic performance indicators from a data perspective but also clearly points out that the main challenge of the current algorithm lies in how to set an optimal discrimination threshold to effectively control false positive responses to background noise while ensuring high recall, thereby achieving more balanced and reliable detection performance.

[0105] Figure 4(a) illustrates the acceleration data preparation stage used in this experiment. The gray scatter points in the figure represent the raw, clean acceleration signal used as a reference. To simulate complex situations that may occur during dynamic monitoring, three typical data anomalies were artificially injected into this clean signal through programming: high-amplitude outliers were injected at multiple discrete time points; a constant negative offset anomaly was injected within a continuous interval; and a linearly increasing drift anomaly was injected within another subsequent interval. After the above processing, the acceleration data to be detected, containing mixed anomalies, is formed as shown by the blue scatter points in the figure. This data will serve as the input for all subsequent algorithm steps.

[0106] Figure 4 (b) details the first stage of the framework's local adaptive preliminary screening process applied to acceleration data. The black scatter points in the figure represent contaminated data to be detected. The algorithm first processes the contaminated data using a moving median filter with a preset window size, generating a baseline that adaptively fits the local nonlinear trend of the acceleration data. Subsequently, by calculating the residuals between the original data points and this local baseline, and applying a rigorous statistical threshold based on the absolute deviation of the median, the outliers marked with red crosses in the figure, representing the highest confidence levels, are identified. According to the program's execution log, this preliminary screening stage identified a total of 30 obvious outliers.

[0107] Figure 4 (c) describes the data filtering steps after the initial screening. The goal of this step is to generate a training sample with relatively few data anomalies for the subsequent model to learn normal acceleration patterns. In the figure, the outliers identified in the previous stage are considered missing data points. The algorithm uses linear interpolation to fill these missing positions with valid data points on both sides of each outlier, ultimately generating the smooth green data curve shown in the figure. This curve eliminates most of the abnormal interference in the original acceleration data, preserving the true trend of the data relatively completely, thus laying a solid foundation for the accurate training of the subsequent model.

[0108] Figure 4(d) shows the final diagnostic results of the entire anomaly identification framework on acceleration data. In the figure, the final baseline learned and generated by the core model accurately reconstructs the normal dynamic pattern of the data; this curve is unaffected by any outliers. Based on this baseline, the framework performs final anomaly detection and classification on the original contaminated data. The diagnostic log shows that 27 anomalies were precisely detected in this stage and classified according to their characteristics. As shown in the figure, outliers are marked with orange crosses, stable deviations are marked with green squares, and progressively changing drift anomalies are marked with yellow diamonds. The final diagnostic report confirms that the model successfully identified 5 outliers, 11 deviations, and 11 drift points, intuitively verifying that this invention also possesses high accuracy and high reliability for dynamic signals such as acceleration.

[0109] Figure 4 (e) The performance metrics of the algorithm after adjusting key parameters were quantitatively evaluated using bar charts. The three metrics were precision, recall, and F1 score. Precision represents the proportion of simulated data anomalies identified out of all data anomalies; recall represents the proportion of simulated anomalies identified; and the F1 score is a combined score of the two. The evaluation results show that the parameter adjustments significantly impacted the recognition performance of different types of anomalies. The most significant change was in outlier detection. While the recall remained at a perfect 1.00, ensuring the model could identify all true outliers, the precision plummeted to 0.30. This reveals that the model introduced a large number of false positives when identifying outliers, incorrectly labeling many normal points as outliers. In contrast, for biased anomalies, the algorithm remained very robust, with both precision and recall at 1.00, demonstrating strong robustness to this pattern. The identification of drift-type anomalies maintained a precision of 1.00, ensuring the reliability of the detection results. However, the recall rate dropped to 0.75, indicating some missed detections. Overall, although the precision for outliers decreased significantly, the model's overall F1 score remained at a good level of 0.86, demonstrating the effectiveness of its overall framework. This also points out that after increasing model smoothness, balancing noise suppression with accurate outlier identification has become a new key to optimization.

[0110] For in-depth diagnosis Figure 4 The reasons for the changes in performance indicators in (e) Figure 4(f) The specific classification results of each type of data anomaly are presented in detail in the form of a confusion matrix. The off-diagonal elements of this matrix precisely reveal the sources of misclassification by the model, providing microscopic evidence for performance analysis. The root cause of the significant drop in outlier precision is clearly visible: a total of 7 non-outlier samples were incorrectly predicted as outliers, including 6 normal samples and 1 drift sample. This result indicates that after adjusting the parameters to make the SBL baseline smoother, the model became overly sensitive to some normal data fluctuations, thus producing misclassifications. At the same time, the matrix also quantifies the missed detection of drift data anomalies: out of a total of 16 true drift samples, 3 were misclassified as normal, 1 was misclassified as an outlier, and 12 were ultimately correctly identified, which corresponds perfectly to the recall rate of 0.75 for this type of data anomaly in the bar chart. This confusion matrix not only validates the macroscopic indicators but also provides a detailed perspective on error analysis, indicating that the algorithm has reached a new balance between its ability to distinguish normal data noise and its ability to capture drift features under the current parameters, providing a concrete basis for further optimization of the model.

[0111] Figure 5 (a) illustrates the displacement data preparation stage used in this experiment. The gray scatter points in the figure represent the original pure displacement signal used as a reference. To simulate complex situations that may occur in the real world, three typical data anomalies were artificially injected into the pure signal through programming: high-amplitude outliers were injected at multiple discrete time points; a constant negative offset was injected within a continuous interval; and a linearly increasing drift was injected within another subsequent interval. After the above processing, the displacement data to be detected, containing mixed anomalies, is formed as shown by the blue scatter points in the figure. This data will serve as the input for all subsequent algorithm steps.

[0112] Figure 5 (b) details the first stage of the framework's local adaptive preliminary screening process applied to displacement data. The black scatter points in the figure represent contaminated data to be detected. The algorithm first processes the contaminated data using a moving median filter with a preset window size, generating a baseline that adaptively fits the local nonlinear trend of the displacement data. Subsequently, by calculating the residuals between the original data points and this local baseline, and applying a rigorous statistical threshold based on the absolute deviation of the median, the outliers marked with red crosses in the figure, representing the highest confidence levels, are identified. According to the program's execution log, this preliminary screening stage identified a total of 40 obvious outliers.

[0113] Figure 5(c) describes the data filtering steps after the initial screening. The goal of this step is to generate a training sample with relatively few data anomalies for the subsequent model to learn normal displacement patterns. In the figure, the outliers identified in the previous stage are considered missing data points. The algorithm uses linear interpolation to fill these missing positions with valid data points on both sides of each outlier, ultimately generating the smooth green data curve shown in the figure. This curve eliminates most of the abnormal interference in the original displacement data, preserving the true trend of the data relatively completely, thus laying a solid foundation for the accurate training of the subsequent model.

[0114] Figure 5 (d) shows the final diagnostic results of the entire anomaly identification framework on displacement data. In the figure, the final baseline learned and generated by the core model accurately reconstructs the normal dynamic pattern of the data; this curve is unaffected by any outliers. Based on this baseline, the framework performs final anomaly detection and classification on the original contaminated data. The diagnostic log shows that 33 outliers were precisely detected in this stage and classified according to their characteristics. As shown in the figure, outliers are marked with orange crosses, stable deviations are marked with green squares, and gradually changing drifts are marked with yellow diamonds. The final diagnostic report confirms that the model successfully identified 6 outliers, 11 deviations, and 14 drift points, intuitively verifying that this invention also has high accuracy and high reliability for displacement monitoring data.

[0115] Figure 5 (e) presents the detailed performance evaluation results of the proposed algorithm in handling mixed data anomalies. This figure uses a bar chart to visually compare the precision, recall, and F1 score of the model in identifying outliers, biases, drift, and overall anomalies. Precision represents the proportion of simulated data anomalies identified out of all data anomalies; recall represents the proportion of simulated anomalies identified; and the F1 score is a combined score of the two. As can be seen from the figure, the model exhibits perfect identification ability for bias anomalies, with all indicators reaching 1.00, demonstrating its extremely high sensitivity and accuracy for this type of patterned anomaly. Regarding outlier detection, the model achieves a recall of 1.00, indicating that all true outliers are successfully detected, ensuring the completeness of the detection; however, its precision is 0.75, revealing a small number of false positive samples that misclassify normal points as outliers. For drift anomalies, the model achieves a precision of 1.00, ensuring the reliability of the detection results, but the recall is 0.88, indicating that some drift samples were not identified. In terms of overall performance, the model achieved a high F1 score of 0.95, which fully demonstrates that the algorithm has high accuracy and robustness in complex monitoring scenarios and can effectively cope with different types of data anomalies.

[0116] To further analyze the classification performance and error sources of the model, Figure 5 (f) A confusion matrix for data anomaly classification was plotted. The vertical axis of the matrix represents the true data anomaly type of the samples, and the horizontal axis represents the model's predicted type. The values ​​in the matrix show the distribution of samples among different anomaly types. The results clearly show that the model correctly located the vast majority of samples: the diagonal data shows that 789 normal points, 3 outliers, 11 biased points, and 14 drift points were all successfully classified. The off-diagonal cells precisely reveal the details of the model's misclassification. The main classification errors are reflected in two aspects: first, one normal sample was misclassified as an outlier, which is the direct cause of the decrease in outlier precision; second, two true drift samples were missed and classified as normal, which explains why the recall rate of drift data anomalies did not reach perfection. This matrix, by quantifying the classification details of each data anomaly type, not only re-verifies the overall efficiency and reliability of the algorithm, but also provides clear data support and improvement directions for subsequent algorithm optimization for specific misclassification scenarios.

[0117] In summary, this method demonstrates superior performance across all test scenarios. Specifically, for strain data, the model successfully identified 5 outliers, 11 deviations, and 12 drifts; for acceleration data, it identified 5 outliers, 11 deviations, and 11 drifts; and for displacement data, it identified 6 outliers, 11 deviations, and 14 drifts. These results clearly demonstrate that, regardless of the type of monitoring signal, this invention can accurately detect and distinguish different types of data anomalies. This achievement strongly validates the robustness and high accuracy of the proposed method. Its success can be attributed to a two-stage strategy combining local adaptive preliminary screening and SBL precise diagnosis, ensuring accurate modeling of normal behavior and sensitive inference of abnormal states. In conclusion, this invention is a diagnostic method capable of accurately identifying multiple types of data anomalies in various types of monitoring data.

[0118] Figure 6 , Figure 7 The regression analysis of monitoring data dominated by outliers is presented using the data correction method based on the Bernoulli sparse Bayesian model and the two-stage elimination and reconstruction method.

[0119] like Figure 6 (a) shows the regression model trained on this complete dataset after replacing outliers in the original data with corrected values. Its performance metrics are RMSE = 1.9320 and MAE = 1.4866. =0.9336; and Figure 7(a) shows the final high-precision model trained on training data with relatively few outliers after identifying and removing outliers. Its performance metrics are RMSE=1.8901 and MAE=1.4866. =0.9336;

[0120] like Figure 6 (b), (c), (d) and Figure 7 As shown in (b), (c), and (d), the data correction method and the two-stage elimination and reconstruction method all exhibit similar health status in terms of model convergence analysis, weight distribution, and parameter determinism. The LML curves all converge quickly and stably. Figure 6 The inferred noise standard deviation is 2.44943. Figure 7 The inference noise standard deviation is 2.01766, the weight distribution is sparse and moderate, and the Gamma parameter exhibits high determinism. This indicates that neither method suffers from model collapse or overconfidence in this outlier scenario.

[0121] In summary, by comparing the key performance indicators of the data correction method and the two-stage elimination and reconstruction method, the root mean square error and mean absolute error of the two-stage elimination method are both lower than those of the data correction method. Higher. This clearly demonstrates that when dealing with monitoring data dominated by outliers, the two-stage elimination and reconstruction method, by effectively isolating the interference of data anomalies, can obtain a final regression model with higher accuracy and better performance than the data correction method.

[0122] Figure 8 , Figure 9 This demonstrates regression analysis of monitoring data with long segments of data anomalies, including bias and offset values, using a two-stage elimination and reconstruction method and a data correction method based on a Bernoulli sparse Bayesian model.

[0123] like Figure 8 As shown in (a), the Bernoulli sparse Bayesian model can track the baseline in continuous data regions, but exhibits severe non-physical oscillations in regions of missing data caused by the removal of anomalies in continuous data. In contrast, Figure 9 The model in (a) provides a smooth, stable, and accurate regression curve that follows the main trend of the data. The original contamination data is effectively identified and corrected to the SBL baseline, and the final model is trained on this complete and continuous dataset. This method fundamentally avoids the problem of numerical oscillations by replacing rather than deleting outliers, and the regression results have clear physical meaning.

[0124] like Figure 8As shown in (b), the Log Marginal Likelihood (LML) curve appears as a horizontal line with an abnormal numerical scale, indicating that the optimization process of the Bernoulli sparse Bayesian model stalls or encounters numerical instability in the initial stage. This confirms the ill-conditioned solution space caused by missing data, which prevents the optimization algorithm from converging effectively. Figure 9 The LML curve in (b) also exhibits a rapidly converging horizontal trajectory, but the essential difference lies in the validity of the final result. In the two-stage elimination and reconstruction method, rapid convergence leads to an ill-conditioned solution; while in the data correction method, it leads to a stable and reasonable model. This indicates that the data correction method provides a well-conditioned dataset, enabling the optimization process to successfully converge to an effective global optimum.

[0125] like Figure 8 As shown in (c), although the weight distribution of the monitoring data reflects sparsity, a few data points exhibit extremely large positive and negative weight spikes. This is a leverage effect forced upon the model to overcome missing data segments, and is a clear signal of model instability. In comparison, Figure 9 The weight distribution in (c) is also sparse, but the key is that the absolute values ​​of all non-zero weights remain on a very moderate order of magnitude. This moderate weight distribution contrasts sharply with the extreme cases of the elimination method and is a core indicator of the model's stability, proving that this is a well-conditional regression model.

[0126] like Figure 8 As shown in (d), the Gamma value of all monitoring data correlation vectors is 1.0, but this reflects an overconfidence. The Bernoulli sparse Bayesian model exhibits high determinism in extreme weight choices that lead to global collapse, revealing that in the overconfidence problem, simply relying on maximizing model evidence may converge to an invalid solution. Figure 9 In (d), the Gamma parameter for all relevant vectors is also 1.0, which is a positive sign. It indicates that the model has a high degree of determinism for each of the selected basis functions used to construct the smooth curve, and the data itself provides sufficient evidence for these choices, representing the performance of a healthy, well-converged model.

[0127] In summary, from four aspects—the smoothness of the regression curve, the effectiveness of model convergence, the moderation of the weight distribution, and the health of the parameter determination— Figure 8 All subgraphs consistently demonstrate the stability and robustness of the data correction method in handling long-segment data anomalies, and... Figure 9 The two-stage elimination and reconstruction method exhibits a stark contrast in its ill-conditioning and instability. This demonstrates that the data correction method outperforms the two-stage elimination and reconstruction method when dealing with monitoring data containing long periods of data anomalies such as bias and drift values.

Claims

1. A method for multi-type data anomaly diagnosis and processing in structural health monitoring based on Bernoulli sparse Bayesian model, characterized in that, It provides a unified framework for data anomaly diagnosis and regression modeling data anomaly processing based on Bernoulli sparse Bayesian, and reconstructs the data anomaly diagnosis problem as a probability inference task, that is, the probability of each data point being normal or abnormal is inferred by Bernoulli sparse Bayesian model, to realize the identification and classification of anomaly, and finally the data anomaly is processed by two regression analysis strategies, the specific steps are as follows: A structural health monitoring (SHM) system is installed on the engineering structure, the SHM system includes a plurality of different types of monitoring sensors arranged at different positions of the engineering structure, and a computer for recording data of the monitoring sensors; The monitoring data in the healthy state is added with analog quantity, which is used to test the identification accuracy of the model for irregular errors and to verify the accuracy and practicability of the model; Firstly, the model is used to preliminarily screen the monitoring data, a moving median filter is applied to the monitoring data to generate a local reference line, which is used to robustly fit the local nonlinear trend of the monitoring data, then the residual between the monitoring data and the local reference line is calculated and applied to a strict statistical threshold based on the median absolute deviation, and finally the preliminary screening of the data anomaly points is completed, providing a data sequence with relatively less data anomaly for subsequent fine modeling; The data sequence after preliminary screening is used to infer whether the data points in the original monitoring data are abnormal, and the data anomaly is classified according to different residual patterns of the abnormal data, which can be divided into three types of data anomaly: outlier, deviation value and drift value; After completing the anomaly identification, the theoretical model proposes two regression analysis data anomaly processing strategies and compares them: the first one is two-stage elimination reconstruction method, which is suitable for abnormal monitoring data dominated by outliers, first a preliminary sparse Bayesian (SBL) model is trained on the original data containing anomalies to identify anomalies, then all samples judged as abnormal are completely removed from the data set, and finally a more accurate SBL model is trained on the data set from which part of the abnormal data has been removed, to effectively isolate the negative impact of data anomaly on the accuracy of the final model; The second one is data correction method, which is suitable for monitoring data with long segment data anomaly such as deviation value and drift value, through the good data filling characteristics of sparse Bayesian algorithm, the filling value conforming to the nonlinear trend of the data can be generated to avoid the problem caused by the discontinuity of the data sequence, so as to improve the accuracy of the regression model; The data processed by regression analysis can be considered as reliable real monitoring data.

2. The method according to claim 1, wherein, The Bernoulli sparse Bayesian model adopts SBL theoretical model and Bernoulli likelihood model for data anomaly value identification and analysis, the basic form of SBL theoretical model is: (1) where is a basis function vector, typically employing Gaussian kernel functions , maps the input data into a high-dimensional feature space. is a corresponding weight vector. is noise that is subject to a zero-mean Gaussian distribution, i.e. ; Therefore, the likelihood function of the target value is: (2) wherein is a design matrix whose elements ; The core of SBL is to put a weight A hierarchical prior is introduced. Specifically, for each weight a separate, zero-mean Gaussian prior is assigned: (3) where is the control weight of the variance of the hyperparameters by putting the hyperparameters and the noise precision under a hyper-prior of Gamma distribution and iteratively optimizing these hyperparameters using the Type II maximum likelihood method (Evidence Maximization), SBL can automatically push most of the irrelevant basis functions to infinity, so that their weights have their posterior distributions concentrated at zero, achieving the sparsity of the model; After the training is completed, for a new input point , SBL can not only give the predicted mean , but also the predicted variance , constituting a complete predictive probability distribution: (4) (5) (6) and are the posterior mean and covariance of the weights respectively, and the prediction variance constitutes the confidence interval of the prediction value, which is the key basis for subsequent data anomaly probability identification; The basic form of Bernoulli likelihood model is: Assume each data point there is a binary latent variable where denotes that the point is an outlier, denotes that it is normal; The distribution of the latent variable is subject to Bernoulli distribution: (7) wherein is the prior probability that a data point is anomalous; In the present invention, the latent variables are not directly estimated but their states are inferred a posteriori from the predictions of the SBL model, in particular, the predicted mean and the predicted standard deviation of the SBL model for each input can be used to construct a confidence interval , if the observed value falls outside this interval, it is judged as an outlier, this decision process can be formalized as a decision on the latent variable ​ (8) Compared with the traditional fixed threshold method, the Bernoulli classification method based on the SBL probability output is more objective and has good robustness, and can adaptively cope with dynamic data changes.

3. The method according to claim 2, wherein, The overall idea of the method is to first screen and then accurately identify through probability inference, aiming to achieve high-precision and high-robustness identification of data outliers: The method first uses a moving median method to construct a local adaptive baseline, which has good robustness and can effectively track the local center trend of the data without being disturbed by extreme outliers. By calculating the deviation between the original data points and the moving median baseline, the most significant data outliers can be identified and temporarily removed. The main goal of this stage is to provide a data sequence with relatively few data outliers for subsequent fine modeling; After preliminary screening, the next step is to accurately infer whether the data is a data outlier; The invention firstly defines the problem from the perspective of probability theory, and the core assumption is that for each data point in the time series there is an implicit state variable behind it that cannot be directly observed : When the point is normal, it is represented by a dot; When the point is representative of an outlier; This binary state variable can be considered as sampling from a Bernoulli distribution: (9) wherein is the prior probability of a point being abnormal data, and the goal of the entire data anomaly detection is to infer the most likely state behind the observation data ; The model takes the data after the first stage of preliminary screening as a representative sample of the normal state, and trains this data through the SBL regression model, which essentially builds a conditional probability model of the normal state With the powerful nonlinear fitting capability of SBL, the model can learn the complex trends hidden behind the data After training is complete, the final baseline generated by the model can be considered as the expected value of the normal probability distribution ; After having the precise normal model, the state inference can be performed for each point in the original contaminated data. For a point to be tested , the residual between it and the corresponding normal expectation is calculated : (10) The residual can be viewed as a key statistic, in that if a point belongs to the normal state ( ), then its residual should be small, consistent with the noise distribution predicted by the model; conversely, if the residual is significantly large, then the probability of the point being generated by the normal model is extremely low, and thus there is good reason to believe that its state is abnormal ( ). Therefore, the threshold judgment method used here can be regarded as an efficient approximation of the Bayesian decision, which uses the confidence interval predicted by the SBL model as a criterion to determine whether a point is normal or abnormal, thereby completing the inference of the implicit Bernoulli state of the data point. Therefore, the threshold judgment method used here can be regarded as an efficient approximation of the Bayesian decision, which uses the confidence interval predicted by the SBL model as a criterion to determine whether a point is normal or abnormal, thereby completing the inference of the implicit Bernoulli state of the data point. After completing the Bernoulli state inference, the application further subdivides all points marked as by analyzing the residual stability of the continuous data anomaly section, and divides the data anomaly into three types of outliers, bias values and drift values, wherein: (1) Outliers are isolated data points that appear instantaneously and deviate significantly from the normal data stream, usually caused by external electromagnetic interference, sensor transient failure or real rare physical events; (2) Bias values refer to the systematic and fixed deviation between sensor measurements and true values, usually caused by incorrect physical installation or uncalibrated equipment; (3) Drift values refer to the slow and one-way cumulative changes in sensor measurement errors over time, which are usually caused by sensor aging or continuous surface contamination.

4. The method according to claim 3, wherein, Through the Bernoulli sparse Bayesian model, regression analysis is performed on the data of each type of outlier identified: Based on the above abnormal identification method, the invention constructs and compares the following two data abnormality processing strategies suitable for different situations, namely the two-stage rejection reconstruction method and the data correction method: The two-stage rejection reconstruction method is a processing strategy that can effectively isolate the negative impact of data outliers and pursue the highest prediction accuracy of the final model. This method first trains a preliminary SBL model on the original monitoring data containing data outliers. Then, all data anomaly samples are located as much as possible according to the identification framework, and they are completely removed from the data set to form a training subset with relatively few data anomalies. On this training subset, a higher-precision SBL model is trained. Finally, the high-precision model is used to predict all original time points to reconstruct a smooth and continuous high-precision data curve. This method is suitable for data anomaly monitoring data dominated by outliers; The data correction method can generate filling values that conform to the nonlinear trend of the data through the good data filling characteristics of the sparse Bayesian algorithm. This method replaces rather than deletes data outliers, avoiding numerical oscillation problems caused by discontinuity in the data sequence, thereby improving the accuracy of the regression model. This method is suitable for monitoring data with long segments of data anomalies such as bias values and drift values; The data processed through regression analysis can be considered as reliable real monitoring data.

5. The memory is configured to store program codes, and the processor is configured to invoke the program codes to execute the method of any one of claims 1-4.

6. A computer-readable storage medium, characterized in that, The computer program stored in the computer readable storage medium is adapted to be loaded and executed by the processor to execute the method of any one of claims 1-4.