Industrial fault detection method and system based on dynamic drift perception and diffusion enhancement
By constructing a DDA-DE model and combining dynamic drift sensing and diffusion enhancement methods, the problem of distinguishing between data drift and anomalies in industrial fault detection was solved, achieving efficient and adaptive fault detection, reducing false alarm rate, and improving detection performance.
Patent Information
- Application Number
- CN202511535580.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-27
AI Technical Summary
Existing industrial fault detection methods struggle to effectively distinguish between data drift and anomalies, exhibiting high false alarm rates. Furthermore, their models lack adaptability and cannot adjust thresholds in real time, leading to decreased detection performance.
A method based on dynamic drift sensing and diffusion enhancement is adopted. An unsupervised fault detection DDA-DE model is constructed. Dynamic drift sensing DDA is used to process the data stream to establish a statistical distribution baseline and drift threshold. Diffusion enhancement anomaly detection DE is combined to determine the initial model parameters and anomaly threshold, thereby realizing adaptive optimization of the model and fault detection.
It improves the model's accuracy in detecting data drift and anomalies, reduces the false alarm rate, enhances the model's robustness and adaptability under complex working conditions, and significantly improves detection performance.
Smart Images

Figure CN120995184A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of fault detection, in particular to an industrial fault detection method and system based on dynamic drift perception and diffusion enhancement. BACKGROUND
[0002] Industrial data stream anomaly detection technology has become the core support for predictive maintenance and process optimization in industrial control systems by analyzing equipment time series data (such as vibration spectrum, temperature fluctuation, etc.) in real time. However, the multi-source data drift phenomenon (such as sensor calibration offset, feature distribution gradual change caused by equipment aging, etc.) commonly existing in industrial sites seriously weakens the reliability of the detection model. Figure 1 The time series and statistical feature display of the FIT101 (used to display the water inflow state) sensor in the SWaT data set is shown, and it can be found that the statistical features of the training set and the test set change greatly, which seriously affects the model monitoring effect. When the production line data stream has implicit drift, the anomaly detection algorithm often has false positives, causing the operation and maintenance personnel to fall into a vicious cycle of frequent false alarms and ineffective maintenance.
[0003] To solve the data drift problem, related methods can be roughly divided into two categories: one is applied in the data preprocessing stage, and the other is implemented in the modeling stage. In the modeling stage, the D3R model is quite representative. To deal with the data drift problem in time series anomaly monitoring, they proposed an anomaly monitoring method based on sequence decomposition. This method uses a data-time hybrid attention mechanism to dynamically decompose long-period multivariate time series, thereby eliminating the influence of trend items on the monitoring results. In the data preprocessing stage, the simple adaptive strategy is a typical representative. This strategy uses statistical estimation methods to remove the trend estimates of time series data, and then carries out anomaly monitoring, and uses the detrended sequence to update the model parameters.
[0004] However, the above work still faces the following challenges in solving the drift problem existing in the industrial scene: First, the decoupling of drift and anomaly is difficult to achieve. Existing work often uses decomposition to achieve the decoupling of drift and anomaly. However, the model needs to process multiple sub-sequences generated by decomposition, resulting in a significant increase in computational burden; ideally, the residual after decomposition should be white noise, but in practice it often contains unextracted patterns (such as unidentified short-period fluctuations); when there is high-order coupling between seasonal and trend components, simple decomposition will lose interaction information. In addition, since drift and anomaly differ significantly in terms of root cause and manifestation, relying solely on statistical data features to determine real-time changes in drift and anomaly is far from enough.
[0005] In industrial scenarios, drift usually originates from the progressive changes of system or environment, while anomaly is mostly caused by sudden equipment failure or external attack. In terms of performance, drift is embodied as the shift of data distribution or the change of time series pattern, while anomaly is embodied as the mutation of local data points, non-continuous fluctuation or business logic conflict. Due to the significant differences in the root causes and forms of performance, if the analysis of them is performed in a single process, the model will be difficult to effectively distinguish between anomaly and drift. In addition, it is far from enough to rely only on statistical data features to judge real-time changes of drift and anomaly.
[0006] Secondly, the model lacks adaptability. The existing anomaly detection framework lacks a continuous learning mechanism and is difficult to capture the dynamic evolution rule of data distribution in the concept drift scenario. Specifically, when facing seasonal fluctuations or long-term system evolution, the static model parameters cannot be iteratively updated through domain knowledge fusion, resulting in gradual degradation of detection performance. This limitation is due to the lack of an adaptive parameter optimization mechanism that synchronizes with the time-varying industrial process.
[0007] Thirdly, fixed thresholds cause high false alarm rates. Specifically, due to the influence of multiple factors such as equipment operating state, environmental periodic fluctuations, etc., the morphological features of data drift exhibit significant time-varying and nonlinear characteristics. When the system drifts, the anomaly index of the monitoring model will be lifted due to the distribution shift. At this time, using a static threshold to determine cannot effectively distinguish between drift and real anomaly events, and will also significantly increase the false alarm probability under normal working conditions. This phenomenon highlights the need for a dynamic threshold adjustment strategy or adaptive mechanism design to improve the robustness of the system to complex working conditions. SUMMARY
[0008] In order to solve the problem and deficiency that the coupling relationship between dynamic perception of data distribution change and anomaly mode is difficult to perceive in the prior art, the present application provides an industrial fault detection method and system based on dynamic drift perception and diffusion enhancement. Through real-time distance monitoring and adaptive model incremental learning mechanism, the co-detection of data drift and anomaly events is realized.
[0009] In the first aspect, the present application provides an industrial fault detection method based on dynamic drift perception and diffusion enhancement, which adopts the following technical scheme: An industrial fault detection method based on dynamic drift perception and diffusion enhancement, comprising: Obtaining industrial flow data; Constructing an unsupervised fault detection DDA-DE model, wherein the industrial data stream is processed by a dynamic drift perception DDA to establish a statistical distribution baseline and a drift threshold; an initial model parameter and an anomaly threshold are determined by a diffusion enhancement anomaly detection DE; and the robustness of the unsupervised fault detection model to concept drift is enhanced based on a diffusion strategy; Model optimization is performed on the constructed unsupervised fault detection DDA-DE model; Fault detection is performed by using the optimized model.
[0010] Further, the constructed unsupervised fault detection DDA-DE model comprises a double-branch detection architecture based on the essential difference between concept drift (statistical feature change) and abnormal event (feature correlation destruction), which can preserve the feature correlation within the data. The drift detection module adopts a covariate offset quantization model based on a sliding window, introduces an industrial enhanced Mahalanobis distance metric, and dynamically tracks the feature space distribution change in real time. The fault detection branch integrates a diffusion enhancement strategy, enhances the distribution offset robustness through a multi-scale feature propagation mechanism, and suppresses noise interference with the help of information bottleneck constraint. Based on the detection results of the two branches, the model parameters are dynamically calibrated, and the fault detection threshold is optimized through an adaptive threshold adjustment mechanism.
[0011] Further, the dynamic drift perception DDA processing of industrial data streams to establish a statistical distribution baseline and a drift threshold comprises extracting statistical descriptors from historical data by using DDA, wherein the baseline distribution parameters including a mean vector of feature vectors and a covariance matrix are calculated to realize real-time distance calculation, so as to continuously update the drift threshold; and the mean and covariance matrix are updated gradually through incremental calculation, and the incremental calculation process is as follows: , , wherein, is the mean value at the current moment, is the mean value at the previous moment, is a new data point at the current moment, is the total number of data points at the current moment, is the covariance matrix at the current moment, is the covariance matrix at the previous moment, and for a feature vector , the formula for calculating the distance from the feature vector to the historical distribution is: , wherein , represents Hadamard product, and the Hadamard product is normalized to eliminate the scale and unit difference between heterogeneous sensor features by being included in the Mahalanobis distance; The formula for calculating , is the temperature coefficient, is the variance of the ith sensor; finally, threshold determination is performed in the training phase, in which the decision boundary is dynamically adjusted by real-time updating, and the updating is calculated based on the mean and standard deviation of the historical distance distribution, and the calculation formula is: , wherein, is the drift threshold value at the current moment, is the mean of the historical distance at the current moment, is the standard deviation of the historical distance at the current moment, is a constant.
[0012] Further, the determination of the initial model parameters and the anomaly threshold by the diffusion enhanced anomaly detection DE includes dynamic comparative analysis of newly arrived time series based on the statistical features obtained in the training phase, wherein first, the stable statistical data and threshold value are calculated based on the historical normal data, which are used to depict the stable mode of the system in the healthy state, when a new time window arrives, the industrial enhanced Mahalanobis distance between the window and the historical distribution is calculated through the same distance calculation process, and the distance calculation process has tolerance to fault fluctuations and slight noise by introducing weighted and robust adjustment for the characteristics of the industrial system; after the distance value is calculated, the distance value is compared with the set threshold value, if the distance value exceeds the set threshold value, i.e. , it indicates that the deviation of the current time window from the historical normal distribution is significant, indicating that the system behavior may have statistical drift or state mutation, if , it indicates that the features of the current window are still within the range of the historical distribution.
[0013] Further, the robustness enhancement of the unsupervised fault detection model based on the diffusion strategy against concept drift includes iterative updating of the model parameters of DE through a sliding window input mechanism, derivation of a fault threshold value based on an error score, and calculation of reconstruction error, wherein the time domain noise is injected through a diffusion process, the features are extracted from the noise sequence, and the reconstruction error is calculated, wherein the time domain noise is injected through a diffusion process, the features are extracted from the noise sequence, and the reconstruction error is calculated, wherein the input of the diffusion module is , wherein the original vector at a certain moment is represented as , the vector contaminated by M-step Gaussian noise can be represented as , and the contamination result obtained in the mth step is only related to , and the data after m-step noise contamination is: , wherein is a standard Gaussian noise, . is a noise scheduling parameter, which decreases with time step m; , wherein is the noise rate, which is incremented by time step, The calculation process is: , where and are preset minimum and maximum noise rates, M is the total number of time steps, and the output of the diffusion strategy is finally obtained: , where, , and ; the variable value after the noise added by the diffusion module is position encoded, denoted as: , where, denotes convolution, denotes the position encoded variable value, denotes relative position encoding, denoted as: , denotes the position of the time window at the moment, denotes the dimension of the vector embedding, denotes even dimensions in odd dimensions in .
[0014] Further, the feature extraction from the noise sequence includes inputting the sequence after the word embedding into the feature extraction structure for extracting complex spatiotemporal correlations in the data, wherein the feature extraction module is composed of L layers of blocks, wherein the input of the block module of the lth layer is , with a time length and dimensions, the initial input denotes the embedded noise sequence, and is obtained in the time sequence module, and the calculation process is represented as: , , where, denotes layer normalization, denotes a multi-head attention mechanism, denotes the time sequence hidden representation of the lth layer, denotes a forward feedback network, and is also obtained in the spatial module, denoted as: , , where denotes the spatial hidden representation of the l-th layer, denotes the transpose; finally, the (l+1)-th layer is obtained by the process: , , wherein denotes the hidden representation, denotes the concatenation operation.
[0015] Further, the reconstruction error is calculated, including adding a reconstruction module composed of two fully connected layers at the end of the model to project the output dimension back to the original input dimension, in order to achieve the unsupervised training goal, and the reconstruction process is expressed as: , wherein denotes the reconstruction result of , denotes the fully connected layer; then the reconstruction error between the output result of the model and the original input is calculated as the loss function value, and is back-propagated to the neural network, facilitating the model parameter update; at the same time, the reconstruction error is used as the loss function, facilitating the model to find the complex correlation between the time sequence and each sensor, and the calculation process is: , wherein denotes the Frobenius norm; and the reconstruction error score is generated for the sensor value of each timestamp, the difference between the data with faults and the internal distribution of the normal data is reflected in the reconstruction error score, and the instances with error scores lower than the threshold value are determined as normal, and those exceeding are marked as abnormal.
[0016] Further, the model optimization of the constructed unsupervised fault detection DDA-DE model includes setting the time step m of the noise added in the diffusion module of the DDA-DE to 2000, setting the preset minimum and maximum noise rates β min and β max to 0.2 and 1 respectively, setting the number of layers L in the feature extraction module to 4, setting the embedding dimension to 128, and setting the size of the sliding window is 64, the loss function adopts reconstruction loss, the model adopts small batch stochastic gradient descent strategy to update, each round of training traverses complete training set, the optimizer adopts ADAM, and the average loss of the training set and the verification set is calculated after each epoch is completed in the training process, so as to monitor the convergence of the model; in order to prevent overfitting, dropout is added between the transformer layers, and early termination mechanism is adopted when the verification set error does not improve for a plurality of consecutive rounds, that is, early termination mechanism is adopted; if it is found that the reconstruction error of the verification set does not decrease for 5 consecutive epochs in the training, the learning rate decay mechanism is triggered, and the learning rate is automatically reduced by half to refine the search; the whole training process is iterated continuously, and the model gradually learns the normal mode of the time series; when the parameters of the network converge and the verification set error is the lowest, the model weight is saved.
[0017] Further, the fault detection using the optimized model includes judging whether the data point is faulty according to the reconstruction error score of the sensor value at each time point output by the model , wherein , if the reconstruction error score is less than a preset threshold , and it is determined that no drift occurs, it is determined that the data at the time point is normal; if the reconstruction error score is greater than or equal to the threshold , and it is determined that no drift occurs, it is determined to be faulty; if the reconstruction error score is less than the threshold , and it is determined that drift occurs, it is determined that the data at the time point is normal, and the model is incrementally updated; if the reconstruction error score is greater than or equal to the threshold , and it is determined that drift occurs, the threshold is adjusted and re-evaluated; wherein the judgment method of the preliminary evaluation of whether a fault occurs is: ; If it is determined that drift occurs, the abnormal threshold is adjusted and re-evaluated, and the adjustment method is: , ; Finally, if no fault is detected and no concept drift occurs, it is not recorded; if an abnormality is detected but no concept drift occurs, the fault is recorded and monitoring continues; if no abnormality is detected but concept drift occurs, the DE is incrementally trained to adapt to the change of data distribution; if both an abnormality and concept drift are detected, the threshold is adjusted and re-evaluated.
[0018] In a second aspect, an industrial fault detection system based on dynamic drift perception and diffusion enhancement includes: A data acquisition module configured to acquire industrial flow data; The model construction module is configured to construct an unsupervised fault detection DDA-DE model, wherein dynamic shift perception DDA is used to process industrial data streams to establish a statistical distribution baseline and a shift threshold; diffusion enhanced anomaly detection DE is used to determine initial model parameters and an anomaly threshold; and a diffusion strategy is used to enhance the robustness of the unsupervised fault detection model against concept shift. The model optimization module is configured to optimize the constructed unsupervised fault detection DDA-DE model. The fault detection module is configured to perform fault detection using the optimized model.
[0019] In a third aspect, the present application provides a computer readable storage medium, wherein a plurality of instructions are stored, the instructions being adapted to be loaded by a processor of a terminal device and to execute the industrial fault detection method based on dynamic shift perception and diffusion enhancement.
[0020] In a fourth aspect, the present application provides a terminal device, comprising a processor and a computer readable storage medium, the processor being configured to implement instructions, and the computer readable storage medium being configured to store a plurality of instructions, the instructions being adapted to be loaded by the processor and to execute the industrial fault detection method based on dynamic shift perception and diffusion enhancement.
[0021] In summary, the present application has the following beneficial technical effects: (1) In solving the problem of data non-stationarity, the dynamic decomposition method D3R based on diffusion reconstruction simply disassembles data into trend items and seasonal items, which causes performance degradation and ignores the linkage between the decomposition items. The industrial enhanced Mahalanobis distance real-time shift perception algorithm proposed in the present application can more efficiently detect concept shift phenomena. The parallel design of the shift perception algorithm and the fault detection classifier realizes the cooperative detection of data shift and abnormal events.
[0022] (2) Most existing time series fault detection methods use fixed thresholds to judge faults, which leads to high false alarm rates when data drifts occur due to static parameters or untimely adjustments. The threshold adjustment strategy and model parameter real-time updating method used by the present application through the joint action of the shift perception algorithm and the fault detection classifier can improve the model's self-adaptive ability while effectively reducing the false alarm rate.
[0023] (3) In terms of model detection performance, the present application considers solving the non-stationarity problem from three aspects: decoupling of anomalies and shifts, model adaptive adjustment, and dynamic threshold adjustment, which ultimately makes the model detection effect better. Figure 4 As can be seen from the above, the detection model proposed in the present application is superior to other methods in terms of F1 value, PATE, and AUC, thereby proving that the present application has great advantages in detection performance. Attached Figure Description
[0024] Figure 1 This is a flowchart of the time-series fault detection method according to Embodiment 1 of the present invention; Figure 2 This is a framework diagram of the fault detection model; Figure 3 Dividing the dataset into different parts; Figure 4 The graph shows the drift and fault detection results for the SWaT dataset; Figure 5 The graph shows the drift and fault detection performance under continuous faults in the SWaT dataset. Figure 6 This is a three-layer timing diagram. Detailed Implementation
[0025] The present invention will be further described in detail below with reference to the accompanying drawings.
[0026] Example 1 Reference Figure 1 This embodiment proposes an industrial fault detection method based on dynamic drift sensing and diffusion enhancement, hereinafter referred to as DDA-DE. The specific processing flow includes the following steps: S1 Data Acquisition This step is the cornerstone of the entire process. Its core function is to systematically collect raw signals from industrial scenarios, providing a high-quality, interpretable data foundation for subsequent modeling. Specific content includes: target definition, industrial protocol access, data acquisition, and secure storage. Its core functions are: supporting real-time decision-making, ensuring data credibility, guaranteeing business continuity, and ultimately extracting key monitoring business data from the industrial control system.
[0027] S1.1 Determine the monitoring targets and parameters Identify key monitoring indicators for production line equipment (such as motor vibration frequency, reactor temperature, and pressure threshold), determine the fault type (such as gradual deviation, sudden oscillation, and periodic interruption), and set the acquisition frequency.
[0028] S1.2 Configure Data Source Access Connect to the industrial control system to capture signals from key industrial control equipment such as PLCs in real time, deploy edge gateways to collect sensor Modbus / TCP streams, connect to the SCADA system API to obtain process parameter logs, and synchronously configure industrial firewalls to isolate the network and set up communication certificate encryption.
[0029] S1.3 Data Acquisition High-frequency sensor data is pushed to edge computing nodes in real time, and batch process data is exported at a fixed frequency through the ODBC interface. A local SD card caching and resume transmission mechanism is enabled for power outage conditions.
[0030] S1.4 data storage The time series data streams of the sensors are stored in a database for real-time analysis of the data by the monitoring system.
[0031] S2 model construction To address the challenge of the scarcity and difficulty of labeling of abnormal data in industrial scenarios, the present application proposes an unsupervised fault detection framework. The core method is to use only normal samples to establish a baseline distribution in the training stage, and to achieve fault detection capability by identifying the distribution shift of unseen data in the test stage.
[0032] In the process of constructing the unsupervised fault detection DDA-DE model, first of all, the complexity and non-stationarity of the industrial system operation data are taken into account, and a double-branch detection architecture is proposed to distinguish between "concept drift" and "abnormal event", two phenomena that are similar in appearance but different in nature. Concept drift is manifested as a gradual shift in the statistical characteristics of data distribution, while abnormal events are manifested as a sudden disruption of the correlation structure between features. In order to preserve the inherent feature correlation of the data, the model abandons the traditional single-branch detection approach and instead constructs two parallel but complementary detection paths, allowing drift and anomaly to be decoupled at the mechanism level, thereby avoiding interference with each other and improving the accuracy and interpretability of detection.
[0033] The core task of the drift detection module is to perceive subtle shifts in data distribution in real time under unlabeled conditions. To this end, the model introduces a covariate shift quantification framework based on a sliding window, which dynamically maintains a length-adjustable historical data window and continuously compares the distribution differences between the current window and the reference window in the feature space. Although traditional Mahalanobis distance can measure multi-dimensional distribution deviation, the differences in scale and unit between heterogeneous sensor features affect the results, so the covariance matrix is normalized into Mahalanobis distance (defined as a feature deviation weighted projection based on the inverse covariance matrix). In this way, the drift detection branch can still be sensitive to slight distribution shifts under complex working conditions.
[0034] The fault detection branch focuses on identifying abnormal events that disrupt the inherent correlation between features. Since previous models are more sensitive to concept drift and seek information bottlenecks, which greatly increases training overhead, a diffusion strategy is added to the fault detection module, and in accordance with this design principle, the present application deliberately avoids the variational autoencoder (VAE) architecture, which creates an internal information bottleneck, and instead adopts a more advanced Transformer-based framework. The proposed diffusion Transformer architecture not only captures the temporal features in the industrial data stream more effectively, but also significantly reduces training costs by externalizing the information bottleneck.
[0035] The two-branch detection results are not used in isolation, but drive a dynamic calibration mechanism for model parameters. The drift detection sub-module supports the output of a distribution shift intensity indicator, and when significant drift is detected, triggers a gradual update of the reference distribution to prevent the model from failing due to adherence to outdated benchmarks; at the same time, the fault threshold is adjusted to prevent false positives. The two branches run in parallel, jointly evaluate, and the threshold and model parameters are smoothly adjusted as the system operating state evolves, enabling low false positives during stable periods and timely capture of real anomalies during turbulent periods, achieving continuous optimization of detection performance. Thus, the DDA-DE model completes the closed-loop construction from data input to state discrimination, and can achieve high-precision, self-adaptive unsupervised fault detection in complex industrial environments without any labels.
[0036] As shown in Figure 2 Phase 0 is data preparation, which is the cornerstone of the entire fault detection process. First, the raw industrial time series pulled from DCS, SCADA or edge gateway need to go through multiple rounds of cleaning, using an outlier rejection algorithm based on joint physical constraints and statistical rules to mark and linearly interpolate obvious distorted points; then, according to the equipment operating condition label, non-steady-state segments such as start-stop, maintenance, and load reduction are cut out in their entirety, and only pure normal data under rated operating conditions are retained for training to prevent the model from mislearning the transition state as a normal mode. The test set is deliberately kept complete with a running cycle, including slight drift, gradual degradation and sudden failure, to make the subsequent evaluation more realistic, laying a high-quality data foundation for the unsupervised training and online verification of the DDA-DE model.
[0037] Then, the framework is divided into two stages: Stage 1: training stage; Stage 2: anomaly and drift detection stage. The first stage uses a collaborative training framework, with the dynamic drift awareness (DDA) method and the diffusion enhancement (DE) anomaly detection method trained and optimized in parallel. Specifically, DDA processes the original data stream to establish a statistical distribution baseline and a drift threshold, while DE determines the initial model parameters and the anomaly threshold. The second stage corresponds to the inference stage, during which the integrated DDA-DE framework performs real-time anomaly monitoring. At the same time, the drift detection model performs parallel distribution shift analysis to achieve collaborative dynamic evaluation of data quality. This stage includes dynamic handling mechanisms covering four scenarios: (1) When no fault or drift is detected, normal operation is maintained.
[0038] (2) When a fault is detected but no drift is detected, the fault is marked but the model is not updated.
[0039] (3) When there is no fault but there is drift, the model parameters are incrementally updated using sliding window data.
[0040] (4) When both drift and fault are detected, first adjust the threshold and then verify the fault; once confirmed, mark the fault and continue the subsequent process.
[0041] S2.1 Dynamic Drift Perception Method It needs to be made clear that the essential difference between data drift detection and fault detection lies in the time scale and data change characteristics they focus on. Data drift detection focuses on the gradual change of data distribution over time, usually manifested as the overall shift of statistical characteristics (such as mean, variance, covariance, etc.), reflecting the systematic evolution of data distribution. Fault detection focuses on the local sudden deviation of data at a specific time point, usually manifested as isolated points or short sequences that do not conform to the overall pattern or context, reflecting the transient abnormal behavior of data. The fundamental difference in task purpose provides a theoretical basis and guidance for the structural design of the model, enabling the model to optimize for the unique needs of data drift and fault detection, respectively, and thus achieve more accurate monitoring and analysis in dynamic industrial environments.
[0042] The present invention proposes an online, adaptive, unsupervised drift perception method (DDA) to process industrial flow data, using a fixed sliding window approach, and designs an industrial enhanced Mahalanobis distance to measure the difference between data distributions.
[0043] It is particularly pointed out that the Mahalanobis distance is chosen as the core measurement index for the analysis of industrial sensor data in this study, mainly based on its significant advantages: First, industrial sensor data usually has multi-dimensionality and high correlation between characteristics, while Mahalanobis distance effectively handles dimension relationships by including the covariance matrix, allowing more accurate measurement of the deviation of current data points from historical distributions - especially avoiding misjudgment due to dimension correlation. Second, its scale invariance is crucial for industrial applications - sensor measurements often involve heterogeneous units and scales; Mahalanobis distance eliminates the effects of dimension scale through an internal normalization mechanism, ensuring robust performance when analyzing multi-source heterogeneous data.
[0044] The practical industrial enhancement scheme proposed in this invention mainly reflects two key aspects: First, the distance calculation strategy uses a sliding window-based method, which can dynamically and continuously update the mean and covariance estimates in real time, significantly improving the system's ability to handle common gradual drift phenomena in industrial environments. Second, in view of the inherent differences in sensor characteristics in actual industrial environments - high-frequency sensors (such as temperature, pressure sensors) are more prone to drift than low-frequency sensors (such as valves) - this study introduces an explicit parameter adjustment mechanism in the original Mahalanobis distance formula, prioritizing the processing of easy-drift sensors, thereby improving the applicability and accuracy of the drift detection algorithm in industrial operating scenarios. The implementation steps of the dynamic drift perception (DDA) method are as follows: S2.1.1 Stage One: Training Statistical Data and Thresholds In the training phase, such as Figure 2The first stage shows that DDA extracts statistical descriptors from historical data. This process involves computing baseline distribution parameters, i.e., the mean vector and covariance matrix , of the feature vector to enable real-time distance computation, thus continuously updating the drift threshold. Given the large scale of data in real production environments, the mean and covariance matrix are updated incrementally to achieve progressive update. The incremental computation process is as follows: , , where is the mean at the current time, is the mean at the previous time, is the new data point at the current time, is the total number of data points at the current time, is the covariance matrix at the current time. is the covariance matrix at the previous time. For the feature vector , the formula for computing its distance to the historical distribution is: , where , denotes the Hadamard product (element-wise multiplication). This method effectively eliminates the scale and unit differences among heterogeneous sensor features by normalizing the covariance matrix into Mahalanobis distance (which is defined as a feature bias-weighted projection based on the inverse covariance matrix). The formula for computing , is the temperature coefficient, which controls the steepness of the function curve, is the variance of the i-th sensor.
[0045] The last step is threshold determination during the training phase, where the decision boundary is dynamically adjusted through real-time update, which is calculated based on the mean and standard deviation of the historical distance distribution, as follows: , where is the drift threshold at the current time, is the mean of the historical distance at the current time, is the standard deviation of the historical distance at the current time, is a constant (used to control the degree of deviation of the threshold from the mean).
[0046] S2.1.2 Stage Two: Real-time Drift Detection In the inference phase, the drift threshold is no longer involved in the update, but uses the statistical characteristics obtained in the training phase to dynamically compare and analyze the newly arrived time series. Specifically, first, based on the historical normal data, the stable statistical data and threshold are calculated, which describe the stable mode of the system in the "healthy" state. When a new time window arrives, the industrial enhanced Mahalanobis distance between the window and the historical distribution is calculated through the same distance calculation process. This distance not only considers the correlation between feature dimensions, but also introduces a weighted and robust adjustment for the characteristics of the industrial system, making it more tolerant to fault fluctuations and slight noise. The drift detection result is as follows: , After calculating the distance value, it is compared with the threshold set in advance. If the distance value exceeds the threshold, i.e. , it means that the deviation of the current time window from the historical normal distribution is significant, indicating that the system behavior may have undergone statistical drift or state mutation, which often means that the industrial equipment, production process or sensing system has changed, thereby triggering a drift alarm. On the contrary, if , it means that the features of the current window are still within the range of the historical distribution, and the changes can be considered as normal fluctuations, and the system is in a stable running state. Through this mechanism, the drift detection in the inference phase not only maintains timeliness, but also has stability, providing decision basis for subsequent fault alarm and incremental training of fault detection model.
[0047] S2.2 Diffusion-enhanced fault detection The core goal of the diffusion strategy is to enhance the robustness of the model to concept drift while reducing the training overhead caused by information bottlenecks. Following this design principle, the invention deliberately avoids the variational autoencoder (VAE) architecture that produces internal information bottlenecks, and instead adopts a more advanced Transformer-based framework. The proposed diffusion Transformer architecture not only more effectively captures the time series features in industrial data streams, but also significantly reduces training costs by externalizing information bottlenecks. In addition, to address the challenge of the scarcity of labeled abnormal data in actual production scenarios, the invention uses an unsupervised learning method for fault detection, thereby better adapting to the data characteristics of the industrial environment. The original time series data is input into the diffusion-based fault model (DA). As shown in Figure 2 , the DA model mainly includes a diffusion enhancement module, a word embedding module, a feature extraction module and a reconstruction module, and the specific architecture content and data processing steps are: S2.2.1 Phase one: training parameters and threshold In the training phase, as Figure 2The first stage shows the iterative update of the model parameters of DE through the sliding window input mechanism. Then the fault threshold is derived based on the calculation of the error score. DE contains three steps: injecting time series noise through the diffusion process, extracting features from the noise sequence, and calculating the reconstruction error, the whole process is shown in Figure 3 .
[0048] Injecting time-domain noise through the diffusion process: the input of the diffusion module is , where the original vector at a certain time can be represented as , the M-step Gaussian noise contaminated vector can be represented as , and the contamination result obtained in the mth step is only related to . The result obtained after m-step noise contamination is: , where is the standard Gaussian noise, . is the noise scheduling parameter, which controls the degree of noise addition, and usually decreases with time step m. , where is the noise rate, which increases with time step, The calculation process is: , where and are the preset minimum and maximum noise rates, and M is the total time step. Therefore, we can get the output of the diffusion strategy: , where, , and .
[0049] Position encoding: the variable value after the diffusion module is added with noise is position encoded, and the formula is as follows: , where, represents convolution, represents the position encoded variable value. represents relative position encoding, and the formula is as follows: , represents the position of the time in the time window, represents the dimension of the vector embedding, represents the even dimension in , represents odd dimensions in the middle.
[0050] Feature extraction from noise sequence: the sequence after word embedding is input into the feature extraction structure, aiming to extract the complex spatio-temporal correlation in the data. The feature extraction module is composed of L layers of blocks, where the input of the block module of the l-th layer is , with a time length and dimension, the initial input represents the embedded noise sequence, and in the time module we get , the calculation process of which can be represented as: , , where represents layer normalization, represents multi-head attention mechanism, represents the time hidden representation of the l-th layer, represents the forward feedback network. Similarly, in the spatial module we get , the calculation process of which can be represented as: , , where represents the spatial hidden representation of the l-th layer, represents the transpose. Finally, we get the process of the l+1-th layer : , , where represents the hidden representation, represents the concatenation operation.
[0051] To achieve the goal of unsupervised training, a reconstruction module composed of two fully connected layers is added at the end of the model, aiming to project the output dimension back to the original input dimension. The reconstruction process can be expressed as: , where represents the reconstruction result of , represents the fully connected layer.
[0052] Reconstruction loss calculation: calculate the output result of the model and the original input reconstruction error, which is taken as the loss function value, is back-propagated to the neural network for model parameter updating. Meanwhile, using the reconstruction error as the loss function facilitates the model to find the complex correlations between the time series and sensors. The calculation process is: , where denotes the Frobenius norm.
[0053] S2.2.2 Stage Two: Fault Evaluation Method The model generates a reconstruction error score for each timestamp of sensor values where .
[0054] The intrinsic distribution difference between the faulty data and the normal data is reflected in the reconstruction error score, and the fault score value is significantly higher than that of the normal instance. Based on this discriminant characteristic, the framework implements a threshold detection mechanism: instances with error scores lower than the threshold are determined to be normal, and those exceeding are marked as abnormal. This threshold determination method is not only simple and efficient, but also effectively distinguishes abnormal data from normal data, providing reliable technical support for real-time anomaly detection in industrial scenarios.
[0055] S3. Model Training Optimization During training, the hyperparameters of DDA-DE are configured as follows: the time step m of the noise addition in the diffusion module is set to 2000, the preset minimum and maximum noise rates β min and β max are respectively, the number of layers L in the feature extraction module is 4, the embedding dimension is 128, the size of the sliding window is 64, the loss function uses the reconstruction loss introduced in S2.2.1, the entire model is updated using the mini-batch stochastic gradient descent strategy, the batch size is 64, and the optimizer uses ADAM with a learning rate of 1e-4 and a training period of 20.
[0056] During training, the average loss of the training set and the validation set is calculated after each epoch to monitor the convergence of the model. To prevent overfitting, a dropout of 0.2 is added between the transformer layers, and the training is terminated early when the validation set error does not improve for a certain number of consecutive rounds, i.e., the early stopping mechanism is used. If the validation set reconstruction error does not decrease for 5 consecutive epochs during training, the learning rate decay mechanism is triggered, which automatically reduces the learning rate by half to refine the search. To further ensure the stability of the training, gradient clipping is also performed, which truncates the gradient when its norm exceeds the threshold of 5.
[0057] The whole training process lasts for iterations, and the model gradually learns the normal pattern of the time series. When the parameters of the network converge and the validation set error is the lowest, the model weight is saved. The final model can effectively reconstruct the normal sequence in the inference stage, and for the fault sequence, there will be a large reconstruction error. This feature will be used in the fault detection link to determine the existence of the fault point.
[0058] S4 model deployment The output of the model is the reconstruction error score of the sensor value at each time point , wherein The present application determines whether the data point is faulty according to the drift detection result in step S2 and the threshold value: if the reconstruction error score is lower than the preset threshold value and it is determined that no drift occurs in step S2, it is determined that the data at this time point is normal; if the reconstruction error score is greater than or equal to the threshold value and it is determined that no drift occurs in step S2, it is determined to be faulty; if the reconstruction error score is less than the threshold value and it is determined that drift occurs in step S2, it is determined that the data at this time point is normal, and the model is updated incrementally; if the reconstruction error score is greater than or equal to the threshold value and it is determined that drift occurs in step S2, the threshold value is adjusted and re-evaluated.
[0059] S4.1 preliminarily determines whether a fault occurs. The specific determination method is: , S4.2 If it is determined in step S2 that drift occurs, the abnormal threshold value needs to be adjusted and re-evaluated, and the adjustment method is: ,
[0060] S4.3 final adjustment If no fault is detected and no concept drift occurs, it is not recorded; if an anomaly is detected but no concept drift occurs, the fault is recorded and monitoring continues; if no anomaly is detected but concept drift occurs, the DE is incrementally trained to adapt to changes in data distribution; if both anomaly and concept drift are detected, the threshold value is adjusted according to S4.2 and re-evaluated.
[0061] S4.4 repeats S4.1-4.3 for each sample to be detected, which realizes fault detection of attack behavior in data.
[0062] Experimental verification To verify the effectiveness of DDA-DE, we mainly based on two real datasets for evaluation: (1) The security water treatment (SWaT) dataset, collected from an industrial water treatment plant that produces filtered water, records 11 days of operation data (7 days of normal operation and 4 days of continuous attack state), contains labeled data from 51 sensors and actuators, used to distinguish between normal and abnormal behavior. (2) The water distribution (WADI) dataset, as an extension of SWaT, integrates the three processes of water treatment, storage and distribution network, simulates the complete real water management cycle. The dataset covers 16 days of operation data (including 14 days of normal operation and 2 days of attack state), records data from 123 sensors and actuators. It is worth noting that this study excludes the Mars Science Laboratory (MSL) rover and Soil Moisture Active Passive (SMAP) satellite datasets - these datasets collected by NASA to monitor the status of Mars probe sensors and actuators - cannot be used for rigorous analysis due to inherent defects that have been recorded. To optimize computational efficiency and verify the effectiveness of the method, the continuous variables in the experimental dataset were selectively retained during model construction. The training set was re-divided into training and validation sets in an 8:2 ratio. The adjusted SWaT dataset contains 24 dimensions, with 396,000 samples in the training set, 99,000 samples in the validation set, 449,919 samples in the test set, and 156,915 samples in the validation set. Abnormal samples account for 12.14% of the test set. The WADI dataset contains 59 dimensions, with 627,656 training samples, 156,915 validation samples, and 172,803 test samples. Abnormal samples account for 5.75% of the test set. The specifications of the adjusted dataset are shown in Table 1.
[0063] Table 1 Dataset Details Dataset Dimension Attack Training set Validation set Test set Failure rate (%) SWaT 24 41 396000 99000 449919 12.14 WADI 59 15 627656 156915 172803 5.75 The invention selects recent research results based on deep learning as model benchmarks, D3R and DE*, respectively, and adjusts the method during testing. The baseline method is introduced as follows: D3R: A dynamic decomposition method based on diffusion reconstruction, which decomposes the original non-stationary sequence into trend items and stationary items through hierarchical decoupling mechanism, and combines diffusion model to reconstruct high noise pollution data, to enhance the extraction ability of essential features.
[0064] DE*: Indicates that the DE model designed in this paper is used for modeling stage, combined with the test time adjustment and detrending strategy proposed by Kim et al. The strategy is to use statistical estimation method to remove the trend estimate of time series data, and then carry out online monitoring, and update the model parameters using the detrended sequence.
[0065] The performance evaluation of the model and baseline model of the present application adopts the F1 score based on the correlation, PATE and AUC. The F1 score is obtained by combining the average directional distance between the predicted anomaly and the true event to calculate the precision (P), and the average directional distance between the true event and the predicted anomaly to determine the recall (R) and calculate the F1 score accordingly. The time series anomaly evaluation based on proximity (PATE) is an evaluation index that combines the prediction interval and the time correlation of the anomaly interval, and considers the anomaly interval buffer zone by proximity weighting. The area under the curve (AUC) in the precision-recall (PR) analysis is formally defined as the integral of the PR curve, which quantifies the overall performance of the fault detection model at different thresholds.
[0066] Table 2 compares the experimental results
[0067] The unique parameters of each method follow the optimal configuration in the corresponding paper. As can be seen from the figure, compared with the current most advanced method, the experimental results show that the DDA-DE model surpasses the most advanced D3R model by 2.46% (0.7747→0.7993) in F1 score on the SWaT dataset, and achieves a 3.7% improvement (0.7017→0.7387) on the WADI dataset. It is worth noting that the WADI dataset is usually limited in model performance due to high-dimensional features, low-frequency sampling and hidden attack patterns, while the dense sampling and explicit attack features used in SWaT are more conducive to the performance of existing algorithms. Compared with the test-time adjustment strategy proposed by Kim et al., the present application can effectively model the coupling relationship between trends and seasonal components, and continuously capture evolving trends during deployment, resulting in a 4.96% improvement in F1 score on the SWaT dataset (0.7497→0.7993) and a 2.6% improvement on the WADI dataset (0.7127→0.7387).
[0068] In the PATE index that emphasizes time dependence, the process design of the present application and the inherent advantages of the Transformer architecture enable DDA-DE to exhibit optimal performance, with an improvement of 24.46% (0.6001→0.8447). On the WADI dataset, DDA-DE achieves the optimal 0.3088, and DA* achieves the sub-optimal 0.1924. Under the PATE index, all models perform worse on the WADI dataset than on the SWaT benchmark. It is worth noting that the performance of DDA-DE decreases by more than 50%, which is mainly due to the inherent complexity of WADI as a large-scale water supply system. The spatial dispersion and temporal sparsity of WADI anomaly events exacerbate time series noise and amplify operational variability, thereby fundamentally weakening the model's ability to capture comprehensive temporal proximity.
[0069] DDA-DE achieved the highest AUC scores (88.37%, 57.69%) on both SWaT and WADI datasets. This performance gap confirms the higher requirements on modeling capability posed by the dynamic and high-dimensional characteristics of WADI complex systems, which further exacerbate the challenge of fault detection. These findings establish the enhancement of spatiotemporal modeling capability for complex industrial systems as a core direction for future research.
[0070] The present invention explains the research motivation of the present application from the perspective of visualization. Figure 4 and Figure 5 The drift detection normalization results for all SWaT and WADI datasets are systematically presented, respectively. It is worth noting that although both datasets are derived from similar industrial processes, they exhibit distinct data characteristics - we believe that this difference is the key determinant of the performance difference of the models on these two benchmark datasets.
[0071] Figure 4 Figures 1 and 5 respectively show the normalized detection instances of the DDA-DE framework on the SWaT dataset for fault and continuous fault, revealing its dual-threshold adaptive mechanism: when both drift detection and fault detection outputs exceed the threshold, the system implements dynamic threshold recalibration; in the scenario where the drift detection signal exceeds the preset threshold but is not accompanied by an abnormal signal, selective incremental model updating is triggered.
[0072] Figure 6 The architecture of the present application is presented through a three-layer time series diagram: the bottom layer is the original time series (containing labeled drift points and anomaly points), the middle layer is the drift detection output (containing drift confidence curves and threshold trigger markers), and the top layer is the fault detection score curve (containing dynamic threshold lines adjusted in real time with drift results).
[0073] The drift detection results show that drift is triggered at the beginning of the test set, indicating a significant distribution shift between the test set and the training set normal points, clearly indicating that the time series data is non-stationary - its statistical properties drift over time. In addition, it can be observed from the bottom layer diagram that the fault detection score does not respond sensitively to the drift event.
[0074] Embodiment 2 The present embodiment provides an industrial fault detection system based on dynamic drift perception and diffusion enhancement, comprising: The data acquisition module is configured to: A computer-readable storage medium, wherein a plurality of instructions are stored, the instructions are adapted to be loaded and executed by the processor of the terminal device, and the instructions are a kind of based on dynamic drift perception and diffusion enhancement of industrial fault detection method.
[0075] A terminal device comprises a processor and a computer readable storage medium, the processor is used for realizing instructions; the computer readable storage medium is used for storing a plurality of instructions, the instructions are suitable for being loaded by the processor and executing the kind of industrial fault detection method based on dynamic drift perception and diffusion enhancement.
[0076] The above are preferred embodiments of the present application, not limited by the protection scope of the present application, therefore: all equivalent changes made according to the structure, shape, principle of the present application should be covered in the protection scope of the present application.
Claims
1. An industrial fault detection method based on dynamic drift sensing and diffusion enhancement, characterized in that, include: Acquire industrial flow data; An unsupervised fault detection model, DDA-DE, is constructed. Dynamic drift sensing DDA is used to process industrial data streams to establish a statistical distribution baseline and drift threshold. Diffusion-enhanced anomaly detection DE is used to determine the initial model parameters and anomaly threshold. The robustness of the unsupervised fault detection model to concept drift is enhanced based on a diffusion strategy. The constructed unsupervised fault detection DDA-DE model is optimized. Fault detection is performed using the optimized model.
2. The industrial fault detection method based on dynamic drift sensing and diffusion enhancement according to claim 1, characterized in that, The construction of the unsupervised fault detection DDA-DE model includes establishing a dual-branch detection architecture based on the essential difference between concept drift and abnormal events to preserve the feature correlation within the data. This architecture includes a drift detection module and a fault detection module. The drift detection module adopts a covariate offset quantization model based on a sliding window and dynamically tracks changes in feature space distribution in real time by introducing an industrial-enhanced Mahalanobis distance metric. The fault detection module integrates a diffusion enhancement strategy, enhances the robustness of distribution offset through a multi-scale feature propagation mechanism, and suppresses noise interference by leveraging information bottleneck constraints. The model parameters are dynamically calibrated based on the results of the two detection branches, and the fault detection threshold is optimized through an adaptive threshold adjustment mechanism.
3. The industrial fault detection method based on dynamic drift sensing and diffusion enhancement according to claim 2, characterized in that, The method of using Dynamic Drift Sensing Direct Amplification (DDA) to process industrial data streams to establish statistical distribution baselines and drift thresholds includes extracting statistical descriptors from historical data using DDA, wherein the baseline distribution parameters include the mean vector of the feature vector. Covariance Matrix This is used to achieve real-time distance calculation, thereby continuously updating the drift threshold; and incremental calculation is used to progressively update the mean and covariance matrix. The incremental calculation process is as follows: , , in, This is the mean at the current moment. This is the average value at the previous moment. For the new data point at the current moment, The total number of data points at the current moment. Let be the covariance matrix at the current time. Let covariance be the covariance matrix at the previous time step, and for the eigenvectors... The formula for calculating its distance from the historical distribution is: , in , The Hadamard product is used to eliminate scale and unit differences between features of heterogeneous sensors by normalizing the covariance matrix and incorporating it into the Mahalanobis distance. The calculation formula is: , For temperature coefficient, Let be the variance of the i-th sensor; finally, the threshold is determined during the training phase, where the decision boundary is dynamically adjusted through real-time updates, calculated based on the mean and standard deviation of the historical distance distribution, using the following formula: , in, The drift threshold at the current moment, The average historical distance at the current moment. The standard deviation of the historical distance at the current moment. It is a constant.
4. The industrial fault detection method based on dynamic drift sensing and diffusion enhancement according to claim 3, characterized in that, The method of using diffusion-enhanced anomaly detection (DE) to determine initial model parameters and anomaly thresholds includes using statistical features obtained during the training phase to perform dynamic comparative analysis on newly arrived time series. First, stable statistical data and thresholds are calculated based on historical normal data to depict the stable pattern of the system in a healthy state. When a new time window arrives, the industrial-enhanced Mahalanobis distance between the window and the historical distribution is calculated using the same distance calculation process. By introducing weighted and robust adjustments tailored to the characteristics of the industrial system, it exhibits tolerance for fault fluctuations and minor noise. After calculating the distance value, it is compared with a set threshold. If the distance value exceeds the set threshold, i.e. This indicates that the current time window deviates significantly from the historical normal distribution, suggesting that the system behavior may experience statistical drift or abrupt state changes. This indicates that the characteristics of the current window are still within the historical distribution range.
5. The industrial fault detection method based on dynamic drift sensing and diffusion enhancement according to claim 4, characterized in that, The robustness enhancement of the unsupervised fault detection model for concept drift based on the diffusion strategy includes iteratively updating the model parameters of the DE through a sliding window input mechanism, deriving the fault threshold based on the calculation of error scores, and including injecting temporal noise through the diffusion process, extracting features from the noise sequence, and calculating the reconstruction error. The injection of temporal noise through the diffusion process includes, assuming the input of the diffusion module is... The original vector contained therein at a certain moment is represented as The vectors of M-step Gaussian noise pollution can be expressed as follows: Furthermore, the pollution result obtained in step m is only consistent with... Related data The result obtained after m steps of noise pollution is: , in Standard Gaussian noise, , This is a noise scheduling parameter that decreases with time step m. , in It is the noise rate, which increases with time step. The calculation process is as follows: , in and These are the preset minimum and maximum noise rates, M is the total number of time steps, and finally, the output of the diffusion strategy is obtained: , in, , and The variable values after noise addition by the diffusion module are positionally encoded and represented as follows: , in, Represents convolution. This represents the variable value after position encoding. Relative position encoding is represented as: , Indicates the position of the moment within the time window. This represents the dimension of the vector embedding. express Even-numbered dimensions in express Odd dimensions in.
6. The industrial fault detection method based on dynamic drift sensing and diffusion enhancement according to claim 5, characterized in that, The feature extraction from the noisy sequence includes inputting the word-embedded sequence into a feature extraction structure to extract complex spatiotemporal relationships in the data. The feature extraction module consists of L layers of blocks, where the input to the l-th layer block is... It has a time length as well as Dimension, initial input The embedded noise sequence is obtained in the timing module. The calculation process is expressed as follows: , , in, Presentation layer standardization This indicates a multi-head attention mechanism. This represents the temporal hiding representation of the l-th layer. This represents a feedforward network, also obtained in the spatial module. , represented as: , , in This represents the spatial hiding representation of the l-th layer. This indicates transpose; finally, the (l+1)th layer is obtained. The process is as follows: , , in This indicates that the symbol is hidden. This indicates a splicing operation.
7. The industrial fault detection method based on dynamic drift sensing and diffusion enhancement according to claim 6, characterized in that, The calculation of reconstruction error includes, to achieve the unsupervised training objective, adding a reconstruction module consisting of two fully connected layers at the end of the model to project the output dimension back to the original input dimension. The reconstruction process is described as follows: , in express The reconstruction results This represents a fully connected layer; then the model's output is calculated. With the original input The reconstruction error is used as the loss function value and backpropagated back to the neural network to facilitate model parameter updates. Simultaneously, using the reconstruction error as the loss function helps the model find complex correlations between time series and various sensors. The calculation process is as follows: , in Denotes the Frobenius norm; A reconstruction error score is generated for the sensor values at each time stamp. ,in The inherent distributional differences between faulty and normal data are reflected in the reconstruction error score. Error scores below a certain threshold are considered acceptable. The instances were judged to be normal, exceeding [a certain threshold]. Then it is marked as an exception.
8. The industrial fault detection method based on dynamic drift sensing and diffusion enhancement according to claim 7, characterized in that, The optimization of the constructed unsupervised fault detection DDA-DE model includes configuring the hyperparameters of DDA-DE during training, setting the time step m for adding noise in the diffusion module to 2000, and presetting the minimum and maximum noise rates β. min and β max The feature extraction module has 4 layers L and an embedding dimension of 4. The size of the sliding window is 128. The loss function is 64, the reconstruction loss is used, the model is updated using mini-batch stochastic gradient descent, the entire training set is traversed in each training epoch, and the ADAM optimizer is used. During training, the average loss of the training set and the validation set is calculated after each epoch to monitor the convergence of the model. To prevent overfitting, dropout is added between transformer layers, and training is terminated early when the validation set error does not improve after several consecutive epochs, i.e., an early stopping mechanism is adopted. If the validation set reconstruction error does not decrease for 5 consecutive epochs during training, the learning rate decay mechanism is triggered, and the learning rate is automatically reduced by half to refine the search. The entire training process is iterative, and the model gradually learns the normal pattern of the time series. When the network parameters converge and the validation set error is at its lowest, the model weights are saved.
9. The industrial fault detection method based on dynamic drift sensing and diffusion enhancement according to claim 8, characterized in that, The fault detection using the optimized model includes calculating the reconstruction error score of the sensor values at each time point based on the model output. ,in The system determines whether a data point is faulty based on both the drift detection results and a threshold: if the reconstruction error score is lower than a preset threshold... If no drift is determined, then the data at that time point is considered normal. If the reconstruction error score is greater than or equal to the threshold If no drift is detected, it is considered a fault; if the reconstruction error score is less than the threshold... If a drift is detected, the data at that time point is considered normal, and the model is incrementally updated; if the reconstruction error score is greater than or equal to the threshold... If drift is detected, the threshold is adjusted. And reassess; the initial assessment method for determining whether a fault has occurred is as follows: ; If drift is detected, the abnormal threshold is adjusted and reassessed. The adjustment method is as follows: , ; Finally, if no fault is detected and no concept drift occurs, no record is made; if an anomaly is detected but no concept drift occurs, the fault is recorded and monitoring continues; if no anomaly is detected but concept drift occurs, the DE is incrementally trained to adapt to changes in data distribution; if both anomaly and concept drift are detected, the threshold is adjusted and re-evaluated.
10. An industrial fault detection system based on dynamic drift sensing and diffusion enhancement, characterized in that, include: The data acquisition module is configured to acquire industrial flow data; The model building module is configured to build an unsupervised fault detection DDA-DE model, wherein dynamic drift sensing DDA is used to process industrial data streams to establish a statistical distribution baseline and drift threshold; diffusion-enhanced anomaly detection DE is used to determine initial model parameters and anomaly thresholds; and the unsupervised fault detection model is enhanced for concept drift robustness based on a diffusion strategy. The model optimization module is configured to optimize the constructed unsupervised fault detection DDA-DE model. The fault detection module is configured to perform fault detection using an optimized model.
Citation Information
Patent Citations
Industrial Internet of Things equipment fault diagnosis system for self-adaptive triggering incremental learning
CN114895656A
Conceptual drift-oriented adaptive interpretable industrial control system anomaly detection method
CN116991137A
Fan fault diagnosis method based on multi-working-condition data synthesis and Bayesian anomaly probability enhancement
CN120030454A
Roadbed settlement data identification method based on artificial intelligence
CN120145201A
Equipment maintenance fault intelligent analysis method and system based on Internet of Things
CN120387553A
Cited By
Continuous learning model hot updating method for industrial product defect detection
CN121542815A
Semi-real-time prediction model training method based on data distribution drift hierarchical triggering
CN121834358A
Distributed modeling wind turbine generator state monitoring method
CN121959133A
Baseline deviation identification method and system for low-dimensional discrete perception in industrial process
CN122110713A
Gear machinery multifunctional composite measurement method based on depth domain adaptation
CN122192758A