Time series data anomaly detection method based on variable time converter
By introducing a VT-GAN model based on a variable time converter in time series data anomaly detection, the existing methods are solved inadequate real-time and strong noise sensitivity in industrial scenarios, and more efficient anomaly detection performance and better spatiotemporal interaction modeling are achieved.
Patent Information
- Application Number
- CN202510623076.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
Existing time series data anomaly detection methods face problems such as insufficient real-time, strong noise sensitivity and limited diversity of abnormal patterns in industrial scenarios, and traditional deep learning models are difficult to fully address these challenges.
Using a time series data anomaly detection method based on a variable time converter (VT-GAN), the VT-GAN model is constructed, including generators and discriminators, and connected through an adversarial training framework. The interactive process includes two parts: data flow and gradient backpropagation, and optimizes the loss function of the VT-GAN model for data anomaly detection.
Through parallel sub-generators and gradient penalty constraints, the pattern crash problem is alleviated, sample diversity is improved, space-time interaction is explicitly modeled, and the problem of insufficient modeling of variable coupling effect is solved, achieving more efficient anomaly detection performance.
Smart Images

Figure CN120144930A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of time series data processing, and more specifically, relates to a method for detecting anomalies in time series data based on a variable time transformer. Background Art
[0002] In the context of industry and intelligent manufacturing, production equipment is widely equipped with sensors to achieve real-time monitoring of the operating status. The multi-source time series data collected by these sensors naturally carry key information about the health status and performance parameters of the equipment. Although anomaly detection techniques based on multi-source time series can effectively achieve early fault warning and remaining useful life prediction by analyzing multi-source sensor data such as vibration, temperature, and pressure, existing methods face three core challenges in industrial scenarios: (1) complex time series dependencies and variable coupling effects; (2) high noise interference caused by sensor errors and communication delays; (3) scarcity of labeled anomaly data in dynamic environments.
[0003] Chinese Patent Document CN118982155A discloses a method for predicting and regulating urban local carbon emission hotspots based on a generative adversarial network, including the following steps: S1, data collection and preprocessing; S2, constructing a short-term memory module using hourly carbon emission data for short-term spatio-temporal feature encoding; S3, based on daily and weekly carbon emission data, using a multi-scale convolutional network to extract periodic change features for mid-term spatio-temporal feature encoding; S4, using a long time series dependence model to analyze monthly and quarterly carbon emission data to obtain long-term spatio-temporal features for global encoding of the overall urban carbon emission distribution; S5, constructing and training a generative adversarial network; S6, using a spatio-temporal attention mechanism to perform weighted fusion of spatio-temporal features at different levels, and combining real-time data to generate a regulation strategy; S7, monitoring the regulation effect in real time for adaptive adjustment and optimization.
[0004] With the performance degradation caused by the aging of industrial equipment, multi-source time series anomaly detection is crucial for achieving equipment health management (PHM) and preventive maintenance. However, existing methods face challenges such as insufficient real-time performance, strong noise sensitivity, and limited anomaly pattern diversity in complex industrial scenarios; traditional deep learning models are difficult to comprehensively address these challenges. The recurrent neural network RNN has the defect of insufficient parallelization ability, while methods based on Transformer usually cannot model variable-specific time series patterns. Although the generative adversarial network GAN performs excellently in data generation, it faces problems of mode collapse and gradient instability in high-dimensional time series scenarios. In addition, traditional meta-learning methods lack an explicit mechanism for aligning multi-source time series features, as well as a high false alarm rate and low anomaly detection coverage. Summary of the Invention
[0005] The present invention aims to overcome at least one defect of the above-mentioned prior art, and provides a method for detecting anomalies in time series data based on a variable time transformer, so as to solve the problems that the traditional model is difficult to meet the millisecond-level response requirement due to the sequence calculation characteristics, the computational complexity of the global attention mechanism in high-dimensional time series data limits the real-time performance, and the false alarm rate is high and the anomaly detection coverage rate is low.
[0006] The detailed technical solution of the present invention is as follows: A method for detecting anomalies in time series data based on a variable time transformer, the method comprising: S1. Obtain time steps, observation data composed of sensor variables, and perform preprocessing to obtain real sample data; S2. Construct a VT-GAN model, including a generator and a discriminator, and connect them through an adversarial training framework. The interaction process includes two parts: data flow and gradient backpropagation;
[0007] Further, the observation data is preprocessed before being input into the VT-GAN model, specifically including: The support set and query set of the observation data come from different stages of the device life cycle. The multivariate time series data of the device is , with a total of N devices. Each sample includes sensor readings of dimensions at time steps. c represents the current moment, and represents the set of real numbers; The support set represents data sampled from E consecutive time windows in the early operation stage of the device. e represents the time window in the early operation stage traced back from the current moment to the past, with a total of E windows, which are used to construct the support set; The query set represents data sampled from consecutive time windows in the later operation stage of the same device. m represents the time window in the later operation stage extended from the current moment to the future, with a total of M windows, which are used to construct the query set; There is a time offset between the query set and the support set, and a bidirectional LSTM, that is, BiLSTM alignment loss is introduced: In formula (1), is the time series alignment loss, and respectively represent the feature vectors of the current moment c in the support set and the query set; The observed data is finally pre - processed to obtain real sample data , which constitutes the input data set.
[0008] Furthermore, the S2 specifically includes: S21. Construct the generator of the VT - GAN model, including a short - term generator, a medium - term generator, and a long - term generator; Input the noise vector , and respectively generate hourly - level time - series patterns through the short - term generator ; generate daily - level trend patterns through the medium - term generator ; generate weekly - level cycle patterns through the long - term generator ; Finally, output the fused fake samples through the gated fusion module .
[0009] S22. Construct the discriminator of the VT - GAN model, including an embedding layer, a temporal self - attention module, a variable self - attention module, a cross - attention module, and a gated fusion module; For the discriminator, input real samples and generated samples , map the input to high - dimensional features through the embedding layer ; respectively capture univariate time - series dependencies through the temporal self - attention module, model cross - variable interactions through the variable self - attention module, and explicitly fuse spatio - temporal features through the cross - attention module; Finally, dynamically weight the outputs of the temporal self - attention module, the variable self - attention module, and the cross - attention module through the gated fusion module to generate discriminative features , and output the discrimination probability according to the discriminative features.
[0010] Furthermore, the short - term generator G1 is stacked by 3 layers of dilated causal convolutions, and the size of each convolutional kernel is , and the dilation factor is increasing, where is the number of layers, and the receptive field gradually expands to time steps; The medium - term generator G2 includes 6 layers of dilated causal convolutions stacked, and the size of each convolutional kernel is , the dilation factor of the first 3 layers is , and the dilation factor of the last 3 layers is fixed at , and the receptive field is 63 steps; The long - term generator G3 uses 12 layers of dilated causal convolutions with residual connections, and the size of each convolutional kernel is ; The core operation of each generator is implemented by dilated causal convolutions, and the processing function is as follows: (2); In formula (2), represents the output feature of the th layer of the s-th generator, and is a one-dimensional causal convolution, and is the dimension of the latent space, and represents the output feature of the
[0011] -1th layer of the s-th generator. Furthermore, the fused fake samples are output through the gated fusion module (3); In formula (3), is the output of the th generator, and is the hidden state of each generator, and means the dimension of the generator hidden state, that is, the length of the hidden feature vector of each generator, which is extracted from the last layer of convolutional features; is a learnable parameter matrix that dynamically calculates the fusion weights of each branch, and ensures that the weights satisfy through Softmax normalization.
[0012] Furthermore, the univariate time series dependencies are captured respectively through temporal self-attention, the cross-variable interactions are modeled through variable self-attention, and the spatio-temporal features are explicitly fused through cross-attention, specifically including: The temporal self-attention module first splits the embedded feature into independent time series along the variable dimension, For each variable , the query matrix , the key matrix , and the value matrix are calculated: (4); In formula (4), , , is a learnable projection matrix; Then the temporal self-attention output is calculated, and is the unified dimension of the Q / K / V vectors in the temporal self-attention: (5); Concatenation result of the output of the temporal self-attention module along the variable dimension: (6); The variable self-attention module for each time step , introducing a learnable variable correlation matrix , enhancing prior knowledge guidance: (7); Where , , is a learnable projection matrix; Then calculate the output of the variable self-attention: (8); Concatenation result of the output of the variable self-attention module along the time dimension: (9); The cross-attention layer uses the temporal self-attention output as the query source and the variable self-attention output as the key-value source: (10); Among them, , , is the projection matrix.
[0013] Then, calculate the cross-attention output : (11); Dynamically integrate multi-modal attention features, introducing learnable gating weights to generate discriminative features : (12); (13); In formulas (12) to (13), is the gating parameter matrix, ensuring = 1 through Softmax normalization, is the temporal attention weight, is the variable attention weight, is the cross-attention weight.
[0014] Furthermore, the discriminative probability is output according to the discriminative features, specifically including: Apply dilated causal convolution to the fused feature to enhance local feature extraction: (14); In formula (14), is the output feature map after dilated causal convolution processing of the fused feature Z, i.e., the dilated convolution feature. is the hidden layer feature dimension, representing the final encoding length of each time step-variable pair; Compress the feature along the time and variable dimensions, and extract the global feature vector from the dilated convolution feature : (15); Output the discriminant probability : (16); In formula (16), is the Sigmoid function, , is a learnable parameter, is the learnable parameter of the fully connected layer; On the other hand, the discriminant probability is passed to the generator, enabling the generator to generate data closer to the real data. Then the discriminator discriminates between true and false to perform adversarial training.
[0015] Furthermore, optimize the loss function of the VT-GAN model, specifically including: Introduce a contrast learning mechanism into the contrastive adversarial loss, and further combine the temperature parameter adjustment and the dynamic mining strategy of abnormal samples. The loss function formula is as follows: (17); In formula (17), is the real sample, i.e., the generated sample, is the real abnormal sample, is the discriminant probability of the generated sample, is the discriminant probability of the abnormal sample, i.e., the abnormal probability; is the temperature parameter, controlling the steepness of the probability distribution.
[0016] In another aspect of the present invention, there is provided a device for an anomaly detection method of time series data based on a variable time transformer. The device includes: At least one processor; And a memory that stores instructions. When the instructions are executed by the at least one processor, the at least one processor executes an anomaly detection method of time series data based on a variable time transformer as described above.
[0017] In another aspect of the present invention, there is also provided a computer-readable storage medium storing executable instructions that, when executed, cause the machine to execute a method for anomaly detection of time series data based on a variable time transformer as described above.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The method for anomaly detection of time series data based on a variable time transformer provided by the present invention adopts a GAN variant enhanced by a residual network, effectively alleviates the mode collapse problem and improves the sample diversity through a parallel sub-generator and gradient penalty constraints; the time self-attention, variable self-attention and cross-attention layers are introduced into the VTT, and the spatio-temporal interaction is explicitly modeled through learnable gating weights, solving the defect of insufficient modeling of variable coupling effects by traditional methods.
[0019] (2) The method for anomaly detection of time series data based on a variable time transformer provided by the present invention realizes the fast parameter adaptation of the LSTM network in a small-sample anomaly scenario (less than 10-step gradient update) based on the model-agnostic meta-learning framework, and the training time is greatly reduced compared with the standard meta-learning method. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flowchart of the method for anomaly detection of time series data based on a variable time transformer according to the present invention.
[0021] Figure 2 is a schematic diagram of the VT-GAN model in Embodiment 1 of the present invention.
[0022] Figure 3 is a schematic diagram of the generator part in Embodiment 1 of the present invention.
[0023] Figure 4 is a schematic diagram of the discriminator part in Embodiment 1 of the present invention.
[0024] Figure 5 is a schematic diagram of the time series anomaly detection results of three dimensions on the SMD dataset in Embodiment 1 of the present invention.
[0025] Figure 6 is a schematic diagram of the decreasing trend of the average training loss of the VT-GAN model on the SWaT dataset in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The present invention will be further described below with reference to the drawings and embodiments.
[0027] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0028] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "comprising" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0029] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0030] Embodiment 1 Refer Figure 1 , this embodiment provides a method for detecting anomalies in time series data based on a variable time converter. The method includes: S1. Obtain observation data and perform preprocessing: Obtain from industrial equipment time steps, observation data composed of sensor variables, and perform preprocessing to obtain real sample data. The health status monitoring of industrial equipment depends on the anomaly detection of multivariate time series data (Multivariate Time Series, MTS); The preprocessing can be methods such as handling missing values and data standardization. Preferably, in this embodiment, the observation data is preprocessed before inputting into the VT-GAN model, specifically including: In the existing work, there is a situation where the time series patterns do not match because the support set and the query set come from different stages of the equipment life cycle. Therefore, diversified time series tasks are proposed: The support set and the query set of the observation data come from different stages of the equipment life cycle. Any device The multivariate time series data of is , a total of N devices, each sample includes time steps of dimensional sensor readings. c represents the current moment, represents the set of real numbers; Support set represents the data sampled from E consecutive time windows in the early operation stage of the device. e represents the time window in the early operation stage that traces back from the current moment to the past. There are a total of E windows, which are used to construct the support set; Query set Denote the data sampled from the later operation stage of the same device for \(m\) consecutive time windows, where \(m\) represents the time window extending from the current moment to the future, i.e., the later operation stage, and there are a total of \(M\) windows, which are used to construct the query set; There is a time offset between the query set and the support set. To address the distribution differences caused by the time offset, a bidirectional LSTM, i.e., BiLSTM alignment loss, is introduced. The BiLSTM alignment loss acts on the feature space inside the VT-GAN model, rather than directly modifying the input data: (1); In formula (1), is the temporal alignment loss, represent the feature vectors at the current moment \(c\) in the support set and the query set respectively. This loss forces the model to learn the common temporal patterns across time periods and alleviates the influence of phase offsets; The observed data is finally preprocessed to obtain the real sample data , which constitutes the input data set.
[0031] S2. Construct the VT-GAN model. The VT-GAN model includes a generator and a discriminator, and is connected through an adversarial training framework. The interaction process includes two parts: data flow and gradient backpropagation. The VT-GAN model is as Figure 2 shown. The generator generates time series data from the latent space, and the discriminator can determine whether these data come from the real time series data set or are generated by the generator.
[0032] S21. The generator of the constructed VT-GAN model includes: a short-term generator, a medium-term generator, a long-term generator, and a gated fusion module: For the generator, as Figure 3 shown, the input noise vector , where \(d\) is the dilation factor: The short-term generator generates hourly temporal patterns through dilated causal convolutions ; the medium-term generator generates daily trend patterns through dilated causal convolutions ; the long-term generator generates weekly periodic patterns through dilated causal convolutions , and finally outputs the fused fake samples .
[0033] To capture the dynamic patterns at different time scales during the degradation process of industrial equipment, this work designs a multi-generator architecture. Each generator focuses on feature extraction and generation at a specific time scale. The multi-generator architecture diagram is as Figure 3As shown in the figure. Through hierarchical expansion with the dilation factor, the short-term generator captures local fluctuations, the medium-term generator models the daily trends, and the long-term generator captures the weekly cycles, covering the entire process of equipment degradation. By virtue of the local connection property of dilated causal convolution, the problem of vanishing gradients in traditional GANs is alleviated, and the Wasserstein loss is combined to further ensure the convergence of training.
[0034] Specifically, the short-term generator G1 is stacked by 3 layers of dilated causal convolution, and the size of each convolution kernel is , and the dilation factor increases according to (where is the number of layers), and the receptive field is gradually expanded to time steps. The short-term generator is used to capture the high-frequency changes in the equipment state. High-frequency changes such as sudden temperature rise and pressure peaks are quickly responded to local features through a shallow convolutional network.
[0035] The medium-term generator G2 consists of 6 layers of dilated causal convolution stacked, and the size of each convolution kernel is . The dilation factor of the first 3 layers is , and the dilation factor of the last 3 layers is fixed at 8. The receptive field is 63 steps, ensuring a stable modeling of the daily pattern. The medium-term generator is used to identify the progressive degradation of equipment performance, balancing local features and medium- to long-term dependencies through a medium-depth network structure.
[0036] The long-term generator G3 adopts 12 layers of dilated causal convolution with residual connection (Residual Connection), as shown in Figure 3 . It includes a fully connected layer and a convolutional layer connected by three residual blocks. The size of each convolution kernel is , and the dilation factor of each layer is . The final receptive field covers time steps. The long-term generator is used to analyze the macroscopic evolution law of the entire life cycle of the equipment, capturing complex correlations across time steps through a deep network.
[0037] Preferably, the core operation of each generator is implemented by dilated causal convolution, and the processing function is as follows: (2); In formula (2), represents the output feature of the -th layer of the s-th generator. is a one-dimensional causal convolution, ensuring that the output only depends on the current and historical inputs and avoiding leakage of future information. is the size of the convolution kernel, is the dilation factor, which grows exponentially with the network depth
[0038] Preferably, the fused fake samples are output by setting up a gated fusion network, which specifically includes: To integrate the outputs of multiple generators and generate the final synthetic samples, a dynamic gated fusion mechanism is proposed: (3); In formula (3), is the output of the -th generator, is the hidden state of each generator, means the dimension of the generator hidden state, i.e., the length of the hidden feature vector of each generator, which is extracted from the last layer of convolutional features through global average pooling; is a learnable parameter matrix that dynamically calculates the fusion weights of each branch .
[0039] A gated network is introduced for the characteristics of time series data, and Softmax normalization is used to ensure that the weights satisfy .
[0040] S22. The discriminator of the constructed VT-GAN model includes: an embedding layer, a temporal self-attention module, a variable self-attention module, cross-attention, and a gated fusion module; For the discriminator, as Figure 4 shown, the real sample and the generated sample are input, and the input is mapped to high-dimensional features through the embedding layer; the temporal self-attention is used to capture the univariate time series dependencies respectively, the variable self-attention is used to model the cross-variable interactions, and the cross-attention is used to explicitly fuse the spatio-temporal features; Finally, the gated fusion module dynamically weights the outputs of the temporal self-attention module, the variable self-attention module, and the cross-attention module to generate discriminative features , and the discriminative probability is output according to the discriminative features.
[0041] Preferably, the discriminator uses a dynamic hybrid attention discriminator, as Figure 4 shown; The dynamic hybrid attention discriminator parallelly models the time, variable, and spatio-temporal cross dependencies through three dedicated components, and uses learnable gated weights to dynamically fuse the multi-modal features, significantly improving the adaptability of the model to complex working conditions.
[0042] For the dynamic hybrid attention discriminator, the real sample and the generated sample are input, and the input is mapped to high-dimensional features ; Capture the univariate time series dependencies through temporal self-attention respectively, model the cross-variable interactions through variable self-attention, explicitly fuse the spatio-temporal features through cross-attention, and finally dynamically weight the outputs of each attention to generate discriminative features And calculate the anomaly score.
[0043] Preferably, to extract the deep spatio-temporal features of multivariate time series, first process the multivariate time series through the input embedding layer, map the original input to a high-dimensional latent space, enhance the feature separability using a learnable parameter matrix, and provide a structured input for the subsequent attention mechanism. This embedding-attention combined architecture effectively captures the deep spatio-temporal pattern features while maintaining computational efficiency: The discriminator maps the input to a high-dimensional hidden space through the embedding layer to obtain the embedding features : (18); In formula (18), and are the learnable parameters of the embedding layer, is the dimension of the hidden space, usually set to 64 or 128 to balance computational efficiency and representational ability. This embedding process enhances the feature separability through linear transformation and provides a structured input for the subsequent attention mechanism.
[0044] The dynamic hybrid attention module consists of three parallel parts, modeling: temporal, variable, and spatio-temporal cross dependencies respectively.
[0045] The goal of temporal self-attention is to capture the long-term dependencies across time steps within a single variable. First, split the embedding features along the variable dimension into independent time series , where . For each variable , calculate the query matrix , the key matrix , and the value matrix : (4); In formula (4), , , is the learnable projection matrix, is the unified dimension of the Q / K / V vectors in temporal self-attention.
[0046] Calculate the output of temporal self-attention : (5); The concatenation result along the variable dimension of the output of the temporal self-attention module: (6); The goal of Variable Self-Attention is to model the dynamic coupling relationship between different variables at the same time step. For each time step , a learnable variable correlation matrix is introduced to enhance the prior knowledge guidance: (7); where , , is a learnable projection matrix.
[0047] Calculate the variable self-attention output: (8); The concatenation result of the variable self-attention module output along the time dimension: (9); The goal of the Cross-Attention Layer is to explicitly fuse the features of the time and variable dimensions and model the cross-space-time interaction pattern. Using the time self-attention output as the query source and the variable self-attention output as the key-value source: (10); where , , is a projection matrix.
[0048] Then, calculate the cross-attention module output : (11); Dynamically integrate the multi-modal attention features, and introduce learnable gating weights to generate the fused feature, namely the discriminant feature : (12); (13); In formulas (12) to (13), is the gating parameter matrix. Through Softmax normalization, it is ensured that = 1, is the time attention weight, is the variable attention weight, is the cross-attention weight.
[0049] Through the discriminant feature The discrimination probability can be obtained, for example, directly through a fully connected layer; preferably, in this embodiment, the fused feature is first applied with dilated causal convolution to enhance local feature extraction: (14); In formula (14), is the output feature map after processing the fused feature Z by dilated causal convolution, that is, the dilated convolution feature, is the hidden layer feature dimension in the model, representing the final encoding length of each time step-variable pair.
[0050] Compress the features along the time and variable dimensions, and extract the global feature vector from the dilated convolution feature : (15); Output the discrimination probability : (16); In formula (16), is the Sigmoid function, , is a learnable parameter, is the learnable parameter of the fully connected layer; On the other hand, the discrimination probability is passed to the generator, so that the generator generates data closer to the real data, and then the discriminator discriminates between true and false for adversarial training.
[0051] S3. Optimize the loss function and perform detection and discrimination: Combine the temperature parameter adjustment to optimize the loss function of the VT-GAN model and perform data anomaly detection based on the constructed VT-GAN model.
[0052] Specifically, optimizing the loss function of the VT-GAN model specifically includes: Traditional adversarial loss only optimizes the discrimination probability of generated samples and lacks explicit constraints on the distribution boundaries of positive and negative samples.
[0053] The contrastive adversarial loss solves the problem of blurred distribution boundaries in traditional adversarial training by introducing a contrastive learning mechanism to explicitly optimize the discriminator's ability to distinguish normal samples, generated samples, and abnormal samples. Further combined with the temperature parameter adjustment and the dynamic mining strategy of abnormal samples, this loss function significantly improves the anomaly detection performance of the model in industrial scenarios with few labels and high noise. The loss function formula is as follows: (17); In formula (17) is the real sample, i.e., the generated sample, is a real abnormal sample or a sample with a high reconstruction error (false negative sample), is the discrimination probability of the generated sample, is the discrimination probability of the abnormal sample, i.e., the anomaly probability; is the temperature parameter, with a default = 0.1, which controls the steepness of the probability distribution. By introducing , the model can adaptively mine potential abnormal patterns in an unsupervised scenario, alleviating the cold start problem in anomaly detection.
[0054] The experiments of this embodiment are proved as follows: The VT-GAN model was compared with the state-of-the-art models for multivariate time series anomaly detection, including MERLIN, DAGMM, OmniAnomaly, MSCRED, USAD, and TranAD. All models were trained using the PyTorch-1.7.1 library in this embodiment. At the same time, the AdamW optimizer was used with an initial learning rate of 0.01, a meta-learning rate of 0.02, and a step scheduler of 0.5 to train the models, including using the following hyperparameter values: VTT: 4-layer encoder, 8-head attention, hidden layer dimension 256.
[0055] GAN: 3 generators, with dilation factors of 1 / 8 / 16 respectively, and the gradient penalty coefficient λ = 10.
[0056] MAML: 5 inner loop steps, 10 outer loop steps, learning rate 0.01.
[0057] To train the model, the training time series was divided into 80% training data and 20% validation data.
[0058] Six publicly available datasets were used in this experiment. To directly compare with existing methods, these widely used public datasets were selected.
[0059] 1) MIT-BIH Supraventricular Arrhythmia Database MBA: It is a collection of electrocardiogram records from four patients, including multiple instances of two different types of abnormalities.
[0060] 2) Soil Moisture Active Passive Dataset SMAP: It is a dataset of soil samples and telemetry information using a Mars probe by NASA.
[0061] 3) Server Machine Dataset SMD: This is a five-week dataset of stacked traces of resource utilization from 28 machines in a computing cluster.
[0062] 4) Safe Water Treatment Dataset SWaT: This dataset was collected from a real water treatment plant during 7 days of normal operation and 4 days of abnormal operation. This dataset consists of sensor values and actuator operations.
[0063] 5) Mars Science Laboratory Dataset MSL: It is a dataset similar to SMAP, but corresponding to the sensor and actuator data of the rover itself.
[0064] 6) XJTU-SY Bearing Dataset XJTU: This dataset includes the complete run-to-failure data of 15 rolling bearings, which were obtained by conducting many accelerated degradation experiments.
[0065] In Table 1, P: Precision, R: Recall, L: Latency (milliseconds), F1: F1 score under complete training data. The VT-GAN model has the optimal P value and F1 value on both datasets.
[0066] Table 1: Performance comparison between the present invention and baseline methods on two datasets MBA and XJTU
[0067] In this embodiment, accuracy, recall, ROC / AUC, and F1 score are used as core evaluation metrics, and training time (h) and parameter count (M) are used as auxiliary evaluation metrics.
[0068] Precision, recall, F1 score, and training latency (in milliseconds) are used as core evaluation metrics, and training time (in hours) and parameter count (in millions) are used as auxiliary evaluation metrics.
[0069] Table 1 compares the performance of the present invention with multiple baseline methods on two datasets MBA and XJTU. The evaluation metrics include precision (P), recall (R), latency time (L(ms)), and F1 score (F1): Judging from the results, the overall performance of VT-GAN is the best. It achieved the highest precision of 0.9846 and the highest F1 score of 0.9825 on the MBA dataset. At the same time, on the XJTU dataset, it also led other methods with a precision of 0.9547 and an F1 score of 0.9477, and the latency time remained within 30 milliseconds.
[0070] In contrast, although MERLIN has the fastest latency of 5 ms (MBA) and 8 ms (XJTU), its recall performance is poor, especially on the MBA dataset with only 0.4923. Among other methods, DAGMM and MSCRED are prominent in terms of precision but have high latency (34 - 36 ms), while OmniAnomaly and USAD perform excellently in terms of recall (0.9477 - 0.9656).
[0071] It is worth noting that the two datasets exhibit different characteristics: the precision of the MBA dataset has a large span (0.8561 - 0.9846), and the F1 scores vary significantly (0.6564 - 0.9825); while the performance of each method on the XJTU dataset is more balanced, and the F1 scores are concentrated in the range of 0.8055 - 0.9477.
[0072] Overall, VT - GAN demonstrates the best balance, achieving leading positions in all key metrics on both datasets while maintaining low latency. The best precision and F1 scores are obtained by VT - GAN.
[0073] As Figure 5 shown, it presents the time - series anomaly detection results in three dimensions on the SMD dataset: taking the 19th dimension as an example, the horizontal axis represents the timestamp (from 0 to approximately 26000), the vertical axis represents the data value range and the anomaly score (from 0 to 1), the black curve represents the smoothed result of the predicted value; the red dashed line is the anomaly label; the blue semi - transparent area indicates the time when the anomaly occurs; the green curve represents the anomaly score. Figure 5 They are the presentations of the 18th dimension, 19th dimension, and 20th dimension respectively. The upper sub - figure shows the time series of the original data, and it can be seen that the data has significant fluctuations at some time points; the lower sub - figure shows the anomaly scores of the corresponding time series, and the anomaly scores at some time points are extremely high, indicating that these points are identified as anomaly points. There is a significant anomaly peak at timestamp = 15000, the blue shaded area covers this interval, and the anomaly score also reaches a peak here, confirming this anomaly point.
[0074] In these three dimensions, the anomaly at timestamp = 15000 is successfully detected, and the anomaly score reaches a peak at this position, indicating that the model has consistency in the detection results of the same anomaly point in different dimensions. The predicted value, i.e., the red line, is close to the true value, i.e., the black line, at most time points, but there are significant deviations near the anomaly points, which helps the model identify anomaly points. The model can effectively identify the same anomaly points in different dimensions, indicating its good generalization ability and stability. The anomaly score can accurately capture anomaly points and remain at a low level during normal periods, indicating that the model has strong anomaly detection capabilities.
[0075] The contrastive adversarial loss significantly enhances the discriminator's ability to distinguish normal samples, generated samples, and abnormal samples by introducing a contrastive learning mechanism, solving the problem of blurred distribution boundaries in traditional adversarial training. Combined with temperature scaling and a dynamic abnormal sample mining strategy, this loss function significantly improves the anomaly detection performance in industrial scenarios with limited labels and high noise.
[0076] As Figure 6 shown, it presents the downward trend of the average training loss of the VT-GAN model on the SWaT dataset: the vertical axis represents the average training loss, gradually decreasing from the initial value of 0.05 to 0.02, indicating that the model loss steadily decreases during training and the parameters are continuously optimized; the horizontal axis represents the number of training epochs, and each epoch corresponds to a complete traversal of the dataset once, with a total of 4 rounds of training. Near the 4th round, the loss drops to 0.0072, indicating that the model obtains better performance in the later stage of training. The loss value monotonically decreases with the increase in the number of rounds, indicating that the model of the present invention does not overfit and the optimization process is stable.
[0077] The ablation experiments are as follows: Through systematic ablation experiments on the core components of the VT-GAN model, as shown in Table 2: the specific contributions of each module to the model performance are deeply analyzed. The experimental results show that the full-version VT-GAN performs excellently on both datasets (MBA: P = 0.9877, F1 = 0.9825, L = 28ms; XJTU: P = 0.9742, F1 = 0.9502, L = 24ms). Specifically: The multi-generator architecture is proven to be the most crucial factor in improving the model performance. When it is removed, the F1 score shows the most significant decline: 4.5% reduction in MBA and 8.35% reduction in XJTU; this verifies the core role of this module in capturing diverse data patterns. The dynamic hybrid attention mechanism demonstrates special false alarm suppression ability. Its absence will lead to a significant reduction in precision: 6.31% decrease in MBA and 5.18% decrease in XJTU; but the F1 score still remains at a relatively high level, indicating that this module mainly focuses on improving precision rather than overall balance.
[0078] The model-agnostic meta-learning (MAML) component shows dual advantages: not only can it improve the accuracy, but also significantly optimize the training efficiency. When this module is removed, the latency time increases significantly: 53.6% increase in MBA and 41.7% increase in XJTU; at the same time, there is a moderate decrease in the F1 score, which fully reflects its key role in accelerating the model convergence (especially in cold start scenarios). The gradient penalty mechanism mainly contributes to the model stability. Its absence will lead to a consistent but relatively small performance decline in all metrics: 1.83% - 2.82% reduction in the F1 score; and the latency time remains close to the original level.
[0079] These findings together indicate that although the components contribute to the performance of VT-GAN in different ways, the multi-generator architecture lays the foundation for the model's effectiveness, dynamic attention focuses on precision optimization, MAML ensures training efficiency, and gradient penalty ensures the stability of the optimization process. The synergistic effect of these modules ultimately enables VT-GAN to achieve the current optimal performance level.
[0080] Table 2 Ablation Experiments: Comparison of Precision (P), F1-Score (F1), and Latency (L) between VT-GAN and Its Ablated Versions
[0081] In summary, for the problem of multivariate time series anomaly detection in industrial equipment health management, this invention proposes a model VT-GAN based on the deep fusion of variable-time Transformer (VTT) and generative adversarial network (GAN), and achieves fast cross-device adaptation through the model-agnostic meta-learning (MAML) framework. Through systematic theoretical analysis and experimental verification, the dynamic hybrid attention mechanism explicitly models spatio-temporal interaction relationships through the gated fusion of temporal self-attention, variable self-attention, and cross-attention layers, reducing the false alarm rate to 0.07% and 0.09% on the SWaT and SMD datasets respectively. The multi-generator adversarial architecture designs a parallel generator group, combined with dilated causal convolution to cover multi-time scale patterns, increasing the sample diversity by 23.4% and the F1-score by 12.7% on the XJTU dataset. By temporal task meta-ization and bidirectional LSTM alignment loss, the time shift problem is solved, and it can converge with only 10 gradient updates in the cold start scenario. The current MAML framework relies on the similarity of in-device task distributions, and in the future, cross-device transfer strategies based on graph neural networks will be explored to handle larger-scale heterogeneous device networks.
[0082] Embodiment 2 This embodiment provides a device for implementing a method for anomaly detection of time series data based on a variable-time transformer, the device includes: At least one processor; And a memory, the memory stores instructions, when the instructions are executed by the at least one processor, enabling the at least one processor to execute a method for anomaly detection of time series data based on a variable-time transformer as described above.
[0083] In this embodiment, the electronic device includes but is not limited to: personal computer, server computer, workstation, desktop computer, laptop computer, notebook computer, mobile computing device, smart phone, tablet computer, cellular phone, personal digital assistant (PDA), handheld device, messaging device, wearable computing device, consumer electronic device, etc.
[0084] Embodiment 3 This embodiment also provides a computer-readable storage medium storing executable instructions that, when executed, cause the machine to perform a method for anomaly detection of time series data based on a variable time transformer as described above.
[0085] Specifically, a system or device equipped with a readable storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer or processor of the system or device is caused to read and execute the instructions stored in the readable storage medium.
[0086] In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments, so the computer-readable code and the readable storage medium storing the computer-readable code constitute a part of this specification.
[0087] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer or a cloud via a communication network.
[0088] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) including computer-usable program code.
[0089] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0090] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.
[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.
[0092] Obviously, the above-described embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific embodiments of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. A method for detecting anomalies in time series data based on a variable time converter, characterized in that: The method comprises: S1. Obtained from industrial equipment time steps, The observation data is composed of sensor variables, and preprocessed to obtain real sample data; S2. Build the VT-GAN model, including the generator and the discriminator, and connect them through the adversarial training framework. The interaction process includes two parts: data flow and gradient back propagation. S3. Optimize the loss function of the VT-GAN model by adjusting the temperature parameters and perform data anomaly detection based on the constructed VT-GAN model.
2. The method for detecting anomalies in time series data based on a variable time converter according to claim 1, characterized in that: Before inputting into the VT-GAN model, the observed data is preprocessed to obtain real sample data, including: The support set and query set of observation data come from different stages of the device life cycle. The multivariate time series data is , a total of N devices, each sample include time step dimensional sensor readings, c represents the current moment, represents the set of real numbers; Support Set It represents the data of E continuous time windows sampled from the early operation stage of the device, e represents the time window from the current moment to the past, that is, the early operation stage, a total of E windows, which are used to construct the support set; Query Set Indicates sampling from the later operation stage of the same device The data of continuous time windows, m represents the time window extending from the current moment to the future, that is, the time window of the later operation stage, a total of M windows are used to construct the query set; There is a time offset between the query set and the support set, so the bidirectional LSTM, or BiLSTM, alignment loss is introduced: (1); In formula (1), is the timing alignment loss, Represent the feature vectors of the current moment c in the support set and query set respectively; After preprocessing, the observed data finally obtains the real sample data , forming the input data set.
3. A time series data anomaly detection method based on a variable time converter according to claim 1 or 2, characterized in that: The S2 specifically includes: S21. Construct the generator of the VT-GAN model, including a short-term generator, a medium-term generator, and a long-term generator; Input noise vector , d is the expansion factor, and the hourly time series patterns are generated by the short-term generator ; Generate daily trend patterns through the medium-term generator ; Generate weekly cycle patterns through long-term generators ; Finally, the gated fusion module outputs the fused fake samples ; S22, constructing the discriminator of the VT-GAN model, including an embedding layer, a temporal self-attention module, a variable self-attention module, a cross-attention module, and a gated fusion module; For the discriminator, input the real sample and generate samples , the input is mapped into high-dimensional features through the embedding layer ; The temporal self-attention module is used to capture the temporal dependency of single variables, the variable self-attention module is used to model the cross-variable interaction, and the cross-attention module is used to explicitly fuse the spatiotemporal features; Finally, the outputs of the temporal self-attention module, the variable self-attention module, and the cross-attention module are dynamically weighted by the gated fusion module to generate discriminative features. , output the discriminant probability according to the discriminant features.
4. The method for detecting anomalies in time series data based on a variable time converter according to claim 3, characterized in that: The short-term generator G1 is composed of 3 layers of dilated causal convolution stacks, each with a convolution kernel size of , the expansion factor is Incremental, is the number of layers, and the receptive field gradually expands to time steps; The mid-term generator G2 consists of 6 layers of dilated causal convolution stacks, each with a convolution kernel size of , the expansion factors of the first three layers are , the dilation factor of the last three layers is fixed to 8, and the receptive field is 63 steps; The long-term generator G3 uses 12 layers of dilated causal convolutions with residual connections, and the size of each convolution kernel is , each layer expansion factor ; The core operation of each generator is implemented by dilated causal convolution, and the processing function is as follows: (2); In formula (2), represents the sth generator The output features of the layer, is a one-dimensional causal convolution, is the latent space dimension.
5. The method for detecting anomalies in time series data based on a variable time converter according to claim 4, characterized in that: Output the fused fake samples through the gated fusion module , including: In order to integrate the outputs of multiple generators and generate the final synthetic samples, a dynamic gating fusion mechanism is proposed: (3); In formula (3), For the The output of the generator is the generated sample, is the hidden state of each generator, The meaning of is the dimension of the generator hidden state, that is, the length of the hidden feature vector of each generator, extracted from the last layer of convolutional features; is a learnable parameter matrix that dynamically calculates the fusion weights of each branch , through Softmax normalization to ensure that the weights meet .
6. The method for detecting anomalies in time series data based on a variable time converter according to claim 5, characterized in that: The method captures the temporal dependency of a single variable through temporal self-attention, models cross-variable interactions through variable self-attention, and explicitly fuses spatiotemporal features through cross-attention, specifically including: The temporal self-attention module first embeds the features Divide by variable dimension Independent time series , , for each variable , calculate the query matrix , the bond matrix , value matrix : (4); In formula (4), , , are the learnable projection matrices for query, key, and value, respectively. is the unified dimension of Q / K / V vectors in temporal self-attention; Then calculate the temporal self-attention output : (5); The temporal self-attention module outputs the concatenation results along the variable dimension: (6); The variable self-attention module performs , introducing a learnable variable correlation matrix , enhanced prior knowledge guidance: (7); In formula (7), , , is the learnable projection matrix; Then calculate the variable self-attention output: (8); The variable self-attention module outputs the concatenation result along the time dimension: (9); Cross-attention layer outputs temporal self-attention As query source, variable self-attention output As a key-value source: (10); In formula (10), , , is the projection matrix; Then, calculate the cross attention output : (11); Dynamically integrate multimodal attention features and introduce learnable gating weights to generate discriminative features : (12); (13); In formulas (12)~(13), is the gate parameter matrix, which is ensured by Softmax normalization. =1, is the temporal attention weight, is the variable attention weight, is the cross attention weight.
7. The method for detecting anomalies in time series data based on a variable time converter according to claim 6, characterized in that: Output the discriminant probability according to the discriminant features, including: Fusion features Apply dilated causal convolution to enhance local feature extraction: (14); In formula (14), The output feature map after the dilated causal convolution processing fusion feature Z is the dilated convolution feature. is the hidden feature dimension, which indicates the final encoding length of each time step-variable pair; Compress features along the time and variable dimensions, from dilated convolution features The global feature vector extracted from : (15); Output discriminant probability : (16); In formula (16), is the Sigmoid function, , is a learnable parameter, is the learnable parameter of the fully connected layer; On the other hand, the discriminant probability It is passed to the generator, so that the generator generates data that is closer to the real data, and the discriminator then distinguishes the true from the false and conducts adversarial training.
8. The method for detecting anomalies in time series data based on a variable time converter according to claim 7, characterized in that: Optimize the loss function of the VT-GAN model, including: Contrastive adversarial loss introduces a contrastive learning mechanism, combined with temperature parameter adjustment and abnormal sample dynamic mining strategy. The loss function formula is as follows: (17); In formula (17), For real samples, That is, generate samples, is a real abnormal sample, is the discriminant probability of the generated sample, is the discrimination probability of abnormal samples, that is, the abnormal probability; is the temperature parameter, which controls the steepness of the probability distribution.
9. A device for implementing a time series data anomaly detection method based on a variable time converter, characterized in that: The device comprises: processor; and a memory having stored thereon a computer program executable on said processor; Wherein, when the computer program is executed by the processor, the steps of a method for detecting anomalies in time series data based on a variable time converter as described in any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Generative adversarial network-based urban local carbon emission hotspot prediction and regulation method
CN118982155A
Transform-based multivariable time sequence anomaly detection method
CN116796272A
Multi-dimensional time sequence anomaly detection method based on time variable double-attention mechanism
CN118378139A
Industrial control time series data anomaly detection method
CN119916792A
Data anomaly detection method and apparatus
WO2023123941A1
Cited By
Line loss data anomaly monitoring method and system based on time sequence anomaly detection
CN120354327A
Agricultural product quality traceability anomaly detection method
CN120410342A
A method for detecting abnormalities in agricultural product quality traceability
CN120410342B
Data processing method, device and system and computer readable storage medium
CN120449947A
Industrial time sequence event analysis method and device based on causal regularization and medium
CN120654104A