Anomaly Detection Method for Time Series Data Based on Variable Time Transformer

By building a VT-GAN model based on a variable time converter, the real-time and noise sensitivity problems of traditional models in industrial scenarios are solved, efficient multi-time series anomaly detection is achieved, and the coverage and real-time response capabilities of anomaly detection are improved.

CN120144930BActive Publication Date: 2025-07-22QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510623076.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-22
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

Traditional models face the problems of insufficient real-time, strong noise sensitivity and limited diversity of abnormal patterns in industrial scenarios, which are difficult to meet the millisecond response requirements. Moreover, the calculation complexity of the global attention mechanism in high-dimensional timing data limits real-time, with high false alarm rate and low abnormal detection coverage.

Method used

The time series data anomaly detection method based on a variable time converter is adopted, and the VT-GAN model is constructed, and the space-time interaction relationship is explicitly modeled using the multi-generator architecture, time self-attention, variable self-attention and cross-attention layers, and the loss function is optimized to improve detection performance.

Benefits of technology

Effectively alleviate the problem of mode crash, improve sample diversity, achieve rapid parameter adaptation, reduce false alarm rate, improve abnormal detection coverage, and meet the real-time detection needs of industrial equipment health management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144930B_ABST
    Figure CN120144930B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of time series data processing. More specifically, it relates to a method for detecting anomalies in time series data based on a variable time transformer. The method proposes an anomaly detection model VT-GAN, which designs a parallel generator group. Each generator extracts pattern features of different time scales through dilated causal convolution. In the VTT architecture, temporal self-attention, variable self-attention, and cross-attention layers are fused to explicitly model spatio-temporal interaction relationships through learnable gating weights. The loss function of the VT-GAN model is adjusted and optimized by combining temperature parameters, and data anomaly detection is carried out based on the constructed VT-GAN model. The present invention solves the problems that traditional models are difficult to meet the millisecond-level response requirements due to the sequence calculation characteristics, the computational complexity of the global attention mechanism in high-dimensional time series data limits the real-time performance, and the false alarm rate is high and the anomaly detection coverage rate is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of time series data processing, and more specifically, relates to a method for detecting anomalies in time series data based on a variable time transformer. Background Art

[0002] In the context of industry and intelligent manufacturing, production equipment is widely deployed with sensors to achieve real-time monitoring of the operating state. The multivariate time series data collected by these sensors naturally carry key information about the health status and performance parameters of the equipment. Although anomaly detection techniques based on multivariate time series can effectively achieve early fault warning and remaining useful life prediction by analyzing multi-source sensor data such as vibration, temperature, and pressure, existing methods face three core challenges in industrial scenarios: (1) complex temporal dependencies and variable coupling effects; (2) high noise interference caused by sensor errors and communication delays; (3) scarcity of labeled anomaly data in dynamic environments.

[0003] Chinese patent document CN118982155A discloses a method for predicting and regulating urban local carbon emission hotspots based on a generative adversarial network, including the following steps: S1, data collection and preprocessing; S2, constructing a short-term memory module using hourly carbon emission data for short-term spatio-temporal feature encoding; S3, based on daily and weekly carbon emission data, using a multi-scale convolutional network to extract periodic change features for medium-term spatio-temporal feature encoding; S4, using a long temporal dependence model to analyze monthly and quarterly carbon emission data to obtain long-term spatio-temporal features for global encoding of the overall urban carbon emission distribution; S5, constructing and training a generative adversarial network; S6, using a spatio-temporal attention mechanism to perform weighted fusion of spatio-temporal features at different levels, and combining real-time data to generate a regulation strategy; S7, monitoring the regulation effect in real time and performing adaptive adjustment and optimization.

[0004] With the performance degradation caused by the aging of industrial equipment, multivariate time series anomaly detection is crucial for realizing equipment health management (PHM) and preventive maintenance. However, existing methods face challenges of insufficient real-time performance, strong noise sensitivity, and limited anomaly pattern diversity in complex industrial scenarios; traditional deep learning models are difficult to fully address the above challenges. The recurrent neural network RNN has the defect of insufficient parallelization ability, while methods based on Transformer usually cannot model variable-specific temporal patterns. Although the generative adversarial network GAN performs excellently in data generation, it faces problems of mode collapse and gradient instability in high-dimensional time series scenarios. In addition, traditional meta-learning methods lack an explicit mechanism for aligning multivariate time series features, as well as a high false alarm rate and low anomaly detection coverage. Summary of the Invention

[0005] The present invention aims to overcome at least one defect of the above-mentioned prior art, and provides a method for detecting anomalies in time series data based on a variable time transformer, so as to solve the problems that the traditional model is difficult to meet the millisecond-level response requirement due to the sequence calculation characteristics, the calculation complexity of the global attention mechanism in high-dimensional time series data limits the real-time performance, and the false alarm rate is high and the anomaly detection coverage rate is low.

[0006] The detailed technical solution of the present invention is as follows:

[0007] A method for detecting anomalies in time series data based on a variable time transformer, the method includes:

[0008] S1. Obtain time steps, observation data composed of

[0009] sensor variables from industrial equipment, and perform preprocessing to obtain real sample data;

[0010] S2. Construct a VT-GAN model, including a generator and a discriminator, and connect them through an adversarial training framework. The interaction process includes two parts: data flow and gradient backpropagation;

[0011] S3. Optimize the loss function of the VT-GAN model and perform data anomaly detection based on the constructed VT-GAN model.

[0011] Further, preprocess the observation data before inputting it into the VT-GAN model. Specifically, it includes:

[0012] The support set and query set of the observation data come from different stages of the equipment life cycle. The multivariate time series data of the equipment is , with a total of N devices. Each sample includes time steps of dimensional sensor readings. c represents the current moment, represents the set of real numbers;

[0013] The support set represents the data sampled from E consecutive time windows in the early operation stage of the equipment. e represents the time window in the early operation stage that traces back from the current moment to the past, with a total of E windows, which is used to construct the support set;

[0014] The query set represents the data sampled from consecutive time windows in the later operation stage of the same equipment. m represents the time window in the later operation stage that extends from the current moment to the future, with a total of M windows, which is used to construct the query set;

[0015] There is a time offset between the query set and the support set. Introduce a bidirectional LSTM, that is, BiLSTM alignment loss:

[0016] (1);

[0017] In formula (1), is the temporal alignment loss, respectively represent the feature vectors of the support set and the query set at the current moment c;

[0018] The observed data is finally preprocessed to obtain real sample data , which constitutes the input data set.

[0019] Furthermore, the specific steps of S2 include:

[0020] S21. Construct the generator of the VT-GAN model, including a short-term generator, a medium-term generator, and a long-term generator;

[0021] Input the noise vector , and generate the hourly time series pattern through the short-term generator; generate the daily trend pattern through the medium-term generator; generate the weekly cycle pattern through the long-term generator;

[0022] Finally, output the fused fake samples .

[0023] S22. Construct the discriminator of the VT-GAN model, including an embedding layer, a temporal self-attention module, a variable self-attention module, a cross-attention module, and a gated fusion module;

[0024] For the discriminator, input the real sample and the generated sample , and map the input to high-dimensional features through the embedding layer; capture the univariate time series dependencies through the temporal self-attention module, model the cross-variable interactions through the variable self-attention module, and explicitly fuse the spatio-temporal features through the cross-attention module;

[0025] Finally, dynamically weight the outputs of the temporal self-attention module, the variable self-attention module, and the cross-attention module through the gated fusion module to generate discriminative features , and output the discriminative probability according to the discriminative features.

[0026] Furthermore, the short-term generator G1 is stacked by 3 layers of dilated causal convolution, and the size of each convolution kernel is , and the dilation factor is increasing, is the number of layers, and the receptive field is gradually expanded to time steps;

[0027] The medium-term generator G2 consists of 6 layers of dilated causal convolution stacked, with the convolution kernel size of each layer being , the dilation factor of the first 3 layers being , and the dilation factor of the last 3 layers being fixed at , with a receptive field of 63 steps;

[0028] The long-term generator G3 uses 12 layers of dilated causal convolution with residual connections, and the convolution kernel size of each layer is , and the dilation factor of each layer is ;

[0029] The core operation of each generator is implemented by dilated causal convolution, and the processing function is as follows:

[0030] (2);

[0031] In formula (2), represents the output feature of the th layer of the sth generator, is a one-dimensional causal convolution, is the latent space dimension, represents the output feature of the -1th layer of the sth generator.

[0032] Furthermore, the fused fake samples are output through the gated fusion module, specifically including:

[0033] To integrate the outputs of multiple generators and generate the final synthetic samples, a dynamic gated fusion mechanism is proposed:

[0034] (3);

[0035] In formula (3), is the output of the th generator, is the hidden state of each generator, means the dimension of the generator hidden state, that is, the length of the hidden feature vector of each generator, which is extracted from the last layer of convolutional features; is a learnable parameter matrix, dynamically calculating the fusion weights of each branch, and ensuring that the weights satisfy through Softmax normalization.

[0036] Furthermore, the time self-attention is used to capture the univariate time series dependencies respectively, the variable self-attention is used to model the cross-variable interactions, and the cross-attention is used to explicitly fuse the spatio-temporal features, specifically including:

[0037] The time self-attention module first splits the embedded feature along the variable dimension into a single independent time series , , for each variable , calculate the query matrix , the key matrix , and the value matrix :

[0038] (4);

[0039] In formula (4), , , is a learnable projection matrix;

[0040] Then calculate the output of the temporal self-attention , is the unified dimension of the Q / K / V vectors in the temporal self-attention:

[0041] (5);

[0042] The concatenation result of the output of the temporal self-attention module along the variable dimension:

[0043] (6);

[0044] The variable self-attention module for each time step , introduces a learnable variable correlation matrix , to enhance the prior knowledge guidance:

[0045] (7);

[0046] where , , is a learnable projection matrix;

[0047] Then calculate the output of the variable self-attention:

[0048] (8);

[0049] The concatenation result of the output of the variable self-attention module along the time dimension:

[0050] (9);

[0051] The cross-attention layer uses the output of the temporal self-attention as the query source and the output of the variable self-attention as the key-value source:

[0052] (10);

[0053] Among them, , , is the projection matrix.

[0054] Then, calculate the cross-attention output :

[0055] (11);

[0056] Dynamically integrate multi-modal attention features, introduce learnable gating weights to generate discriminant features :

[0057] (12);

[0058] (13);

[0059] In formulas (12) to (13), is the gating parameter matrix, and through Softmax normalization, ensure that = 1, is the temporal attention weight, is the variable attention weight, is the cross-attention weight.

[0060] Furthermore, according to the discriminant features, output the discriminant probability, specifically including:

[0061] Apply dilated causal convolution to the fused feature to enhance local feature extraction:

[0062] (14);

[0063] In formula (14), is the output feature map after processing the fused feature Z by dilated causal convolution, that is, the dilated convolution feature, is the hidden layer feature dimension, representing the final encoding length of each time step-variable pair;

[0064] Compress the features along the time and variable dimensions, and extract the global feature vector from the dilated convolution feature :

[0065] (15);

[0066] Output the discriminant probability :

[0067] (16);

[0068] In formula (16), is the Sigmoid function, , are learnable parameters, are the learnable parameters of the fully connected layer;

[0069] On the other hand, the discrimination probability is passed to the generator, enabling the generator to generate data closer to the real data. Then the discriminator discriminates between true and false for adversarial training.

[0070] Furthermore, the loss function of the VT-GAN model is optimized, specifically including:

[0071] The contrastive adversarial loss introduces a contrastive learning mechanism and further combines temperature parameter adjustment and dynamic abnormal sample mining strategy. The loss function formula is as follows:

[0072] (17);

[0073] In formula (17), is the real sample, i.e., the generated sample, is the real abnormal sample, is the discrimination probability of the generated sample, is the discrimination probability of the abnormal sample, i.e., the abnormal probability; is the temperature parameter, which controls the steepness of the probability distribution.

[0074] In another aspect of the present invention, there is provided an apparatus for an anomaly detection method of time series data based on a variable time transformer, and the apparatus includes:

[0075] At least one processor;

[0076] And a memory, the memory stores instructions, and when the instructions are executed by the at least one processor, the at least one processor executes an anomaly detection method of time series data based on a variable time transformer as described above.

[0077] In another aspect of the present invention, there is also provided a computer-readable storage medium, which stores executable instructions, and when the instructions are executed, the machine executes an anomaly detection method of time series data based on a variable time transformer as described above.

[0078] Compared with the prior art, the beneficial effects of the present invention are:

[0079] (1) A method for anomaly detection of time series data based on a variable time transformer provided by the present invention adopts a GAN variant enhanced by a residual network, effectively alleviates the mode collapse problem and improves the sample diversity through a parallel sub-generator and gradient penalty constraints; time self-attention, variable self-attention and cross-attention layers are introduced into the VTT, and spatio-temporal interactions are explicitly modeled through learnable gating weights, solving the defect of insufficient modeling of variable coupling effects by traditional methods.

[0080] (2) A method for anomaly detection of time series data based on a variable time transformer provided by the present invention is based on a model-agnostic meta-learning framework to achieve fast parameter adaptation of the LSTM network in small-sample anomaly scenarios (<10-step gradient updates), and the training time is significantly reduced compared with the standard meta-learning method. Brief Description of the Drawings

[0081] Figure 1 is a flowchart of the method for anomaly detection of time series data based on a variable time transformer described in the present invention.

[0082] Figure 2 is a schematic diagram of the VT-GAN model in Embodiment 1 of the present invention.

[0083] Figure 3 is a schematic diagram of the generator part in Embodiment 1 of the present invention.

[0084] Figure 4 is a schematic diagram of the discriminator part in Embodiment 1 of the present invention.

[0085] Figure 5 is a schematic diagram of the time series anomaly detection results of three dimensions on the SMD dataset in Embodiment 1 of the present invention.

[0086] Figure 6 is a schematic diagram of the downward trend of the average training loss of the VT-GAN model on the SWaT dataset in Embodiment 1 of the present invention. Detailed Description of the Embodiments

[0087] The present invention will be further described below in conjunction with the drawings and embodiments.

[0088] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0089] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "comprise" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0090] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0091] Embodiment 1

[0092] Refer Figure 1 , this embodiment provides a method for detecting anomalies in time series data based on a variable time converter, and the method includes:

[0093] S1. Obtain observation data and perform preprocessing: Obtain time steps, observation data composed of sensor variables, and perform preprocessing to obtain real sample data. The health status monitoring of industrial equipment depends on the anomaly detection of multivariate time series data (Multivariate Time Series, MTS);

[0094] The preprocessing can be methods such as handling missing values and data standardization. Preferably, in this embodiment, the observation data is preprocessed before inputting into the VT-GAN model, specifically including:

[0095] In the existing work situation, there is a situation where the time series patterns do not match due to the support set and the query set coming from different stages of the equipment life cycle. Therefore, diversified time series tasks are proposed:

[0096] The support set and the query set of the observation data come from different stages of the equipment life cycle. The multivariate time series data of any equipment is , there are a total of N devices, and each sample includes sensor readings of dimensions for time steps. c represents the current moment,

[0097] Support set represents the data sampled from E consecutive time windows in the early operation stage of the device. e represents the time window in the early operation stage that traces back from the current moment to the past, with a total of E windows, which are used to construct the support set;

[0098] Query set Denote the data sampled from the later operation stage of the same device from \(m\) consecutive time windows, where \(m\) represents the time window extending from the current moment to the future, i.e., the later operation stage, and there are a total of \(M\) windows, which are used to construct the query set;

[0099] There is a time offset between the query set and the support set. To address the distribution differences caused by the time offset, a bidirectional LSTM, i.e., BiLSTM alignment loss, is introduced. The BiLSTM alignment loss acts on the feature space inside the VT-GAN model instead of directly modifying the input data:

[0100] (1);

[0101] In formula (1), is the temporal alignment loss, respectively represent the feature vectors at the current moment \(c\) in the support set and the query set. This loss forces the model to learn the common temporal patterns across time periods and alleviates the impact of phase offset;

[0102] The observed data is finally processed to obtain the real sample data , which constitutes the input data set.

[0103] S2. Construct the VT-GAN model. The VT-GAN model includes a generator (Generator) and a discriminator (Discriminator), and is connected through an adversarial training framework. The interaction process includes two parts: data flow and gradient backpropagation. The VT-GAN model is as Figure 2 shown. The generator generates time series data from the latent space, and the discriminator can determine whether these data come from the real time series data set or are generated by the generator.

[0104] S21. The generator of the constructed VT-GAN model includes: a short-term generator, a medium-term generator, a long-term generator, and a gated fusion module:

[0105] For the generator, as Figure 3 shown, the input noise vector , where \(d\) is the dilation factor:

[0106] The short-term generator generates hourly temporal patterns through dilated causal convolutions ; the medium-term generator generates daily trend patterns through dilated causal convolutions ; the long-term generator generates weekly periodic patterns through dilated causal convolutions , and finally outputs the fused fake samples .

[0107] In order to capture the dynamic patterns of different time scales in the degradation process of industrial equipment, this work designs a multi-generator architecture, each generator focuses on feature extraction and generation at a specific time scale. The multi-generator architecture is shown in the figure Figure 3 As shown in the figure. Through the hierarchical expansion of the expansion factor, the short-term generator has local fluctuations, the medium-term generator models the daily trend, and the long-term generator captures the weekly cycle, covering the entire process of device degradation. By expanding the local connection characteristics of the causal convolution, the gradient vanishing problem in the traditional GAN is alleviated, and the training convergence is further guaranteed by combining the Wasserstein loss.

[0108] Specifically, the short-term generator G1 is composed of three layers of dilated causal convolution stacks, and the size of each convolution kernel is , the expansion factor is Incremental, is the number of layers, and the receptive field gradually expands to The short-term generator is used to capture high-frequency changes in the device state, such as sudden temperature rise and pressure peak, and quickly respond to local features through a shallow convolutional network.

[0109] The mid-term generator G2 consists of 6 layers of dilated causal convolution stacks, each with a kernel size of , the expansion factors of the first three layers are ,The expansion factor of the last three layers is fixed to 8, and the receptive field is 63 steps, ensuring the stable modeling of the day-level patterns.,The medium-term generator is used to identify the progressive degradation of device,performance, and balance local features with medium- and long-term dependencies through a medium-depth,network structure.

[0110] The long-term generator G3 uses 12 layers of dilated causal convolution with residual connection, such as Figure 3 As shown in Figure 2, it includes fully connected layers and convolutional layers connected by three residual blocks, and the size of each convolution kernel is , each layer expansion factor , the final receptive field covers The long-term generator is used to analyze the macroscopic evolution law of the equipment throughout its life cycle and capture the complex correlations across time steps through a deep network.

[0111] Preferably, the core operation of each generator is implemented by dilated causal convolution, and the processing function is as follows:

[0112] (2);

[0113] In formula (2), represents the sth generator The output features of the layer. It is a one-dimensional causal convolution, which ensures that the output depends only on the current and historical inputs to avoid future information leakage. is the convolution kernel size, is the dilation factor, which exponentially increases with the network depth to gradually expand the receptive field. This design optimizes the adaptability to multi-scale temporal patterns by hierarchically and dynamically adjusting the dilation factor.

[0114] Preferably, the gated fusion network outputs the fused fake samples, specifically including:

[0115] To integrate the outputs of multiple generators and generate the final synthetic samples, a dynamic gated fusion mechanism is proposed:

[0116] (3);

[0117] In formula (3), is the output of the th generator, is the hidden state of each generator, The meaning of is the dimension of the generator hidden state, i.e., the length of the hidden feature vector of each generator, which is extracted from the last layer of convolutional features through global average pooling;

[0118] is a learnable parameter matrix that dynamically calculates the fusion weights of each branch .

[0119] A gated network is introduced for the characteristics of temporal data, and Softmax normalization is used to ensure that the weights satisfy .

[0120] S22. The discriminator of the constructed VT-GAN model includes: an embedding layer, a temporal self-attention module, a variable self-attention module, cross-attention, and a gated fusion module;

[0121] For the discriminator, as Figure 4 shown, the input real sample and the generated sample , are mapped to high-dimensional features through the embedding layer; the temporal self-attention is used to capture the univariate temporal dependencies respectively, the variable self-attention is used to model the cross-variable interactions, and the cross-attention is used to explicitly fuse the spatio-temporal features;

[0122] Finally, the gated fusion module dynamically weights the outputs of the temporal self-attention module, the variable self-attention module, and the cross-attention module to generate the discriminant features , and the discriminant probability is output according to the discriminant features.

[0123] Preferably, the discriminator uses a dynamic hybrid attention discriminator, as Figure 4 shown;

[0124] The dynamic hybrid attention discriminator parallelly models temporal, variable, and spatio-temporal cross-dependencies through three dedicated components, and dynamically fuses multi-modal features with learnable gating weights, significantly enhancing the adaptability of the model to complex working conditions.

[0125] For the dynamic hybrid attention discriminator, the input real samples and the generated samples are mapped into high-dimensional features through the embedding layer ; the temporal self-attention is respectively used to capture the single-variable time-series dependencies, the variable self-attention is used to model the cross-variable interactions, the cross-attention is used to explicitly fuse the spatio-temporal features, and finally the outputs of each attention are dynamically weighted to generate discriminative features and the anomaly score is calculated.

[0126] Preferably, to extract the deep spatio-temporal features of the multivariate time series, the multivariate time series is first processed by the input embedding layer to map the original input into a high-dimensional latent space, the learnable parameter matrix is used to enhance the feature separability, and a structured input is provided for the subsequent attention mechanism. This embedding-attention combined architecture effectively captures the deep spatio-temporal pattern features while maintaining the computational efficiency:

[0127] The discriminator maps the input into a high-dimensional hidden space through the embedding layer to obtain the embedding features :

[0128] (18);

[0129] In formula (18), and are the learnable parameters of the embedding layer, is the dimension of the hidden space, which is usually set to 64 or 128 to balance the computational efficiency and the representation ability. This embedding process enhances the feature separability through linear transformation and provides a structured input for the subsequent attention mechanism.

[0130] The dynamic hybrid attention module consists of three parallel parts, which respectively model: temporal, variable, and spatio-temporal cross-dependencies.

[0131] The goal of the temporal self-attention is to capture the long-term dependencies across time steps within a single variable. First, the embedding features are sliced into independent time series by the variable dimension, where , for each variable , the query matrix , the key matrix , and the value matrix are calculated:

[0132] (4);

[0133] In formula (4), , , is a learnable projection matrix, is the unified dimension of the Q / K / V vectors in the temporal self-attention.

[0134] Calculate the output of the temporal self-attention :

[0135] (5);

[0136] The concatenation result of the output of the temporal self-attention module along the variable dimension:

[0137] (6);

[0138] The goal of the Variable Self-Attention is to model the dynamic coupling relationship between different variables at the same time step. For each time step , introduce a learnable variable correlation matrix to enhance the prior knowledge guidance:

[0139] (7);

[0140] where , , is a learnable projection matrix.

[0141] Calculate the output of the variable self-attention:

[0142] (8);

[0143] The concatenation result of the output of the variable self-attention module along the time dimension:

[0144] (9);

[0145] The goal of the Cross-Attention Layer is to explicitly fuse the features of the time and variable dimensions and model the cross-time-space interaction pattern. Using the output of the temporal self-attention as the query source and the output of the variable self-attention as the key-value source:

[0146] (10);

[0147] where , , is the projection matrix.

[0148] Then, calculate the output of the cross-attention module :

[0149] (11);

[0150] Dynamically integrate multi-modal attention features, introduce learnable gating weights to generate fusion features, i.e., discriminative features :

[0151] (12);

[0152] (13);

[0153] In formulas (12) to (13), is the gating parameter matrix, and through Softmax normalization, ensure = 1, is the temporal attention weight, is the variable attention weight, is the cross-attention weight.

[0154] Through the discriminative feature the discriminant probability can be obtained. For example, directly perform discrimination through a fully connected layer; preferably, in this embodiment, first apply dilated causal convolution to the fusion feature to enhance local feature extraction:

[0155] (14);

[0156] In formula (14), is the output feature map of the dilated causal convolution processing the fusion feature Z, i.e., the dilated convolution feature, is the hidden layer feature dimension in the model, representing the final encoding length of each time step-variable pair.

[0157] Compress the features along the temporal and variable dimensions, and extract the global feature vector from the dilated convolution feature :

[0158] (15);

[0159] Output the discriminant probability :

[0160] (16);

[0161] In formula (16), is the Sigmoid function, , is the learnable parameter, Are the learnable parameters of the fully connected layer;

[0162] On the other hand, the discrimination probability is passed to the generator, enabling the generator to generate data closer to the real data. Then the discriminator discriminates between true and false to conduct adversarial training.

[0163] S3. Optimize the loss function and perform detection and discrimination: Combine the temperature parameter adjustment to optimize the loss function of the VT-GAN model and perform data anomaly detection based on the constructed VT-GAN model.

[0164] Specifically, optimizing the loss function of the VT-GAN model specifically includes:

[0165] Traditional adversarial loss only optimizes the discrimination probability of the generated samples and lacks explicit constraints on the distribution boundaries of positive and negative samples.

[0166] The contrastive adversarial loss, by introducing a contrastive learning mechanism, explicitly optimizes the discriminator's ability to distinguish normal samples, generated samples, and abnormal samples, solves the problem of blurred distribution boundaries in traditional adversarial training, and further combines the temperature parameter adjustment and the dynamic mining strategy of abnormal samples. This loss function significantly improves the anomaly detection performance of the model in industrial scenarios with few annotations and high noise. The loss function formula is as follows:

[0167] (17);

[0168] In formula (17) is the real sample, i.e., the generated sample, is the real abnormal sample or the sample with high reconstruction error (pseudo-negative sample), is the discrimination probability of the generated sample, is the discrimination probability of the abnormal sample, i.e., the anomaly probability; is the temperature parameter, default = 0.1, which controls the steepness of the probability distribution. By introducing , the model can adaptively mine potential abnormal patterns in the unsupervised scenario and alleviate the cold start problem in anomaly detection.

[0169] The experiments of this embodiment are proved as follows:

[0170] The VT-GAN model was compared with the state-of-the-art models for multivariate time series anomaly detection, including MERLIN, DAGMM, OmniAnomaly, MSCRED, USAD, TranAD. All models were trained using the PyTorch-1.7.1 library in this embodiment. At the same time, the AdamW optimizer was used with an initial learning rate of 0.01, a meta-learning rate of 0.02, and a step scheduler of 0.5 to train the models, including using the following hyperparameter values:

[0171] VTT: 4-layer encoder, 8-head attention, hidden layer dimension 256.

[0172] GAN: 3 generators, dilation factors are 1 / 8 / 16 respectively, gradient penalty coefficient λ = 10.

[0173] MAML: number of inner loop steps is 5, number of outer loop steps is 10, learning rate is 0.01.

[0174] To train the model, the training time series is divided into 80% training data and 20% validation data.

[0175] In this experiment, six publicly available datasets were used. To directly compare with existing methods, these widely used public datasets were selected.

[0176] 1) MIT-BIH Supraventricular Arrhythmia Database MBA: It is a collection of electrocardiogram records from four patients, including multiple instances of two different types of abnormalities.

[0177] 2) Soil Moisture Active Passive Dataset SMAP: It is a dataset of soil samples and telemetry information used by NASA with a Mars rover.

[0178] 3) Server Machine Dataset SMD: This is a five-week dataset of stacked traces of resource utilization from 28 machines in a computing cluster.

[0179] 4) Safe Water Treatment Dataset SWaT: This dataset was collected from a real water treatment plant operating normally for 7 days and abnormally for 4 days. The dataset consists of sensor values and actuator operations.

[0180] 5) Mars Science Laboratory Dataset MSL: It is a dataset similar to SMAP, but corresponding to the sensor and actuator data of the rover itself.

[0181] 6) XJTU-SY Bearing Dataset XJTU: This dataset includes complete run-to-failure data of 15 rolling bearings, which were obtained by conducting many accelerated degradation experiments.

[0182] In Table 1, P: Precision, R: Recall, L: Latency (milliseconds), F1: F1 score under complete training data. The P value and F1 value of the VT-GAN model are optimal on two datasets.

[0183] Table 1: Performance comparison between the present invention and baseline methods on two datasets MBA and XJTU

[0184]

[0185] In this embodiment, accuracy, recall rate, ROC / AUC, and F1 score are used as the core evaluation metrics, and training time (h) and the number of parameters (M) are used as auxiliary evaluation metrics.

[0186] Precision, recall rate, F1 score, and training latency (in milliseconds) are used as the core evaluation metrics, and training time (in hours) and the number of parameters (in millions) are used as auxiliary evaluation metrics.

[0187] Table 1 compares the performance of the present invention with multiple baseline methods on two datasets, MBA and XJTU. The evaluation metrics include precision (P), recall rate (R), latency time (L(ms)), and F1 score (F1):

[0188] From the results, VT-GAN shows the best overall performance, achieving the highest precision of 0.9846 and the highest F1 score of 0.9825 on the MBA dataset. At the same time, on the XJTU dataset, it also leads other methods with a precision of 0.9547 and an F1 score of 0.9477, and the latency time remains within 30 milliseconds.

[0189] In contrast, although MERLIN is the fastest model with a latency time of 5 ms (MBA) and 8 ms (XJTU), its recall rate performance is poor, especially only 0.4923 on the MBA dataset. Among other methods, DAGMM and MSCRED are prominent in terms of precision performance but have a higher latency (34 - 36 ms), while OmniAnomaly and USAD perform excellently in terms of recall rate (0.9477 - 0.9656).

[0190] It is worth noting that the two datasets show different characteristics: the precision rate span of the MBA dataset is relatively large (0.8561 - 0.9846), and the F1 score difference is significant (0.6564 - 0.9825); while the performance of each method on the XJTU dataset is more balanced, and the F1 score is concentrated in the range of 0.8055 - 0.9477.

[0191] Overall, VT-GAN demonstrates the best balance, achieving leading advantages in all key metrics of the two datasets while maintaining a low latency. The best precision and F1 scores are both obtained by VT-GAN.

[0192] Such as Figure 5As shown, the time series anomaly detection results of the present invention in three dimensions on the SMD dataset are presented: Taking the 19th dimension as an example, the horizontal axis represents the timestamp (from 0 to approximately 26,000), the vertical axis represents the data value range and the anomaly score (from 0 to 1), and the black curve represents the smoothed result of the predicted value; the red dashed line is the anomaly label; the blue translucent area is the time when the anomaly is marked; the green curve represents the anomaly score. Figure 5 They are the presentations of the 18th dimension, 19th dimension, and 20th dimension respectively. The upper sub - figure shows the time series of the original data, and it can be seen that there are significant fluctuations in the data at some time points; the lower sub - figure shows the anomaly scores of the corresponding time series, and the anomaly scores at some time points are extremely high, indicating that these points are identified as anomaly points. There is a significant anomaly peak at timestamp = 15,000. The blue shaded area covers this interval, and the anomaly score also reaches a peak here, confirming this anomaly point.

[0193] In these three dimensions, the anomaly at timestamp = 15,000 is successfully detected, and the anomaly score reaches a peak at this position, indicating that the model has consistency in the detection results of the same anomaly point in different dimensions. The predicted value, i.e., the red line, is close to the true value, i.e., the black line, at most time points, but there are significant deviations near the anomaly points, which helps the model identify anomaly points. The model can effectively identify the same anomaly points in different dimensions, indicating that it has good generalization ability and stability. The anomaly score can accurately capture the anomaly points and remain at a low level during normal periods, indicating that the model has strong anomaly detection ability.

[0194] By introducing a contrast learning mechanism, the contrastive adversarial loss significantly enhances the discriminator's ability to distinguish normal samples, generated samples, and abnormal samples, and solves the problem of blurred distribution boundaries in traditional adversarial training. Combining temperature scaling and dynamic abnormal sample mining strategies, this loss function significantly improves the anomaly detection performance in industrial scenarios with limited labels and high noise.

[0195] As Figure 6 shown, it presents the decreasing trend of the average training loss of the VT - GAN model on the SWaT dataset: The vertical axis represents the average training loss, which gradually decreases from the initial value of 0.05 to 0.02, indicating that the model loss steadily decreases during training and the parameters are continuously optimized; the horizontal axis represents the number of training epochs. Each epoch corresponds to a complete traversal of the dataset once, and a total of 4 epochs of training are carried out. Near the 4th epoch, the loss drops to 0.0072, indicating that the model obtains better performance in the later stage of training. The loss value decreases monotonically with the increase of epochs, indicating that the model of the present invention does not overfit and the optimization process is stable.

[0196] The ablation experiment is as follows:

[0197] Through systematic ablation experiments on the core components of the VT-GAN model, as shown in Table 2: The specific contributions of each module to the model performance were deeply analyzed. The experimental results show that the full-version VT-GAN performs excellently on both datasets (MBA: P = 0.9877, F1 = 0.9825, L = 28 ms; XJTU: P = 0.9742, F1 = 0.9502, L = 24 ms). Specifically:

[0198] The multi-generator architecture was proven to be the most crucial factor in improving the model performance. When it was removed, the most significant drop in the F1 score occurred: a 4.5% decrease in MBA and an 8.35% decrease in XJTU; this verified the core role of this module in capturing diverse data patterns. The dynamic hybrid attention mechanism demonstrated special false alarm suppression ability. Its absence would lead to a significant decrease in precision: a 6.31% decrease in MBA and a 5.18% decrease in XJTU; but the F1 score still remained at a relatively high level, indicating that this module mainly focused on improving precision rather than overall balance.

[0199] The model-agnostic meta-learning (MAML) component showed dual advantages: not only could it improve the accuracy but also significantly optimize the training efficiency. When this module was removed, the latency increased significantly: a 53.6% increase in MBA and a 41.7% increase in XJTU; meanwhile, there was a moderate decrease in the F1 score, which fully demonstrated its key role in accelerating model convergence (especially in cold start scenarios). The gradient penalty mechanism mainly contributed to the model stability. Its absence would lead to a consistent but relatively small performance drop in all metrics: a 1.83% - 2.82% decrease in the F1 score; meanwhile, the latency remained close to the original level.

[0200] These findings together indicate that although the ways in which each component contributes to the performance of VT-GAN are different, the multi-generator architecture laid the foundation for the model effectiveness, the dynamic attention focused on precision optimization, MAML guaranteed the training efficiency, and the gradient penalty ensured the stability of the optimization process. The synergistic effect of these modules ultimately enabled VT-GAN to reach the current optimal performance level.

[0201] Table 2 Ablation Experiments: Comparison of Precision (P), F1 Score (F1), and Latency (L) between VT-GAN and Its Ablated Versions

[0202]

[0203] In summary, for the problem of multivariate time series anomaly detection in industrial equipment health management, the present invention proposes a model VT-GAN based on the deep fusion of variable-time Transformer (VTT) and generative adversarial network (GAN), and realizes fast cross-device adaptation through the model-agnostic meta-learning (MAML) framework. Through systematic theoretical analysis and experimental verification, the dynamic hybrid attention mechanism explicitly models spatio-temporal interaction relationships through the gated fusion of time self-attention, variable self-attention, and cross-attention layers, reducing the false alarm rate to 0.07% and 0.09% on the SWaT and SMD datasets respectively. The multi-generator adversarial architecture designs a parallel generator group, combines dilated causal convolution to cover multi-time scale patterns, and the diversity of generated samples is increased by 23.4%, and the F1 score is increased by 12.7% on the XJTU dataset. By temporal task meta-ization and bidirectional LSTM alignment loss, the time shift problem is solved, and it can converge with only 10 gradient updates in the cold start scenario. The current MAML framework relies on the similarity of in-device task distributions, and in the future, cross-device migration strategies based on graph neural networks will be explored to handle larger-scale heterogeneous device networks.

[0204] Embodiment 2

[0205] This embodiment provides an apparatus for implementing a method for anomaly detection of time series data based on a variable-time transformer. The apparatus includes:

[0206] At least one processor;

[0207] And a memory that stores instructions, which when executed by the at least one processor, cause the at least one processor to execute a method for anomaly detection of time series data based on a variable-time transformer as described above.

[0208] In this embodiment, the electronic device includes, but is not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile computing devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable computing devices, consumer electronic devices, etc.

[0209] Embodiment 3

[0210] This embodiment also provides a computer-readable storage medium that stores executable instructions, which when executed cause the machine to execute a method for anomaly detection of time series data based on a variable-time transformer as described above.

[0211] Specifically, a system or device equipped with a readable storage medium can be provided. On this readable storage medium, software program codes for implementing the functions of any one of the above-mentioned embodiments are stored, and the computer or processor of the system or device is made to read and execute the instructions stored in the readable storage medium.

[0212] In this case, the program code read from the readable medium itself can implement the functions of any one of the above-mentioned embodiments. Therefore, the computer-readable code and the readable storage medium storing the computer-readable code constitute a part of this specification.

[0213] Examples of the readable storage medium include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer or a cloud via a communication network.

[0214] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that include computer-usable program codes.

[0215] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows Figure 1 or multiple flows and / or blocks

[0216] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more flows Figure 1 or multiple flows and / or blocks

[0217] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 steps for implementing the functions specified in one block or multiple blocks.

[0218] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for anomaly detection of time series data based on a variable time transformer, characterized in that, The method includes: S1. Obtain from industrial equipment time steps, observation data composed of sensor variables, and perform preprocessing to obtain real sample data; S2. Construct a VT-GAN model, including a generator and a discriminator, which are connected through an adversarial training framework. The interaction process includes two parts: data flow and gradient backpropagation; S3. Adjust and optimize the loss function of the VT-GAN model in combination with temperature parameters and perform data anomaly detection based on the constructed VT-GAN model; The specific content of S2 includes: S21. Construct the generator of the VT-GAN model, including a short-term generator, a medium-term generator, and a long-term generator; Input noise vector , where d is the dilation factor, and the hourly time series patterns are generated by the short-term generator respectively ; the daily trend patterns are generated by the medium-term generator ; the weekly cycle patterns are generated by the long-term generator ; Finally, the fused fake samples are output through the gating fusion module ; S22. Construct the discriminator of the VT-GAN model, including an embedding layer, a temporal self-attention module, a variable self-attention module, a cross-attention module, and a gated fusion module; For the discriminator, the real samples are input and the generated samples are mapped into high-dimensional features through the embedding layer ; the time self-attention module is used to capture univariate time series dependencies, the variable self-attention module is used to model cross-variable interactions, and the cross-attention module is used to explicitly fuse spatio-temporal features respectively; Finally, the gated fusion module dynamically weights the outputs of the temporal self-attention module, the variable self-attention module, and the cross-attention module to generate discriminative features. Based on the discriminative features, the discriminative probability is output. The short-term generator G1 is stacked by 3 layers of dilated causal convolutions, and the size of each convolutional kernel is , and the dilation factor is increasing, where is the number of layers, and the receptive field is gradually extended to The medium-term generator G2 includes a stack of 6 layers of dilated causal convolutions, and the size of the convolutional kernel for each layer is , and the dilation factors for the first 3 layers are , the dilation factors for the last 3 layers are fixed at 8, and the receptive field is 63 steps; The long-term generator G3 adopts 12-layer dilated causal convolution with residual connections, and the size of the convolution kernel for each layer is , and the dilation factor for each layer is ; The core operation of each generator is implemented by dilated causal convolution, and the processing function is as follows: (2); In formula (2), represents the output feature of the th layer of the sth generator, is a one-dimensional causal convolution, is the latent space dimension.

2. The time series data anomaly detection method based on a variable time converter according to claim 1, wherein, Before inputting into the VT-GAN model, preprocess the observed data to obtain real sample data, which specifically includes: The support set and query set of the observed data come from different stages of the device life cycle. The multivariate time series data of the device is , with a total of N devices. Each sample includes sensor readings of dimensions at time steps. Let c represent the current moment, and denote the set of real numbers; Support set It represents sampling data of E consecutive time windows from the early operation stage of the device. e represents the time window of the early operation stage by tracing back from the current moment to the past, with a total of E windows, which are used to construct the support set; Query set Indicates sampling data from the later operation stage of the same device The data of m consecutive time windows, where m represents the time window extending from the current moment to the future, i.e., the later operation stage, with a total of M windows, used to construct the query set; There is a time offset between the query set and the support set, and a bidirectional LSTM (i.e., BiLSTM) is introduced to align the loss: (1); In formula (1), is the temporal alignment loss, respectively represent the feature vectors of the support set and the query set at the current moment c; The observed data is finally preprocessed to obtain the true sample data , which constitutes the input data set.

3. The method for abnormal detection of time series data based on a variable time converter according to claim 2, wherein Output the fused fake samples through the gating fusion module , specifically including: To integrate the outputs of multiple generators and generate the final synthetic samples, a dynamic gated fusion mechanism is proposed: (3); In formula (3), is the output of the th generator, i.e., the generated sample, is the hidden state of each generator, means the dimension of the generator hidden state, i.e., the length of the hidden feature vector of each generator, which is extracted from the last-layer convolutional features; is a learnable parameter matrix that dynamically calculates the fusion weights of each branch , and ensures that the weights satisfy through Softmax normalization.

4. The method for detecting anomalies in time series data based on a variable time converter according to claim 3, wherein, The above respectively captures univariate time series dependencies through temporal self-attention, models cross-variable interactions through variable self-attention, and explicitly fuses spatio-temporal features through cross-attention, which specifically includes: The temporal self-attention module first embeds the features and splits them into independent time series according to the variable dimension , . For each variable , the query matrix , the key matrix , and the value matrix are calculated as follows: (4); In formula (4), , , are learnable projection matrices for query, key, and value respectively, is the unified dimension of the Q / K / V vectors in the temporal self-attention; Then calculate the output of the temporal self-attention : (5); The output of the temporal self-attention module is the concatenation result along the variable dimension: (6); The variable self-attention module is applied to each time step , introducing a learnable variable correlation matrix , enhancing the guidance of prior knowledge: (7); In formula (7), , , is a learnable projection matrix; Then calculate the output of the variable self-attention: (8); The output of the variable self-attention module is the concatenation result along the time dimension: (9); The cross-attention layer takes the temporal self-attention output as the query source, and the variable self-attention output as the key-value source: (10); In formula (10), , , is a projection matrix; Then, calculate the cross-attention output : (11); Dynamically integrate multi-modal attention features and introduce learnable gating weights to generate discriminative features : (12); (13); In Formulas (12) to (13), is a gating parameter matrix, and through Softmax normalization, it is ensured that = 1, is the temporal attention weight, is the variable attention weight, is the cross-attention weight.

5. A method for detecting anomalies in time series data based on a variable time converter according to claim 4, characterized in that Output the discrimination probability according to the discrimination features, which specifically includes: For the fused features Apply dilated causal convolution to enhance local feature extraction: (14); In formula (14), is the output feature map after dilated causal convolution processing of the fused feature Z, that is, the dilated convolution feature. is the hidden layer feature dimension, representing the final encoding length of each time step-variable pair; Compress features along the time and variable dimensions, and extract the global feature vector from the dilated convolutional features :​ (15); Output discrimination probability : (16); In formula (16), is the Sigmoid function, , are learnable parameters, are the learnable parameters of the fully connected layer; On the other hand, the discrimination probability is passed to the generator, enabling the generator to generate data closer to the real data. Then the discriminator discriminates between true and false to conduct adversarial training.

6. The time series data anomaly detection method based on a variable time converter according to claim 5, wherein, Optimize the loss function of the VT-GAN model, which specifically includes: Introduce a contrast learning mechanism into the contrastive adversarial loss, and combine the temperature parameter adjustment and the abnormal sample dynamic mining strategy. The formula of the loss function is as follows: (17); In formula (17), is a real sample, i.e., a generated sample, is a real abnormal sample, is the discrimination probability of the generated sample, is the discrimination probability of the abnormal sample, i.e., the abnormality probability; is the temperature parameter, which controls the steepness of the probability distribution.

7. An apparatus for implementing a method for detecting anomalies in time series data based on a variable time transformer, characterized in that, The device includes: A processor; And a memory, on which a computer program that can run on the processor is stored; Wherein, when the computer program is executed by the processor, it implements the steps of a method for detecting anomalies in time series data based on a variable time transformer as described in any one of claims 1 to 6.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Generative adversarial network-based urban local carbon emission hotspot prediction and regulation method

    CN118982155A

  • Transform-based multivariable time sequence anomaly detection method

    CN116796272A

  • Multi-dimensional time sequence anomaly detection method based on time variable double-attention mechanism

    CN118378139A