Mask generation-based interpretable time sequence anomaly detection method and system
Through the method of time-frequency domain conversion and mask generation modeling, the problems of high-dimensionality, noise interference and insufficient interpretability of industrial sensor data are solved, and efficient, accurate detection and visual interpretation of industrial time series anomalies are achieved.
Patent Information
- Application Number
- CN202510319315.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-22
AI Technical Summary
Existing time series anomaly detection methods face problems such as high dimensionality, noise interference, periodicity and mutation coexistence characteristics, and insufficient interpretability when processing industrial sensor data, resulting in missed detection and difficulty users to understand the causes of abnormalities.
The interpretability time series anomaly detection method based on mask generation is adopted, and the detection of time-frequency domain conversion, feature extraction, vectorization processing and mask generation modeling is constructed, the exception score is calculated and visualized results are generated, which enhances the robustness and interpretability of the detection.
It significantly improves the ability to capture industrial time series abnormal patterns, improves the robustness and interpretability of detection results, can intuitively identify and locate abnormal locations, and enhances user decision support capabilities.
Smart Images

Figure CN120354294A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of time series anomaly detection, and particularly to an interpretable time series anomaly detection method and system based on mask generation. Background Art
[0002] Time series anomaly detection has important application values in the industrial field and is widely used in equipment health monitoring, fault prediction, and production optimization. Existing anomaly detection methods mainly rely on statistical analysis, machine learning, or deep learning techniques.
[0003] During the industrial production process, equipment relies on sensors to collect key parameters such as temperature, pressure, and current to monitor the operating status and detect anomalies. The time series data generated by these sensors is crucial for equipment health management. However, the complexity of industrial sensor data poses many challenges to existing methods for anomaly detection.
[0004] Firstly, industrial sensor data usually has high dimensionality. This poses technical challenges of large data volume, high storage and computational overhead, and existing methods are difficult to process efficiently. For example, in the bearing monitoring system of a certain factory, the simultaneously collected signals may include vibration (acceleration), temperature, and rotational speed, etc. Assuming the vibration sensor records data at a sampling frequency of 1 kHz, the number of data points in a single channel within one minute can reach 60,000, and high-dimensional time series data will be formed after superimposing multiple sensors.
[0005] Secondly, industrial sensor data is vulnerable to noise interference and may have missing data. This data instability reduces the accuracy of anomaly detection, making it difficult for existing methods to effectively cope with. For example, in a steel plant with a high-temperature environment (>1000°C), temperature sensors may experience intermittent failures due to harsh environments, resulting in data interruptions or abnormal peaks.
[0006] In addition, industrial sensor data usually exhibits the coexistence of periodicity and mutability. Existing methods are difficult to simultaneously identify periodic patterns and abnormal mutations, resulting in false detections or missed detections. The operating states of many industrial devices are periodic. For example, the vibration signals of rotating machinery show stable patterns with changes in rotational speed. However, equipment anomalies (such as bearing damage, tool wear) may lead to mutations. For example, the vibration signal amplitude of a normally operating motor generally fluctuates between ±0.5g (the unit of gravitational acceleration), but when the bearing is damaged, the vibration amplitude may suddenly rise to ±5g or higher. Such mutations are usually early signs of equipment failures. If not detected in time, they may lead to equipment damage or even production accidents.
[0007] Finally, existing methods are insufficient in terms of the interpretability of industrial anomaly detection. Users often have difficulty understanding the specific reasons and nature of anomalies, which limits the widespread acceptance and use of anomaly detection technology in industrial practical applications. For example, when the vibration signal sampled at 1 kHz suddenly increases at a specific frequency (such as 200 Hz), the system can reconstruct and display the normal waveform at that frequency, intuitively presenting the anomaly location and its possible normal form to assist maintenance personnel in making accurate judgments. Summary of the Invention
[0008] To address the deficiencies of the prior art and achieve the purpose of efficient and accurate detection of time series data in the industrial field, avoiding false detections and missed detections, and improving the interpretability of anomaly detection, the present invention adopts the following technical solutions:
[0009] An interpretable time series anomaly detection method based on mask generation, comprising the following steps:
[0010] Step 101: Perform time-frequency domain conversion on the time series data of industrial sensors to obtain industrial time-frequency domain signals;
[0011] Step 102: Extract features from the industrial time-frequency domain signals, capture the local patterns of the time series data of industrial sensors in the time dimension, and identify the features of different frequency components in the frequency dimension to obtain potential features;
[0012] Step 103: Vectorize the potential features to obtain discrete potential vectors. Discretizing the potential features can facilitate the subsequent processing and understanding of the model;
[0013] Step 104: Construct and train a prior model, and learn the probability distribution of the discrete potential vectors through mask generation modeling;
[0014] Step 105: Calculate the anomaly score based on the probability distribution of the discrete potential vectors, perform anomaly score mapping processing on the anomaly score, use nearest neighbor interpolation technology to map the final anomaly score obtained in the discrete potential space to the data space, and perform anomaly status determination processing on the anomaly score based on a threshold;
[0015] Step 106: Perform a mask operation on the detected anomaly segment, resample it using the prior model, convert the resampled data back to the time-frequency domain, and then perform inverse time-frequency domain conversion to generate a visual result of the possible normal state.
[0016] Further, in step 101, the time series data of industrial sensors are segmented by applying a window function, and then for the data within each window, short-time Fourier transform is used for time-frequency domain conversion to obtain time-frequency domain signals. The formula is as follows:
[0017] x(t) = {x0, x1, …, x T-1}
[0018]
[0019] X(f, t) ∈ R 2*H*T′
[0020] Where t represents the time index, x(t) represents the time - series data of the industrial sensor, T represents the length of x(t), X(f, t) represents the data after short - time Fourier transform, f represents the frequency component, n represents the time index for traversal, x(n) is the n - th sampling point of x(t), ω represents the window function, ω(n - t) represents the translation of the window function at time t, j represents the imaginary unit; 2 represents the real and imaginary parts channels, H represents the frequency dimension, and T′ represents the length of X(f, t).
[0021] Based on the industrial time - frequency domain signal, it is possible to obtain the performance of the time - series data of the industrial sensor at different times and frequencies, and it helps to capture the transient features and frequency changes in the industrial time series, providing important information for subsequent industrial anomaly detection.
[0022] Furthermore, the operation of feature extraction on the time - frequency domain signal in step 102 is based on the industrial time - frequency domain signal. Through the convolutional layer in the convolutional neural network, it captures the local patterns of the time - series data of the industrial sensor in the time dimension and identifies the features of different frequency components in the frequency dimension. The formula is as follows:
[0023] z = E(X(f, t))
[0024] z ∈ R D*H*W
[0025] Where X(f, t) represents the data after time - frequency domain conversion, f represents the frequency component, t represents the time index, E represents the convolutional operation of the encoder, z represents the latent feature, D represents the latent feature dimension, and W represents the latent time dimension.
[0026] Since industrial sensor data has high dimensionality and existing methods are vulnerable to noise interference, feature extraction can capture the features of industrial time series at different times and frequencies. Moreover, it can reveal the frequency components of the data and their changes over time, thereby enhancing the learning ability of the model on high - dimensional data and at the same time enhancing the robustness to noise, providing rich time - frequency information for subsequent industrial anomaly detection.
[0027] Further, in step 103, when performing vector quantization processing on the potential features, the vector quantizer VQ traverses all vectors in the potential vector z, calculates the distances between the current vector in the potential vector z and all vectors in the codebook C, selects the vector in the codebook C with the shortest distance from the current vector in the potential vector z, and uses its codebook index as the quantization result of the current vector in the potential vector z. The formula is as follows:
[0028] z q = VQ(z)
[0029]
[0030] where z q represents the quantization result of all vectors in the potential vector z. The color similarity between each position in z q represents the similarity between the vectors represented by these positions. C k represents the k-th vector in the codebook C. ||z - C k || represents the distances between all vectors in the potential vector z and all vectors in C. The shorter the distance between vectors, the more similar the vectors are. Use s to represent the codebook index part in z q . Each position in s corresponds to an index in the codebook and is used to represent the quantization result of that position.
[0031] It is achieved by mapping the features to a finite and pre-defined codebook, thereby simplifying the complexity of industrial time series data, reducing the computational cost, and providing an effective feature representation method for anomaly detection and generation tasks.
[0032] Further, in the mask generation modeling of step 104, the mask operation formula is as follows:
[0033] s M = s ⊙ m + [MASK] ⊙ (1 - m)
[0034] where s M represents the randomly masked version of the codebook index part s. m represents the mask matrix, and the elements are 0 or 1. 1 represents retention, and 0 represents masking. [MASK] represents the mask block used to replace the masked positions. ⊙ represents element-wise multiplication;
[0035] The goal of training the prior model is to maximize the mask prediction probability. The probability distribution of the mask prediction is denoted as p θ (s|s M ), where θ represents the parameters of the prior model. The prior model can use a bidirectional Transformer model;
[0036] Each element of the mask prediction probability is represented in softmax form as:
[0037]
[0038] u = f θ (s M )
[0039] where s i represents the i-th position in s, k represents the codebook vector index, K represents the codebook size, u represents the unnormalized score output by the prior model, and (u i ) k represents the score that the i-th position is k; k * represents the true codebook index, and f θ represents the prior model.
[0040] For mask generation modeling, the prior model uses a sliding mask window of arbitrary size to randomly mask a part of the positions in s, and then predicts the original values of these positions along the time dimension;
[0041] The principle of the mask prediction probability p θ (s|s M ) is a process that assigns higher probabilities to possible states and lower probabilities to unlikely and impossible states; the prior model is trained by maximizing p θ (s|s M ) so as to assign higher values, and assign lower values to (u i ) k The former generates higher probabilities for possible normal states, and the latter generates lower probabilities for unlikely and impossible states.
[0042] Using the bidirectional self-attention mechanism of the bidirectional Transformer is beneficial for the model to simultaneously utilize the context information of the current position and enhance the modeling of global time dependencies; using mask generation modeling to randomly mask some positions and predict the original values of these positions enables the model to learn to distinguish normal and abnormal states.
[0043] Furthermore, the training process of the prior model is as follows:
[0044] In the first stage, the encoder E for feature extraction, the vector quantizer VQ, and the decoder Decoder are trained, and the goal is to minimize the reconstruction error. The formula used is as follows:
[0045] E1||x(t)-Decoder(VQ.E(X(f,t)) / )||
[0046] Among them, E1 represents the expected value, t represents the time index, x(t) represents the time series data of industrial sensors, X(f,t) represents the data after time-frequency domain conversion, and f represents the frequency component;
[0047] In the second stage, the prior model is trained with the goal of maximizing the masked prediction probability:
[0048] E1[p θ (s|s M )]
[0049] Among them, the encoder E and the vector quantizer VQ trained in the first stage are set to be untrainable in the second stage.
[0050] Using the two-stage training method enables the model to effectively learn the distribution of normal behavior from time series data and identify anomalies; discrete latent features are extracted through vector quantization technology, and a prior model is constructed using masked generative modeling, enabling the model to simultaneously learn anomaly patterns in both the time and frequency dimensions and improve sensitivity to complex industrial anomalies.
[0051] Furthermore, in step S105, based on the probability distribution of the discrete latent vectors, the anomaly score is calculated as follows:
[0052] Using the negative log-likelihood formula -logp θ (s|s M ), sliding the masked window along the time dimension to calculate the anomaly score a of a specific time step w in s w , and the formula is as follows:
[0053] a w =E1[-logp θ (s :,w-a:w+a |s M(:,w-a:w+a) )]
[0054] Among them, a w ∈R H , -logp θ (s :,w-a:w+a |s M(:,w-a:w+a) )∈R H*2a ,[:,w-a:w+a] represents the time range position from w-a to w+a in the frequency dimension, M(:,w-a:w+a) represents the masking operation on the time range position from w-a to w+a in the frequency dimension, and a represents a flexible positive integer, reflecting the randomness of the masking operation;
[0055] Perform multi-scale score aggregation, use different a values to generate multiple groups of anomaly scores, and perform an averaging operation on the multiple groups of anomaly scores to obtain the final anomaly score a final .
[0056] The calculation of the anomaly score intuitively quantifies the degree of anomaly, facilitating the identification and location of anomalies; the nearest neighbor interpolation technique maps the anomaly score back to the original data space, making the anomaly detection results more intuitive and easier to interpret.
[0057] Further, in the step S105, the anomaly status determination process for the anomaly score is as follows:
[0058] Set the threshold to the nth quantile of a calculated using the training dataset, where n is determined according to the number of existing anomalies in the dataset. The more anomalies there are in the training dataset, the lower n is; in the test dataset, if the anomaly score a of the test dataset final is higher than the threshold, then set this time step as an anomaly. final By setting a threshold and framing the anomalous part, the location of the anomaly can be quickly and accurately determined, providing valuable insights for practical applications.
[0059]
[0060] Further, in the step 106, the time series data x(t), after time-frequency domain conversion, the convolutional operation E of the encoder, and the quantization operation VQ of the vector quantizer, obtains s. The horizontal axis of s is the time axis, and the vertical axis is the frequency axis. The bluish part in s represents the anomalous part, and other colors represent the normal part; then a masking operation is performed on the anomalous part of s to obtain the masked version s of s M , and the gray part in s M represents the masked part; then s M is predicted by the prior model (bidirectional Transformer) to obtain the possible normal state s' of s; finally, through the decoder decoder, the time series data of the possible normal state is obtained
[0061] An interpretable time series anomaly detection system based on mask generation includes a preprocessing unit, an encoding quantization unit, an anomaly detection unit, and an interpretable enhancement unit connected in sequence. Each unit performs time-frequency domain conversion, feature extraction and vectorization processing, constructing a prior model for anomaly detection, and generation of the normal state in sequence according to the described interpretable time series anomaly detection method based on mask generation.
[0062] The technical solutions provided by the embodiments of this specification may include the following beneficial effects:
[0063] The present invention significantly improves the ability to capture abnormal patterns in industrial time series by introducing mask generation modeling; by designing a semantic-preserving convolutional encoder and a vector quantization module, it preserves frequency band independence and time alignment, eliminates the interference of cross-band information mixing, and effectively addresses the high dimensionality of industrial time series data; through the modeling method of discrete latent space, it significantly enhances the sensitivity of the model to complex abnormal patterns in industry, improves the robustness and interpretability of detection results; by generating possible normal states, it realizes the visual interpretation of the abnormal positions in industrial time series data, thereby enhancing the information integrity of industrial anomaly detection and the user decision-making support ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a schematic structural diagram of the system according to an embodiment of the present invention.
[0065] Figure 2 is a flowchart of the method according to an embodiment of the present invention.
[0066] Figure 3 is a schematic diagram of vector quantization processing of latent features in an embodiment of the present invention.
[0067] Figure 4a is a schematic diagram of two-stage training of the prior model in an embodiment of the present invention (the first stage).
[0068] Figure 4b is a schematic diagram of two-stage training of the prior model in an embodiment of the present invention (the second stage).
[0069] Figure 5 is a schematic diagram of anomaly detection of time series in an embodiment of the present invention.
[0070] Figure 6 is a schematic diagram of the visualization result of generating possible normal states in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.
[0072] Time series anomaly detection has important application values in the industrial field and is widely used in equipment health monitoring, fault prediction, and production optimization. Existing anomaly detection methods mainly rely on statistical analysis, machine learning, or deep learning techniques, and these methods can identify abnormal points in time series to a certain extent.
[0073] In the industrial production process, equipment relies on sensors to collect key parameters such as temperature, pressure, and current to monitor the operating status and detect abnormalities. The time series data generated by these sensors is crucial for equipment health management. However, industrial sensor data has complexity, leading to many challenges in anomaly detection with existing methods. First of all, industrial sensor data usually has high dimensionality. This faces technical challenges of large data volume, high storage and computing overhead, and it is difficult for existing methods to process efficiently. Secondly, industrial sensor data is vulnerable to noise interference and may have missing data. This data instability will reduce the accuracy of anomaly detection, making it difficult for existing methods to effectively cope with. In addition, industrial sensor data usually exhibits the coexistence of periodicity and mutability. Existing methods are difficult to identify periodic patterns and abnormal mutations simultaneously, resulting in false detections or missed detections. The operating status of many industrial devices has periodicity. However, equipment anomalies may lead to mutations. Such mutations are usually early signs of equipment failures. If not detected in time, it may lead to equipment damage or even production accidents. Finally, current methods are insufficient in the interpretability of industrial anomaly detection. Users often have difficulty understanding the specific reasons and nature of the anomalies, which limits the wide acceptance and use of anomaly detection technology in industrial practical applications.
[0074] To address the above problems, this application proposes an interpretable time series anomaly detection system based on mask generation. By more deeply and effectively integrating the time-frequency characteristics of industrial time series data, it extracts richer multi-scale and multi-frequency features to improve the ability to capture fine-grained information of industrial time series data, thus facilitating subsequent anomaly detection and interpretation tasks, such as Figure 1 As shown, the system includes a preprocessing unit, an encoding quantization unit, an anomaly detection unit, and an interpretability enhancement unit.
[0075] The preprocessing unit uses the short-time Fourier transform to perform time-frequency domain conversion on industrial time series data, converting time-domain data into time-frequency domain data to obtain time-frequency domain signals.
[0076] The encoding quantization unit performs feature extraction operations on the time-frequency domain signals to obtain latent features; then vectorizes the latent features to obtain discrete latent vectors;
[0077] The anomaly detection unit constructs and trains a prior model, enabling the prior model to learn the probability distribution of discrete latent vectors using mask generation modeling technology; based on the probability distribution of discrete latent vectors, calculates the anomaly score, and performs anomaly score mapping processing and anomaly status determination processing;
[0078] The interpretability enhancement unit is used to generate visualization results of possible normal states.
[0079] The present invention also provides an interpretable time series anomaly detection method based on mask generation, such as Figure 2 shown, including the following steps:
[0080] Step 101: Perform time-frequency domain conversion on the time series data using the short-time Fourier transform, and convert the time-domain data into time-frequency domain data to obtain the time-frequency domain signal.
[0081] The time series data is the time series data collected by industrial sensors;
[0082] The formula used is as follows:
[0083] x(t) = {x0, x1, …, x T-1}
[0084]
[0085] X(f, t) ∈ R 2*H*T′
[0086] where t is the time index, x(t) is the time series data, T is the length of x(t), X(f, t) is the data after the short-time Fourier transform, f is the frequency component, n is the time index for traversal, x(n) is the nth sampling point of x(t), ω is the window function, ω(n - t) is the translation of the window function at time t, j is the imaginary unit; 2 is the real and imaginary parts channel, H is the frequency dimension, and T′ is the length of X(f, t).
[0087] In one embodiment, first apply a window function to the industrial time series data for segmentation, and then perform a fast Fourier transform on the data within each window to obtain the time-frequency domain signal, providing an input for subsequent feature extraction operations.
[0088] Through the above embodiments, the performance of the industrial time series data at different times and frequencies can be obtained; and it helps to capture the transient features and frequency changes in the industrial time series, providing important information for subsequent industrial anomaly detection.
[0089] Step 102: Perform feature extraction operations on the time-frequency domain signal to obtain latent features.
[0090] The formula used is as follows:
[0091] z = E(X(f, t))
[0092] z ∈ R D*H*W
[0093] where E is the convolutional operation of the encoder, z is the latent feature, D is the latent feature dimension, and W is the latent time dimension.
[0094] In one embodiment, first, a time-frequency domain signal obtained through short-time Fourier transform is received, and then, through the convolutional layer in a convolutional neural network (CNN), local patterns of industrial time series are captured in the time dimension, and features of different frequency components are identified in the frequency dimension.
[0095] Industrial sensor data has high dimensionality, and existing methods are vulnerable to noise interference. Through the above embodiment, features of industrial time series at different times and frequencies can be captured; moreover, the frequency components of the data and their changes over time can be revealed, thereby enhancing the learning ability of the model on high-dimensional data and at the same time enhancing the robustness to noise, providing rich time-frequency information for subsequent industrial anomaly detection.
[0096] Step 103: Perform vector quantization processing on the latent features to obtain discrete latent vectors.
[0097] The formula used is as follows:
[0098] z q = VQ(z)
[0099] z q ∈R D*H*W
[0100] s ∈ R H*W
[0101] where VQ is the quantization operation of the vector quantizer, z q is the discrete latent vector; s is the codebook index part of z q , s ∈ R H*W , and each position in s corresponds to an index in the codebook, which is used to represent the quantization result of that position;
[0102] In one embodiment, as Figure 3 shown, the encoder E is used to convert the time-frequency domain signal X(f,t) into latent features z. The vector quantizer VQ traverses all the vectors in z, calculates the distances between the current vector in z and all the vectors in the codebook (CodeBook) C, selects the vector in C with the shortest distance to the current vector in z, and uses its codebook index as the quantization result of the current vector in z. The formula used is as follows:
[0103]
[0104] where argmin is the value of the independent variable that makes the function reach the minimum value, k is the vector index in C, C k is the vector in C, ||z - C k || is the distance between all the vectors in z and all the vectors in C; the quantization result of all the vectors in z is z q , the shorter the distance between vectors means the more similar the vectors are, z qThe color similarity between each position in represents the similarity between the vectors represented by these positions; hereinafter, for the sake of brevity, s is used instead of z q .
[0105] Through the above embodiments, the latent features are discretized to facilitate model processing and understanding; moreover, this process is achieved by mapping the features to a finite and predefined codebook, thereby simplifying the complexity of industrial time series data, reducing the computational cost, and providing an effective feature representation for anomaly detection and generation tasks.
[0106] Step 104: Construct and train a prior model so that the prior model learns the probability distribution of the discrete latent vectors using the masked generation modeling technique.
[0107] The formula utilized by the masking operation in the masked generation modeling technique is as follows:
[0108] s M = s ⊙ m + [MASK] ⊙ (1 - m)
[0109] where s M is the randomly masked version of s, m is the masking matrix with elements 0 or 1, where 1 means keep and 0 means mask; [MASK] is the masked block used to replace the masked positions; ⊙ is element-wise multiplication.
[0110] The goal of training the prior model is to maximize the masked prediction probability, and the probability distribution of the masked prediction is denoted as p θ (s|s M ), where θ represents the parameters of the prior model.
[0111] Each element of the masked prediction probability is represented in softmax form as:
[0112]
[0113] u = f θ (s M )
[0114] where s i represents the i-th position in s, k is the codebook vector index, K is the codebook size, u is the unnormalized score output by the prior model, (u i ) k is the score that the i-th position is k; k * is the true codebook index, and f θ is the prior model.
[0115] The steps of training the prior model include:
[0116] In the first stage, the encoder E, the VQ, and the decoder Decoder are trained with the goal of minimizing the reconstruction error, using the following formula:
[0117] E1||x(t) - Decoder(VQ.E(X(f,t)) / )||
[0118] where E1 is the expected value;
[0119] In the second stage, the prior model is trained with the goal of maximizing the masked prediction probability:
[0120] E1[p θ (s|s M )]
[0121] where the encoder E and the VQ trained in the first stage are set to be non-trainable in the second stage.
[0122] In one embodiment, s is used as the input data, and a bidirectional Transformer model is used to construct the prior model; then, the prior model is trained, and the training dataset and the mask generation modeling are used to learn the prior distribution p(s) of s, and the learned prior distribution is denoted as p θ (s); for the mask generation modeling, specifically, the prior model uses a sliding mask window of any size to randomly mask a part of the positions of s, and then predicts the original values of these positions along the time dimension; the goal of training the prior model is to maximize the masked prediction probability p θ (s|s M );the principle of p θ (s|s M ) is a process of assigning higher probabilities to possible states and lower probabilities to unlikely and impossible states; the prior model is trained by maximizing p θ (s|s M ) so as to assign higher values, and assign lower values to (u i ) k ; the former generates higher probabilities for possible normal states, and the latter generates lower probabilities for unlikely and impossible states.
[0123] In addition, a two-stage training method is adopted for training the prior model, such as Figure 4a 、 Figure 4bAs shown, in the first stage, the time series data x(t) is subjected to the short-time Fourier transform STFT, the convolutional operation E of the encoder, and the quantization operation VQ of the vector quantizer to obtain s, and then through the decoder Decoder and the inverse short-time Fourier transform ISTFT, the reconstructed time series data x'(t) is obtained; in the first stage, E, VQ, and Decoder are trained, and the goal is to minimize the reconstruction error; in the second stage, the time series data x(t) is subjected to STFT, E, and VQ to obtain s, and then through the masking operation, the masked version s of s is obtained M , and finally through the bidirectional Transformer, the predicted value s' of s is obtained; in the second stage, the bidirectional Transformer is trained, and the goal is to maximize the masking prediction probability, where E and VQ trained in the first stage are set to be non-trainable in the second stage, and the non-trainable E and VQ are marked with snowflake symbols.
[0124] Existing anomaly detection methods are difficult to simultaneously capture the anomaly patterns of industrial time series data in the time and frequency dimensions, resulting in limited detection accuracy. Through the above embodiments, the bidirectional self-attention mechanism of the bidirectional Transformer is used, which is beneficial for the model to simultaneously utilize the context information of the current position and enhance the modeling of global time dependence; the masking generation modeling is used to randomly mask some positions and predict the original values of these positions, aiming to let the model learn to distinguish normal and abnormal states; the two-stage training method is used to enable the model to effectively learn the distribution of normal behaviors from time series data and identify anomalies; the discrete latent features are extracted through vector quantization technology, and the prior model is constructed using the masking generation modeling, enabling the model to simultaneously learn the anomaly patterns in the time and frequency dimensions and improve the sensitivity to industrial complex anomalies.
[0125] Step 105: Calculate the anomaly score based on the probability distribution of the discrete latent vector, and perform anomaly score mapping processing and anomaly state determination processing on the anomaly score.
[0126] Calculating the anomaly score based on the probability distribution of the discrete latent vector includes:
[0127] Using the negative log-likelihood formula -logp θ (s|s M ), sliding the masking window along the time dimension to calculate the anomaly score a of a specific time step w in s w , and the formula is as follows:
[0128] a w =E1[-logp θ (s :,w-a:w+a |s M(:,w-a:w+a) )]
[0129] where a w ∈RH , -logp θ (s :,w-a:w+a |s M(:,w-a:w+a) ) ∈ R H*2a ,[:, w - a:w + a] is the time range position from w - a to w + a in the frequency dimension, M(:, w - a:w + a) is the masking operation on the time range position from w - a to w + a in the frequency dimension, and a is a flexible positive integer, reflecting the randomness of the masking operation;
[0130] Perform multi-scale score aggregation, use different values of a to generate multiple sets of anomaly scores, and perform an averaging operation on the multiple sets of anomaly scores to obtain the final anomaly score a final .
[0131] The steps for performing anomaly score mapping processing on the anomaly scores specifically include: using the nearest neighbor interpolation technique to map the final anomaly score obtained in the discrete latent space to the data space.
[0132] The specific steps for performing anomaly status determination processing on the anomaly scores include:
[0133] Set the threshold to the nth quantile of a calculated using the training dataset, where n is determined according to the number of existing anomalies in the dataset. The more anomalies there are in the training dataset, the lower n is; in the test dataset, if the anomaly score a final of the test dataset is higher than the threshold, then set that time step as an anomaly. final
[0134] In one embodiment, as Figure 5 shown, for the time series data x(t), after short-time Fourier transform STFT, convolutional operation E of the encoder, and quantization operation VQ of the vector quantizer, s is obtained. The horizontal axis of s is the time axis and the vertical axis is the frequency axis. The bluish part in s represents the abnormal part, and other colors represent the normal part; then a random masking operation is performed on all parts of s to obtain the masked version s M , s M where the gray part in s represents the masked part; then the predicted value of the masked part is obtained through a bidirectional Transformer, and the anomaly score is calculated through -logp θ (s|s M ) The anomaly score, where the reddish part represents a high anomaly score and the green part represents a low anomaly score; then use the nearest neighbor interpolation technique to map the final anomaly score a final obtained in the discrete latent space to the data space; finally, determine whether the anomaly score a final of each part in the time series data exceeds the threshold. If it exceeds the threshold, then frame the abnormal part with an orange border to determine the abnormal position.
[0135] Through the above embodiments, the calculation of the anomaly score intuitively quantifies the degree of anomaly, facilitating the identification and location of anomalies; the nearest neighbor interpolation technique maps the anomaly score back to the original data space, making the anomaly detection results more intuitive and easier to interpret; finally, by setting a threshold and bounding the anomalous part, the anomaly location can be determined quickly and accurately, providing valuable insights for practical applications.
[0136] Step 106: Generate a visualization result of the possible normal state.
[0137] Generating a visualization result of the possible normal state includes: after detecting an anomalous segment, performing a masking operation on the anomalous segment in s, then resampling using a prior model, transforming the resampled data back to the time-frequency domain through a decoder, and then performing an inverse short-time Fourier transform (ISTFT) to generate a visualization result of the possible normal state.
[0138] In one embodiment, as Figure 6 shown, for the time series data x(t), after short-time Fourier transform (STFT), the convolutional operation E of the encoder, and the quantization operation VQ of the vector quantizer, s is obtained. The horizontal axis of s is the time axis, and the vertical axis is the frequency axis. The bluish part in s represents the anomalous part, and other colors represent the normal parts; then a masking operation is performed on the anomalous part of s to obtain the masked version s M , s M where the gray part in s M represents the masked part; then s
[0139] After passing through a bidirectional Transformer prediction, the possible normal state s′ of s is obtained; finally, through the decoder, the time series data of the possible normal state is obtained.
[0140] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An interpretable time series anomaly detection method based on mask generation, characterized in that It includes the following steps: Step 101: Perform time-frequency domain conversion on the time series data of industrial sensors to obtain industrial time-frequency domain signals; Step 102: Extract features from the industrial time-frequency domain signals, capture the local patterns of the time series data of industrial sensors in the time dimension, and identify the features of different frequency components in the frequency dimension to obtain potential features; Step 103: Vectorize the potential features to obtain discrete latent vectors; Step 104: Construct and train a prior model, and learn the probability distribution of the discrete latent vectors through masked generative modeling; Step 105: Calculate an anomaly score based on the probability distribution of the discrete latent vectors, perform anomaly score mapping processing on the anomaly score, map the final anomaly score obtained in the discrete latent space to the data space, and perform anomaly status determination processing on the anomaly score based on a threshold; Step 106: Perform a masking operation on the detected abnormal segment, resample it using the prior model, convert the resampled data back to the time-frequency domain, and then perform inverse time-frequency domain conversion to generate a visualization result of the possible normal state.
2. The interpretable time series anomaly detection method based on mask generation according to claim 1, wherein: In step 101, the time series data of industrial sensors are segmented by applying a window function, and then for the data within each window, short-time Fourier transform is used for time-frequency domain conversion to obtain time-frequency domain signals. The formula is as follows: x(t) = {x0, x1, …, x T-1} X(f,t) ∈ R 2*H*T′ Where t represents the time index, x(t) represents the time series data of industrial sensors, T represents the length of x(t), X(f,t) represents the data after short-time Fourier transform, f represents the frequency component, n represents the time index for traversal, x(n) is the nth sampling point of x(t), ω represents the window function, ω(n - t) represents the translation of the window function at time t, j represents the imaginary unit; 2 represents the real and imaginary parts channels, H represents the frequency dimension, and T′ represents the length of X(f,t).
3. The interpretable time series anomaly detection method based on mask generation according to claim 1, characterized in that: In step 102, the feature extraction operation on the time-frequency domain signals is based on the industrial time-frequency domain signals. Through the convolutional layer in the convolutional neural network, local patterns of the time series data of industrial sensors are captured in the time dimension, and the features of different frequency components are identified in the frequency dimension. The formula is as follows: z = E(X(f,t)) z belongs to R D*H*W Where X(f,t) represents the data after time-frequency domain conversion, f represents the frequency component, t represents the time index, E represents the convolutional operation of the encoder, z represents the potential feature, D represents the potential feature dimension, and W represents the potential time dimension.
4. The method for interpretable time series anomaly detection based on mask generation according to claim 1, characterized in that: In step 103, when vectorizing the potential features, the vector quantizer VQ traverses all vectors in the latent vector z, calculates the distances between the current vector in the latent vector z and all vectors in the codebook C, selects the vector in the codebook C with the shortest distance to the current vector in the latent vector z, and uses its codebook index as the quantization result of the current vector in the latent vector z. The formula is as follows: z q = VQ(z) Among them, z q represents the quantization result of all vectors in the latent vector z, and C k represents the k-th vector in the codebook C, ||z - C k || represents the distance between all vectors in the latent vector z and all vectors in C; s is used to represent the codebook index part in z q Each position in s corresponds to an index in the codebook, which is used to represent the quantization result at that position.
5. The interpretable time series anomaly detection method based on mask generation according to claim 4, characterized in that: In the masked generative modeling of step 104, the masking operation formula is as follows: s M = s ⊙ m + [MASK] ⊙ (1 - m) where s M represents the randomly masked version of the codebook index part s, m represents the mask matrix, [MASK] represents the masked block used to replace the masked positions; ⊙ represents element-wise multiplication; The goal of training the prior model is to maximize the mask prediction probability, and the probability distribution of the mask prediction is denoted as p θ (s|s M ), where θ represents the parameters of the prior model; Each element of the mask prediction probability is expressed as: u = f θ (s M ) Among them, s i represents the i-th position in s, k represents the codebook vector index, K represents the codebook size, u represents the unnormalized score output by the prior model, (u i ) k represents the score that the i-th position is k; k * represents the true codebook index, and f θ represents the prior model.
6. The method for interpretable time series anomaly detection based on mask generation according to claim 5, wherein: The training process of the prior model is as follows: In the first stage, the encoder E for feature extraction, the vector quantizer VQ, and the decoder Decoder are trained. The goal is to minimize the reconstruction error, and the formula used is as follows: E1||x(t)-Decoder(VQ.E(X(f,t)) / )|| where E1 represents the expected value, t represents the time index, x(t) represents the time series data of industrial sensors, X(f,t) represents the data after time-frequency domain conversion, and f represents the frequency component; In the second stage, the prior model is trained. The goal is to maximize the mask prediction probability: E1[p θ (s|s M )] Among them, the encoder E and the vector quantizer VQ trained in the first stage are set to be untrainable in the second stage.
7. The interpretable time series anomaly detection method based on mask generation according to claim 6, characterized in that: In the step S105, based on the probability distribution of the discrete latent vector, the anomaly score is calculated as follows: Using the negative log-likelihood formula -logp θ (s|s M ) and sliding a masking window along the time dimension to calculate the anomaly score a of a specific time step w in s w , the formula is as follows: a w = E1[-logp θ (s :,w-a:w+a |s M(:,w-a:w+a) )] where[:,w-a:w+a] represents the time range position from w-a to w+a in the frequency dimension, M(:,w-a:w+a) represents the masking operation on the time range position from w-a to w+a in the frequency dimension, and a represents a flexible positive integer; Perform multi-scale score aggregation, use different values of a to generate multiple sets of anomaly scores, and perform an averaging operation on the multiple sets of anomaly scores to obtain the final anomaly score a final .
8. The method for interpretable time series anomaly detection based on mask generation according to claim 7, wherein: In the step S105, the anomaly status determination process for the anomaly score is as follows: Set the threshold to the n-th quantile of a calculated using the training dataset, where n is determined based on the number of existing anomalies in the dataset. The more anomalies there are in the training dataset, the lower n is; final In the test dataset, if the anomaly score a of the test dataset final is higher than the threshold, then set this time step as an anomaly.
9. The interpretable time series anomaly detection method based on mask generation according to claim 6, characterized in that: In the said step 106, the time series data x(t), after time-frequency domain conversion, the convolution operation E of the encoder, and the quantization operation VQ of the vector quantizer, obtains s. The horizontal axis of s is the time axis, and the vertical axis is the frequency axis; then, a masking operation is performed on the abnormal part of s to obtain the masked version s of s M ; then s M is predicted by the prior model to obtain the possible normal state s' of s; finally, through the decoder decoder, the time series data of the possible normal state is obtained 10. An interpretable time series anomaly detection system based on mask generation, comprising a preprocessing unit, an encoding and quantization unit, an anomaly detection unit, and an interpretable enhancement unit connected in sequence, characterized in that: Each unit performs time-frequency domain conversion, feature extraction and vectorization processing, constructs a prior model for anomaly detection, and generates a normal state in sequence according to the interpretable time series anomaly detection method based on masking described in any one of claims 1 to 9.
Citation Information
Patent Citations
Time sequence anomaly detection method based on fast Fourier transform and mask convolution
CN117033933A
Time sequence anomaly detection method and system based on time-frequency mask auto-encoder
CN118468177A
Financial industry-oriented time sequence anomaly detection method and system
CN118839199A
Unsupervised time sequence anomaly detection method based on multi-scale reconstruction
CN119202988A
Method and computer program for monitoring an artificial intelligence
EP4258179A1