Few-sample multi-dimensional time sequence anomaly detection method, system, equipment and medium

By segmenting multidimensional time series data and enhancing it with a pre-trained language model, combined with frequency and time dimension feature extraction, the problem of sparse labeling of multidimensional time series data is solved, and the accuracy and robustness of anomaly detection are improved.

CN120995347APending Publication Date: 2025-11-21NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511123090.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Anomaly detection in multidimensional time series data faces challenges such as scarce samples and limited labeled data. Traditional methods are ineffective, and deep learning is prone to overfitting, making it difficult to build robust and accurate anomaly detection models.

Method used

Multidimensional time series data are segmented, pre-trained language models are used for data augmentation, and features are extracted from both frequency and time dimensions. Anomaly detection is performed by reconstructing features from the frequency and time domains.

Benefits of technology

It improves the accuracy of multidimensional time series anomaly detection, enhances the detection capability in complex scenarios, reduces overfitting, and achieves accurate identification and early warning of early and subtle faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995347A_ABST
    Figure CN120995347A_ABST
Patent Text Reader

Abstract

The invention discloses a few-sample multi-dimensional time sequence anomaly detection method, system and device and a medium, and relates to the field of anomaly detection.The method comprises the steps that a few-sample multi-dimensional time sequence is subjected to block processing, and an initial time sequence of each dimension is obtained; performing exception shielding on the initial time sequence of each dimension to obtain an exception shielding time sequence of each dimension; performing data enhancement on the abnormal shielding time sequence of each dimension by adopting a pre-training language model to obtain an enhanced time sequence of each dimension; performing feature extraction on the enhanced time sequence of each dimension from a frequency dimension and a time dimension to obtain a frequency domain feature and a time domain feature of each dimension; respectively aggregating the frequency domain feature and the time domain feature of each dimension to obtain a reconstruction time sequence of each dimension; and determining an anomaly detection result according to the initial time sequence and the reconstructed time sequence of each dimension. According to the invention, the detection precision of the multidimensional time sequence anomaly is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of anomaly detection, and in particular to a method, system, device, and medium for detecting anomalies in a small number of multidimensional time series samples. Background Technology

[0002] Time series analysis is fundamental to the Industrial Internet of Things (IIoT) and intelligent operation and maintenance (O&M), and is crucial for ensuring production safety and efficiency, especially in fields such as aerospace and financial transactions, where anomaly detection is the primary step in identifying faults and preventing accidents. However, multidimensional time series data suffers from problems such as anomalous sparsity, difficulty in obtaining precisely labeled datasets, and limited labeled training data, making anomaly identification extremely challenging and rendering traditional supervised methods ineffective.

[0003] In recent years, anomaly detection based on statistical models or machine learning has become a core guarantee for industrial equipment monitoring and fault early warning due to its strong real-time performance, good interpretability, and ability to handle time-series characteristics. However, traditional statistical methods are only suitable for simple anomaly scenarios and perform poorly with complex data; while machine learning methods can handle complex scenarios, they rely on effective feature engineering. With the development of multiple fields, the data dimension is increasing, and the demand for early warning of small and rare faults is increasing. Building robust and accurate anomaly detection models when labeled samples are scarce has become a bottleneck restricting the improvement of intelligent levels.

[0004] The field of time series data analysis has seen the emergence of various new technological approaches, with deep learning becoming a powerful tool due to its strong ability to extract time series features. However, real-world datasets exhibit a wide variety of anomalies, and the sheer number of anomalies is difficult to determine, rendering the previous assumption of "uncontaminated datasets" unreasonable. Following the development of deep learning, some researchers have improved neural networks to distinguish between normal and anomalous time series data. However, the powerful modeling capabilities of modern neural networks may reconstruct sparse anomalies in the dataset, leading to overfitting. Therefore, in multi-dimensional time series anomaly detection, obtaining high-quality, uncontaminated data requires in-depth research into the different types and types of anomalous data present in the time series. Summary of the Invention

[0005] The purpose of this application is to provide a method, system, device, and medium for detecting anomalies in multi-dimensional time series with few samples, which can improve the accuracy of anomaly detection in multi-dimensional time series.

[0006] To achieve the above objectives, this application provides the following solution:

[0007] Firstly, this application provides a method for detecting anomalies in small-sample, multi-dimensional time series data, including:

[0008] The multidimensional time series with few samples is divided into blocks to obtain the initial time series for each dimension;

[0009] Anomaly masking is performed on the initial time series of each dimension to obtain the anomaly masked time series of each dimension;

[0010] Pre-trained language models were used to augment the anomaly occlusion time series for each dimension, resulting in augmented time series for each dimension.

[0011] Feature extraction is performed on the enhanced time series of each dimension from both the frequency and time dimensions to obtain the frequency domain features and time domain features of each dimension.

[0012] The frequency domain features and time domain features of each dimension are aggregated to obtain the reconstructed time series of each dimension;

[0013] Based on the initial time series and reconstructed time series for each dimension, the anomaly detection results are determined.

[0014] Secondly, this application provides a few-sample, multi-dimensional time series anomaly detection system, including:

[0015] The preprocessing module is used to divide a small number of multidimensional time series into blocks to obtain the initial time series for each dimension.

[0016] The anomaly masking module is used to mask the anomalies of the initial time series for each dimension, resulting in an anomaly masked time series for each dimension.

[0017] The data augmentation module is used to perform data augmentation on the anomaly occlusion time series of each dimension using a pre-trained language model, so as to obtain the augmented time series of each dimension.

[0018] The time-frequency feature extraction module is used to extract features from the enhanced time series of each dimension from the frequency dimension and the time dimension, respectively, to obtain the frequency domain features and time domain features of each dimension;

[0019] The reconstruction module is used to aggregate the frequency domain features and time domain features of each dimension to obtain the reconstructed time series of each dimension.

[0020] The results determination module is used to determine the anomaly detection results based on the initial time series and the reconstructed time series for each dimension.

[0021] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for detecting anomalies in small-sample multidimensional time series.

[0022] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for detecting anomalies in small-sample multidimensional time series.

[0023] According to the specific embodiments provided in this application, this application has the following technical effects:

[0024] This application provides a method, system, device, and medium for anomaly detection in multidimensional time series with few samples. By masking anomalies in the initial time series of each dimension and using a pre-trained language model for data augmentation, the time-frequency features of the time series are extracted from both frequency and time dimensions, and the reconstruction and anomaly detection of the time series data are completed, effectively improving the accuracy of anomaly detection in multidimensional time series. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is an application environment diagram of a few-sample multidimensional time series anomaly detection method in one embodiment of this application;

[0027] Figure 2 A flowchart illustrating a method for detecting anomalies in a small sample multidimensional time series according to an embodiment of this application;

[0028] Figure 3 This is a schematic diagram of a few-sample multidimensional time series preprocessing process in one embodiment of this application;

[0029] Figure 4 This is a schematic diagram of an abnormal occlusion process in one embodiment of this application;

[0030] Figure 5 This is a schematic diagram of the structure of a pre-trained language model in one embodiment of this application;

[0031] Figure 6 This is a schematic diagram of the structure of a frequency domain feature reconstruction model in one embodiment of this application;

[0032] Figure 7 This is a schematic diagram of an efficient Kolmogorov-Arnold self-attention structure in one embodiment of this application;

[0033] Figure 8 This is a schematic diagram of the functional modules of a small-sample multidimensional time series anomaly detection system provided in an embodiment of this application;

[0034] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0036] This application addresses the challenges in monitoring actual industrial equipment. On the one hand, it requires in-depth experimental analysis and theoretical research on the correlation mechanisms between various time-series characteristics and potential anomaly types, laying a solid theoretical foundation for building robust anomaly detection models. On the other hand, it needs to establish a multi-dimensional time-series anomaly detection model with stronger generalization ability, higher detection accuracy, and lower latency, enabling accurate identification and early warning of early weak fault sources and complex context anomalies in equipment. This provides a scientific basis for setting early warning thresholds, optimizing maintenance strategies, and scheduling resources in actual operation and maintenance. This is not only a core barrier to ensure the safe and stable operation of industrial equipment and avoid catastrophic failures, but also a cutting-edge focus and core challenge in the fields of intelligent operation and maintenance and industrial big data analysis. It plays a crucial driving role in promoting the large-scale implementation of predictive maintenance and improving the level of intelligence in industrial production.

[0037] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0038] The few-sample, multidimensional time series anomaly detection method provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 101 communicates with server 102 via a network. A data storage system can store the data that server 102 needs to process. The data storage system can be set up independently, integrated into server 102, or placed in the cloud or on another server. Terminal 101 can send a few-sample multidimensional time series to server 102. After receiving the few-sample multidimensional time series, server 102 sequentially performs segmentation, anomaly masking, data augmentation, feature extraction, and feature aggregation on the few-sample multidimensional time series to obtain a reconstructed time series for each dimension. Further, based on the initial and reconstructed time series for each dimension, the anomaly detection result is determined. Server 102 can feed back the obtained anomaly detection result to terminal 101. Furthermore, in some embodiments, the few-sample multidimensional time series anomaly detection method can also be implemented independently by server 102 or terminal 101.

[0039] The terminal 101 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 102 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.

[0040] In one exemplary embodiment, such as Figure 2 As shown, a method for detecting anomalies in multidimensional time series data with few samples is provided. This method is executed by a computer device, specifically a terminal or server, or both. In this embodiment, the method is applied to... Figure 1 Taking server 102 as an example, the explanation includes the following steps 201 to 206.

[0041] Step 201: Divide the small sample multidimensional time series into blocks to obtain the initial time series for each dimension.

[0042] In a specific application example, the few-sample multidimensional time series anomaly detection method is used to detect performance anomalies in a server cluster. The few-sample multidimensional time series includes multiple performance indicators of servers in the server cluster at various time points within a set time period, with one dimension corresponding to one performance indicator.

[0043] In the field of server cluster performance monitoring and anomaly detection, large data centers operate tens of thousands of servers, making their stable and efficient operation crucial. In few-sample multidimensional time series analysis, "multidimensional" refers to the simultaneous observation and recording of multiple different variables describing the system state at the same time point t, representing different types of system indicators. Representing a few-sample multidimensional time series as... Where N represents the number of dimensions and T represents the number of time points. For example, at a time point t, measurement samples of 26 dimensions such as CPU utilization, memory utilization, disk read speed, and network traffic are recorded simultaneously, forming a 26-dimensional vector.

[0044] As an optional implementation method, such as Figure 3 As shown, before performing block processing on the few-sample multidimensional time series, normalization and channel independence operations are first performed on the few-sample multidimensional time series.

[0045] Normalization uses simple mean and standard deviation functions to remove non-stationary information from small sample multidimensional time series, as shown in Equation (1).

[0046]

[0047] Where μ(·) and σ(·) represent the mean and standard deviation functions, respectively. This represents the time series with the nth dimension, where n = 1, 2, ..., N. This represents the time series in the nth dimension after normalization.

[0048] Channel-independent operation treats each dimension of the normalized time series as a single-dimensional time series data, sharing the embeddings and weights of each dimension.

[0049] Segmentation processing divides the aforementioned single-dimensional time series into non-overlapping blocks by aggregating adjacent time points.

[0050] Step 202: Perform anomaly masking on the initial time series of each dimension to obtain the anomaly masked time series of each dimension.

[0051] In a specific application example, step 202 includes steps 21 to 24.

[0052] Step 21: For any initial time series, calculate the residual score and similarity score for each block of the initial time series. For example... Figure 4 As shown, the initial time series is decomposed and block similarity is calculated to determine the mapping matrix.

[0053] In this application, the trend term is extracted using moving average decomposition on a block-by-block basis, and the residual score is calculated. Moving average calculates the average of the data at each current time point and the p-1 subsequent time points in the initial time series, thus obtaining a new series that is more stable and easier to observe. Specifically, the residual score of the i-th block is calculated using the following formulas (2) to (4):

[0054] Trend i =AvgPool(P i ,window=p) (2)

[0055]

[0056] Among them, Trend i For the trend term of the i-th segment, Residual i For the residual term of the i-th block, ResScore i Let P be the residual fraction of the i-th block. iLet p be the i-th block of the initial time series, p be the block size, and m be the m-th data point in the block. This application uses the block size p as the window size for the moving average, extracts the trend term of the block, and the components other than the trend term are the residual terms. The average of the p time point data in the residual term of the i-th block is calculated as the residual score of the block.

[0057] The trend term, seasonal term, and residual term correspond to the low-frequency, mid-frequency, and high-frequency components of the signal, respectively. Many industrial sensor data exhibit weak seasonality, and explicit modeling may introduce noise. This application uses a moving average to calculate the residual fraction to capture sudden anomalies and improve computational efficiency.

[0058] Furthermore, the similarity score of the i-th block is calculated using the following formula (5) to quantify local mutation anomalies in the initial time series:

[0059]

[0060] Among them, P i+1 For the (i+1)th block of the initial time series, SimScore i Let be the similarity score of the i-th block.

[0061] Patch similarity assessment provides information about the dynamic evolution of pattern patches. A higher similarity score indicates a more significant local mutation anomaly.

[0062] Step 22: Adaptively fuse the residual score and similarity score of each block of the initial time series to obtain the anomaly score of each block.

[0063] Similarity scores and residual scores are used to identify local and global anomalies, respectively. A sliding window-based mapping matrix maps the similarity and residual scores to the initial time series, and the two are adaptively fused using learnable weight parameters. Specifically, the anomaly score for the i-th block is obtained using the following formula (6):

[0064] AnomalyScore i =W s SimScore i +W r ·ResScore i (6)

[0065] Among them, AnomalyScore i For the anomaly score of the i-th block, W s and W r These are the weighting parameters for the similarity score and the residual score, respectively.

[0066] Step 23: Generate a dynamic mask matrix based on the anomaly score of each block of the initial time series.

[0067] After the operations in steps 21 and 22 above, the anomaly score of each block is obtained. Based on the anomaly score, a higher weight is assigned to the anomaly region for mask selection. The masking mechanism can filter out unimportant future information in the time series data. Specifically, the following formulas (7) to (10) are used to generate the dynamic mask matrix:

[0068] W ano =1+3⊙AnomalyScore i (7)noise~U(0,1) T ⊙W ano (8)

[0069] ids keep =argsort(noise,descending=True) (9)

[0070] mask=1 T -scatter(1,ids keep ,0) (10)

[0071] In the process of generating the dynamic mask matrix, W ano Anomalous regions are assigned higher weights, and "noise" is the anomaly-aware noise generated based on these weights. Low-noise regions represent normal patterns in the time series, while high-noise regions represent potential anomalies within the series. (ids) keep To sort by noise, calculate the indices that need to be retained. 1 T Given a matrix of all ones, the scatter function will extract the matrix's ids. keep The index position is set to 0, and the mask is a dynamic mask matrix.

[0072] Step 24: Perform anomaly masking on the initial time series according to the dynamic masking matrix to obtain anomaly masked time series of the corresponding dimension. Specifically, retain low-noise regions, mask high-noise regions, and fill the masked regions with channel mean to preserve statistical characteristics.

[0073] Step 203: Use a pre-trained language model to perform data augmentation on the anomaly occlusion time series of each dimension to obtain the augmented time series of each dimension.

[0074] This application augments the data of anomaly occlusion time series by transferring the parameters of a pre-trained language model to the target task (such as performance anomaly detection of server clusters).

[0075] Large language models have proven their powerful capabilities as few-shot learners in time-series tasks. Considering the unsupervised learning capabilities and adaptability of GPT-2 to few-shot domains, this application employs the GPT-2 model to augment the data of anomaly-masked time series, resulting in augmented time series.

[0076] Generative Pre-trained Transformer (GPT) is a pre-trained model based on the Transformer Decoder trained on massive amounts of text. Its specific structure is as follows: Figure 5 As shown, the structure within the dashed section has multiple components. GPT-2 uses a learnable embedding matrix to implement positional encoding, assigning a d-dimensional vector to each time point of the input sequence, allowing it to learn the most suitable positional representation for the current task from the training data, where d is the hidden layer size of the pre-trained model. Figure 5 Only the normalization layer and positional encoding are trained, while all other parameters are frozen. The calculation process of GPT-2 is shown in formula (11):

[0077] h (l) =LayerNorm(Attention(h (l-1) )+FFN(Attention(h (l-1) (11)

[0078] Where h represents the Transformer layer stack of GPT-2, l represents the number of model layers, LayerNorm, Attention, and FFN are the normalization layer, attention layer, and feedforward neural network of the Transformer architecture, respectively, and the trainable parameter θ trainable ={θ ln ,θ wpe Optimization is achieved through fine-tuned normalization layers and positional encoding. The final output of GPT-2 is anomaly-free time series data, i.e., augmented time series.

[0079] Step 204: Extract features from the enhanced time series of each dimension from the frequency dimension and the time dimension respectively to obtain the frequency domain features and time domain features of each dimension.

[0080] In a specific application example, step 204 includes steps 41 and 42.

[0081] Step 41: Based on the frequency domain feature reconstruction model, capture the contextual information of the enhanced time series in each dimension to obtain the frequency domain features of each dimension. The frequency domain feature reconstruction model is constructed based on the discrete cosine transform and compression mechanism.

[0082] The frequency domain feature reconstruction model uses discrete cosine transform and compression excitation mechanisms to capture enhanced local and global contextual information of time series from the frequency dimension, merging output features into an output feature map to enhance the correlation between dimensions of time series data. This model improves robustness in processing non-stationary sequence data, providing an optimized time-frequency domain representation for small-sample anomaly detection algorithms. The specific internal structure is as follows: Figure 6 As shown, frequency domain analysis can reveal latent features in time-series signals that are difficult to capture, making it crucial for time-series tasks. The Discrete Cosine Transform (DCT) extracts local features while preserving real-valued outputs, thus addressing the Gibbs phenomenon generated by the Fast Fourier Transform (FFT) and its inherent limitations in processing non-stationary signals.

[0083] Specifically, step 41 includes steps 1) to 5).

[0084] 1) For any augmented time series, divide the augmented time series into multiple frequency blocks according to frequency.

[0085] 2) Perform discrete cosine transform on each frequency block to obtain the frequency domain characteristic representation of each frequency block. The calculation process is shown in formula (12).

[0086]

[0087] Among them, F (j) The frequency domain feature representation of the j-th frequency block. For the j-th frequency block in the n-th dimension, n = 1, 2, ..., N, x t express The data at time point t in the frequency domain is k, which is the frequency index in the frequency domain, ranging from 0 to T-1.

[0088] 3) Concatenate the frequency domain feature representations of each frequency block to obtain the frequency channel vector F. This concatenation can be achieved through stack operations.

[0089] 4) A compression excitation mechanism is used to capture the global context information of the enhanced time series to obtain the frequency domain attention weights.

[0090] The compression activation mechanism is an attention mechanism used in deep convolutional neural networks to enhance the network's representation ability on feature channels. In the compression stage, global average pooling is applied to the enhanced time series input to compress channel features. The activation stage models the dependencies between channels through a two-layer fully connected network. Its calculation process is shown in equations (13) and (14):

[0091]

[0092] Attentionfreq =σ1(W2δ(W1Z) n (14)

[0093] Among them, Z n For channel-level tensors generated by global average pooling, Attention freq For the frequency domain attention weights learned by the model, x n (t) represents the data at time point t in the nth dimension, σ1 is the ReLU activation function, δ is the Sigmoid activation function, and W1 and W2 are fully connected layers with bottleneck structure and nonlinear activation.

[0094] 5) Determine the frequency domain feature X of the corresponding dimension based on the frequency domain attention weights and the frequency channel vectors. freq .

[0095] Specifically, the frequency domain attention weights are multiplied by the frequency channel vectors to obtain frequency domain features, and then information components are extracted from the frequency domain to enhance feature diversity.

[0096] Step 42: Based on the temporal feature reconstruction model, extract the temporal dependencies of the enhanced time series in each dimension to obtain the temporal features of each dimension. The temporal feature reconstruction model is a Kolmogorov-Arnold network that replaces the multilayer perceptron in the Transformer model.

[0097] The efficient Kolmogorov-Arnold network is a lightweight network architecture based on the Kolmogorov-Arnold representation theorem. It reduces memory usage during forward and backward propagation by reconstructing the original linear combination of basis functions into matrix multiplication. Local nonlinear modeling is achieved using dynamic B-spline basis functions and a recursive DeBoor algorithm. Kaiming initialization and noise perturbation strategies ensure efficient parameter allocation. Adaptive grid adjustment dynamically optimizes the control points of the basis functions based on the input data distribution, enhancing adaptability to non-stationary time characteristics. Combining the efficient Kolmogorov-Arnold network with a self-attention mechanism, a nonlinear modeling architecture for temporal feature extraction is established, the detailed internal structure of which is as follows: Figure 7 As shown, the structure of the dashed part has multiple components. By fusing temporal numerical features and timestamp semantic information, a temporal embedding representation is generated. An attention mechanism is used to capture long- and short-term temporal dependencies. The calculation process is shown in formula (15):

[0098]

[0099] Where Q, K, and V are the query, key, and value vectors, respectively. The scaling factor is used. A stable gradient propagation path is constructed using residual connections and Dropout regularization techniques to obtain the attention feature X. atten Under limited sample settings, the parameter efficiency and dynamic mesh adaptive capability of e-KAN avoid overfitting caused by redundant MLP parameters in traditional Transformers. Using e-KANs to replace the traditional multilayer perceptron layers, a feedforward subnetwork is obtained, realizing a high-order nonlinear mapping. Its calculation process is shown in Equation (16):

[0100] X time =LayerNorm(X atten +Dropout(KAN2(σ2(KAN1(X atten (16)

[0101] Among them, X time For frequency domain features, σ2 represents the LeakyReLU activation function. KAN2 and KAN1 are efficient Kolmogorov-Arnold networks that capture abrupt temporal changes through local B-spline basis functions, while the self-attention mechanism models long-term dependencies.

[0102] In a few-sample setting, the parameter efficiency and dynamic mesh adaptation of the efficient Kolmogorov-Arnold network avoid overfitting caused by redundant parameters in traditional Transformers. The calculation process is shown in Equations (17) and (18):

[0103]

[0104] KAN(x)=(Φ L-1 ·Φ L-2 ·...·Φ0)(x) (18)

[0105] in, This represents each input variable x p A unary function Φ mapped to the output space q It represents a continuous function.

[0106] This application combines an efficient Kolmogorov-Arnold network with a self-attention mechanism to form a complementary time series representation, which synergistically enhances the model's nonlinear modeling ability and significantly improves the model's adaptability to the inherent distribution shifts in non-stationary time series data.

[0107] Step 205: Aggregate the frequency domain features and time domain features of each dimension to obtain the reconstructed time series of each dimension, as shown in formula (19).

[0108]

[0109] in, To reconstruct the time series, C1 is the attention coefficient of the frequency domain feature reconstruction result, and C2 is the attention coefficient of the time domain feature reconstruction result.

[0110] Step 206: Determine the anomaly detection results based on the initial time series and reconstructed time series for each dimension.

[0111] In a specific application example, step 206 includes steps 61 to 63.

[0112] Step 61: Based on the initial time series for each dimension, perform inverse normalization on the reconstructed time series for each dimension to obtain the reconstructed data for each dimension.

[0113] This application uses an inverse normalization operation to restore the reconstructed time series to the original data scale, so as to effectively reduce the non-stationarity of the time data. The calculation process is shown in formula (20).

[0114]

[0115] During training, the standard MSE is used as the loss function and the anomaly score. The difference between the reconstructed output and the original input is measured by calculating the average of the squared errors between the reconstructed value and the true value. The loss function is shown in Equation (21).

[0116]

[0117] in, Let X be the value of the loss function. n,t For the initial data at the t-th time point in the n-th dimension, This is the reconstructed data for the nth dimension at the tth time point.

[0118] Step 62: Calculate the anomaly score for each time point based on the reconstructed data for each dimension and the initial time series for each dimension, as shown in formula (22).

[0119]

[0120] Among them, AnomalyScore t The anomaly score at time point t. This is the reconstructed data for the nth dimension.

[0121] Step 63: Determine the anomaly label for each time point based on the anomaly score at each time point to obtain the anomaly detection result; the anomaly label is either abnormal or normal.

[0122] Specifically, cosine is combined with the error distribution of the training and test sets, and anomaly thresholds are set according to quantiles. The anomaly threshold at time point t is set as τ. t If the abnormal score exceeds the abnormal threshold, it is considered abnormal; otherwise, it is considered normal. The calculation process is shown in formula (23).

[0123]

[0124] Where, label t Let be the anomaly label at time point t. By inverse normalization and reconstruction error, this application marks abnormal data as 1 and normal data as 0, and finally outputs the anomaly label for each time point.

[0125] Real-world datasets contain various anomaly types, and the number of anomalies is difficult to determine, making it challenging to obtain a clean dataset for reconstructing model training. Traditional unsupervised deep learning methods assume that the dataset is free from anomaly contamination, which greatly limits the detection accuracy of the model. This application integrates time-frequency feature extraction and few-shot augmentation. Addressing the difficulty of traditional anomaly detection and prediction methods in handling complex and nonlinear temporal dependencies, this application improves model performance by integrating complex influencing factors from both time and frequency dimensions into a multidimensional time-series anomaly detection model, thereby enhancing the accuracy of multidimensional time-series anomaly detection.

[0126] Based on the same inventive concept, this application also provides a few-sample multidimensional time series anomaly detection system for implementing the methods described above. The solution provided by this system is similar to the implementation scheme described in the above methods; therefore, the specific limitations in one or more few-sample multidimensional time series anomaly detection system embodiments provided below can be found in the limitations of the few-sample multidimensional time series anomaly detection method described above, and will not be repeated here.

[0127] In one exemplary embodiment, such as Figure 8 As shown, a few-sample multidimensional time series anomaly detection system is provided, including: a preprocessing module 801, an anomaly masking module 802, a data augmentation module 803, a time-frequency feature extraction module 804, a reconstruction module 805, and a result determination module 806.

[0128] The preprocessing module 801 is used to perform block processing on a small number of multidimensional time series to obtain the initial time series for each dimension.

[0129] The anomaly masking module 802 is used to mask the initial time series of each dimension separately, so as to obtain the anomaly masked time series of each dimension.

[0130] The data augmentation module 803 is used to perform data augmentation on the anomaly occlusion time series of each dimension using a pre-trained language model, so as to obtain the augmented time series of each dimension.

[0131] The time-frequency feature extraction module 804 is used to extract features from the enhanced time series of each dimension from the frequency dimension and the time dimension respectively, so as to obtain the frequency domain features and time domain features of each dimension.

[0132] The reconstruction module 805 is used to aggregate the frequency domain features and time domain features of each dimension to obtain the reconstructed time series of each dimension.

[0133] The result determination module 806 is used to determine the anomaly detection results based on the initial time series and the reconstructed time series for each dimension.

[0134] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 9 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores few-sample, multi-dimensional time series data. The I / O interfaces allow the processor to exchange information with external devices. The communication interface allows communication with external terminals via a network connection. When executed by the processor, the computer program implements a method for detecting anomalies in few-sample, multi-dimensional time series data.

[0135] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0136] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0137] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0138] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0139] In this application, all actions to acquire signals, information, or data are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with the authorization granted by the owner of the relevant device.

[0140] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0141] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0142] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0143] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for detecting anomalies in a small-sample, multi-dimensional time series, characterized in that, The method includes: The multidimensional time series with few samples is divided into blocks to obtain the initial time series for each dimension; Anomaly masking is performed on the initial time series of each dimension to obtain the anomaly masked time series of each dimension; Pre-trained language models were used to augment the anomaly occlusion time series for each dimension, resulting in augmented time series for each dimension. Feature extraction is performed on the enhanced time series of each dimension from both the frequency and time dimensions to obtain the frequency domain features and time domain features of each dimension. The frequency domain features and time domain features of each dimension are aggregated to obtain the reconstructed time series of each dimension; Based on the initial time series and reconstructed time series for each dimension, the anomaly detection results are determined.

2. The method for detecting anomalies in small-sample multidimensional time series according to claim 1, characterized in that, Anomaly masking is performed on the initial time series for each dimension to obtain the anomaly masked time series for each dimension, specifically including: For any initial time series, calculate the residual score and similarity score for each block of the initial time series; The residual score and similarity score of each block of the initial time series are adaptively fused to obtain the anomaly score of each block; A dynamic mask matrix is ​​generated based on the anomaly score of each block of the initial time series; The initial time series is anomaly masked based on the dynamic mask matrix to obtain the anomaly masked time series of the corresponding dimension.

3. The method for detecting anomalies in small-sample multidimensional time series according to claim 1, characterized in that, The pre-trained language model is a pre-trained GPT-2 model.

4. The method for detecting anomalies in small-sample multidimensional time series according to claim 1, characterized in that, Feature extraction is performed on the enhanced time series data for each dimension, from both the frequency and time dimensions, yielding frequency-domain and time-domain features for each dimension. Specifically, these features include: Based on the frequency domain feature reconstruction model, the contextual information of the enhanced time series in each dimension is captured to obtain the frequency domain features of each dimension; the frequency domain feature reconstruction model is constructed based on the discrete cosine transform and compression mechanism. Based on the temporal feature reconstruction model, the temporal dependencies of the augmented time series in each dimension are extracted to obtain the temporal features of each dimension; the temporal feature reconstruction model is to replace the multilayer perceptron in the Transformer model with a Kolmogorov-Arnold network.

5. The method for detecting anomalies in small-sample multidimensional time series according to claim 4, characterized in that, Based on the frequency domain feature reconstruction model, the contextual information of the enhanced time series in each dimension is captured to obtain the frequency domain features of each dimension, specifically including: For any enhanced time series, the enhanced time series is divided into multiple frequency blocks according to frequency; Perform discrete cosine transform on each frequency block to obtain the frequency domain feature representation of each frequency block; The frequency domain feature representations of each frequency block are concatenated to obtain the frequency channel vector; A compressed excitation mechanism is used to capture the global context information of the enhanced time series, and frequency domain attention weights are obtained. Based on the frequency domain attention weights and the frequency channel vectors, the frequency domain features of the corresponding dimensions are determined.

6. The method for detecting anomalies in small-sample multidimensional time series according to claim 1, characterized in that, Based on the initial time series and reconstructed time series for each dimension, the anomaly detection results are determined, specifically including: Based on the initial time series of each dimension, the reconstructed time series of each dimension is denormalized to obtain the reconstructed data for each dimension. Based on the reconstructed data for each dimension and the initial time series for each dimension, calculate the anomaly score for each time point; Based on the anomaly score at each time point, an anomaly label is determined for each time point to obtain the anomaly detection result; the anomaly label is either abnormal or normal.

7. The method for detecting anomalies in small-sample multidimensional time series according to claim 1, characterized in that, The method is used to detect anomalies in the performance of a server cluster; the few-sample multidimensional time series includes multiple performance indicators of the servers in the server cluster at each time point within a set time period, with one dimension corresponding to one performance indicator.

8. A few-sample, multi-dimensional time series anomaly detection system, characterized in that, The system is applied to the few-sample multidimensional time series anomaly detection method according to any one of claims 1-7, the system comprising: The preprocessing module is used to divide a small number of multidimensional time series into blocks to obtain the initial time series for each dimension. The anomaly masking module is used to mask the anomalies of the initial time series for each dimension, resulting in an anomaly masked time series for each dimension. The data augmentation module is used to perform data augmentation on the anomaly occlusion time series of each dimension using a pre-trained language model, so as to obtain the augmented time series of each dimension. The time-frequency feature extraction module is used to extract features from the enhanced time series of each dimension from the frequency dimension and the time dimension, respectively, to obtain the frequency domain features and time domain features of each dimension; The reconstruction module is used to aggregate the frequency domain features and time domain features of each dimension to obtain the reconstructed time series of each dimension. The results determination module is used to determine the anomaly detection results based on the initial time series and the reconstructed time series for each dimension.

9. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the few-sample multidimensional time series anomaly detection method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the method for detecting anomalies in a few-sample multidimensional time series as described in any one of claims 1-7.