Data anomaly detection method and system based on low-dimensional embedding time sequence characteristics
By using variational autoencoder and long-term memory network model in data abnormality detection and detection based on low-dimensional embedded timing characteristics, the problem of large computing resources consumption and difficult to guarantee detection accuracy in the prior art is solved, and efficient and accurate data abnormality detection is achieved.
Patent Information
- Application Number
- CN202510159322.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-23
AI Technical Summary
When using machine learning to detect data abnormalities, the prior art faces the problem of high computing resources and difficult to guarantee detection accuracy.
The data anomaly detection method based on low-dimensional embedding timing characteristics is adopted, and efficient detection of data anomaly is achieved through variational autoencoder and long-term short-term memory network model. This method summarizes the local information of the short window into low-dimensional embedding, and implements the low-dimensional embedding through the long-term short-term memory network model to achieve long-term trend perception.
It improves the accuracy and efficiency of data abnormality detection, effectively saves computing resources, and can more accurately detect data abnormalities that occur in the short term and long term.
Smart Images

Figure CN120030479A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a data anomaly detection method and system based on low-dimensional embedded time series features. Background Art
[0002] In today's information society, data has become an important basis for corporate decision-making. The quality of data is directly related to the accuracy and effectiveness of decision-making. However, as the amount of data continues to grow, how to efficiently and accurately detect data anomalies has become a challenge. Traditional data anomaly detection methods rely on predefined rules and manual review, which are inefficient when processing large-scale data sets and difficult to adapt to the diversity and complexity of data. In order to solve these problems, a large number of researchers have explored the use of machine learning to achieve automated detection of data anomalies.
[0003] In the prior art, when using machine learning to realize automatic detection of data anomalies, it often faces the problem of repeated training and learning for different data, which consumes a lot of computing resources; at the same time, the accuracy of data anomaly detection is difficult to guarantee.
[0004] How to solve the above technical problems is the subject faced by the present invention. Summary of the invention
[0005] In order to address the deficiencies in the prior art, the present invention provides a data anomaly detection method and system based on low-dimensional embedded time series features, which has high accuracy, high efficiency and can effectively save computing resources.
[0006] The technical solution adopted by the present invention to solve the technical problem is: on the one hand, the present invention provides a data anomaly detection method based on low-dimensional embedded time series features, comprising the following steps:
[0007] S1, obtain feature data set;
[0008] S2. Define data anomaly detection model;
[0009] The data anomaly detection model includes a variational autoencoder and a long short-term memory network;
[0010] The variational autoencoder model summarizes the local information of a short window into a low-dimensional embedding, and the long short-term memory network model acts on the low-dimensional embedding generated by the variational autoencoder to achieve long-term trend perception and effectively detect data anomalies occurring in the short and long term.
[0011] S3, training data anomaly detection model;
[0012] S4, performing anomaly detection evaluation on the data anomaly detection model;
[0013] S5. Automatically orchestrate the data anomaly detection process;
[0014] S6. Run the detection process and mark abnormal data.
[0015] The anomaly detection process can run independently or be called through a custom database function.
[0016] Preferably, the step S1 specifically comprises:
[0017] S11, add data source, define basic information of data set and time window, and generate multi-sequence data segments;
[0018] S12, extracting structural information of the data set to obtain basic features of the data set;
[0019] S13: Pre-clean or filter the data set to generate a feature data set.
[0020] Preferably, the step S2 is specifically:
[0021] S21. Define a method for constructing a data anomaly detection model, define an encoder, a decoder, and a long short-term memory network, and initialize each parameter;
[0022] S22. Define an encoder encoding method, a decoder decoding method and a data forward propagation method.
[0023] The forward propagation logic method arranges the data flow and calculation process of each network layer of the defined encoder and decoder to realize the reverse flow of gradients in the network layer and the update of model parameters in the machine learning process.
[0024] Preferably, the step S3 is specifically:
[0025] S31, define a loss function;
[0026] The loss function is:
[0027] loss elbo =loss KL (p∥q)+loss reconstruction (x,y),
[0028] In the formula, loss elbo is the total loss, loss KL (p∥q) is the KL divergence loss, loss reconstruction (x,y) is the reconstruction loss;
[0029] KL divergence loss calculation: According to the KL divergence formula: Among them, p(x) is the original distribution and q(x) is the predicted distribution. In actual calculation, KL divergence loss is first performed on each dimension of the latent space, then the divergence of the latent space dimension of each sample is summed, and finally the mean of the KL divergence loss of all samples is output.
[0030] Calculation of reconstruction loss: Negative log-likelihood loss is used to calculate the reconstruction loss. According to the negative log-likelihood loss formula: Among them, x is the reconstructed data, y is the original data, and the reconstruction loss is achieved by calculating the negative log-likelihood loss of input x and y.
[0031] The reconstruction loss method is called to calculate the difference between the reconstructed data and the original data, and the KL divergence loss method is called to calculate the difference between the distribution of the latent variable and the prior distribution.
[0032] S32. Define multiple time windows in the feature data set, and select matching continuous data segments according to the time windows for association, and train the data anomaly detection model multiple times.
[0033] Preferably, the step S32 is specifically:
[0034] A1. Set the model to training mode and initialize the model hidden state; A2. Load data in batches through the training data loader, obtain data and its labels, and transfer the data and labels to the device where the model is located; A3. The model receives input data and the corresponding sequence length, and then performs forward propagation to generate parameters for reconstructing the data, the mean and logarithmic variance of the latent variables. After each iteration, the hidden state is separated from the computational graph; A4. Call the loss function to calculate the total loss, perform backpropagation on the model parameters based on the total loss to calculate the gradient, and perform gradient clipping on the gradient; A5. The optimizer updates the model parameters according to the gradient. After each parameter update, the gradient information is cleared to prepare for the next iteration; A6. After the training iteration, the training loss array containing the loss value of each batch is returned.
[0035] Preferably, the step S4 is specifically:
[0036] S41, defining anomaly detection score function;
[0037] The anomaly detection score function is the prediction error between the reconstructed window prediction data sequence and the original data sequence, calculated by the loss function;
[0038] To facilitate the actual calculation of the score function, since the prediction error between the reconstructed window predicted data sequence and the original data sequence is linearly related to the KL divergence loss and the reconstruction loss, the prediction error is defined as KL divergence loss + reconstruction loss.
[0039] S42. Define multiple time windows in the feature data set, and select matching continuous data segments according to the time windows for association, evaluate the data anomaly detection model multiple times, and obtain model weights of multiple sequences.
[0040] By utilizing the model weights of multiple sequences, the characteristic changes of the test data in the multi-window time series are strengthened, thereby improving the accuracy of data anomaly detection.
[0041] Preferably, the step S42 is specifically:
[0042] B1. Set the model to evaluation mode and initialize the model hidden state; B2. Load data in batches through the data loader, obtain data and its labels, and transfer the data and labels to the device where the model is located; B3. The model receives input data and the corresponding sequence length, and then performs forward propagation to generate parameters for reconstructed data, the mean and logarithmic variance of latent variables;
[0043] B4. Call the loss function to calculate the anomaly detection score function and set the prediction error threshold based on the calculation result; B5. Extract multiple data segments at a certain moment for anomaly detection evaluation to obtain the optimal prediction error threshold at that moment; B6. After the evaluation iteration is completed, return the training loss array containing the loss value of each batch.
[0044] Preferably, the data anomaly detection process in step S5 includes creating a new anomaly detection process, generating a time window, loading model weights, reading data to be tested, anomaly detection and anomaly marking.
[0045] Preferably, the step S6 is specifically:
[0046] Load the data anomaly detection model weights, read the data to be tested at a given time, call the loss function to calculate the anomaly detection score function; read the best prediction error threshold at a given time, determine whether the data is abnormal based on the calculation results, and mark the abnormal data.
[0047] On the other hand, the present invention provides a data anomaly detection system based on low-dimensional embedded time series features, including a data management module for generating and managing data sets required for training and evaluating data anomaly detection models.
[0048] To detect data from various data sources, the data management module supports seamless integration of multiple data sources. Whether it is structured data (such as table data in SQL database) or unstructured data (text), the system can quickly integrate. Through unified management and automatic synchronization of data sources, users can easily introduce different data sources and create comprehensive data sets for model training and anomaly detection.
[0049] Model management module, used to define data anomaly detection models and manage data anomaly detection model weights;
[0050] The model management module introduces a version control function, which allows users to manage each model release in a versioned manner, and users can select any historical version for testing or deployment. Version control ensures that when problems occur, they can quickly roll back to a stable version, thereby ensuring the reliability of the system.
[0051] Detection management module, used for abnormal detection process definition, process scheduling execution, status monitoring, execution log recording and abnormal process alarm management;
[0052] In actual applications, anomaly detection tasks often need to run concurrently in a multi-node, multi-threaded environment. The detection management module also includes real-time task scheduling and dynamic allocation functions, which can dynamically allocate detection tasks according to task priority, data volume, and system resource conditions. Users can define the execution plan of tasks through the interface, including task start time, detection frequency, number of parallel executions, and other parameters. The system will automatically generate task scheduling plans based on these settings and execute them efficiently in the background. Configure scheduling rules to issue real-time alerts for abnormal task scheduling.
[0053] The training and evaluation module is used to train and evaluate the data anomaly detection model and generate training and evaluation reports;
[0054] In order to provide timely feedback on the results of each training and model evaluation, the system also includes a training evaluation report generation function. After each training and evaluation, the system will automatically generate a detailed report covering parameter selection during the training process, model performance evaluation, abnormal data analysis, etc. The report is presented in the form of charts and text, allowing users to quickly understand the performance of the model.
[0055] Detection execution module, used for data anomaly detection task execution, logging, status monitoring and data uploading;
[0056] The anomaly analysis module is used to calculate the anomaly detection score function, compare the calculation result with the prediction error threshold, determine whether the data is abnormal, analyze the abnormal data returned by the detection task, and generate an abnormal data report based on a predefined format.
[0057] In actual applications, the next step in identifying abnormal data is to analyze and find the cause of the abnormality, so the abnormal analysis module also provides a multi-dimensional analysis function, which supports users to conduct in-depth query and analysis of abnormal data from multiple dimensions (such as time, space, attribute characteristics, etc.). For example, users can view the distribution of abnormal data over a period of time according to the time dimension, and can also perform detailed analysis according to data attribute characteristics (such as equipment, geographic location, etc.) to help users discover potential risk factors in the data. After the abnormal data analysis is completed, users can customize the report format according to their needs and select the analysis dimensions and information that need to be highlighted. Users can also predefine push templates and rules, and push abnormal reports to relevant responsible persons in real time through emails or message notifications to ensure timely communication and processing of abnormal information.
[0058] The beneficial effects of the present invention are: data anomaly detection has high accuracy, high efficiency and can effectively save computing resources. A variational encoder is used to capture the low-dimensional embedding features of data in the latent space on a short window, and a long short-term memory network is used to act on the low-dimensional embedding output of the variational encoder, thereby perceiving the long-term correlation of the data sequence to be tested in the time series, and abnormal data can be detected more accurately. The data anomaly detection model is trained and evaluated, and process orchestration technology is used to realize the automated iteration, execution and output of the anomaly detection process, effectively improving the efficiency of data anomaly detection. The data is pre-cleaned and screened and marked to reduce the scale of the training and evaluation data sets, effectively save computing resources, and improve the convergence speed of the model. At the same time, multi-window weights are used to strengthen the characteristic changes of the applicable data to be tested in the multi-window time series, thereby improving the accuracy of data anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a diagram of the method steps of the present invention.
[0060] Figure 2 It is a system block diagram of the present invention. DETAILED DESCRIPTION
[0061] In order to clearly illustrate the technical features of this solution, this solution is described below through specific implementation methods.
[0062] See also Figure 1 As shown, this embodiment provides a data anomaly detection method based on low-dimensional embedded time series features, including the following steps:
[0063] S1, obtain feature data set;
[0064] S2. Define data anomaly detection model;
[0065] The data anomaly detection model includes a variational autoencoder and a long short-term memory network;
[0066] The variational autoencoder model summarizes the local information of a short window into a low-dimensional embedding, and the long short-term memory network model acts on the low-dimensional embedding generated by the variational autoencoder to achieve long-term trend perception and effectively detect data anomalies occurring in the short and long term.
[0067] S3, training data anomaly detection model;
[0068] S4, performing anomaly detection evaluation on the data anomaly detection model;
[0069] S5. Automatically orchestrate the data anomaly detection process;
[0070] S6. Run the detection process and mark abnormal data.
[0071] Step S1 is specifically as follows:
[0072] S11, add data source, define basic information of data set and time window, and generate multi-sequence data segments;
[0073] S12, extracting structural information of the data set to obtain basic features of the data set;
[0074] S13: Pre-clean or filter the data set to generate a feature data set.
[0075] In the embodiment of the present application, during the data set definition process, the user first needs to enter the basic information of the data set in the data management module, and then form a complete data set by screening marks or uploading data. For example, a terminal monitoring data set is shown in the following table.
[0076]
[0077] Table 1. Definition of basic information of dataset
[0078]
[0079] Table 2 does not contain abnormal continuous data segments
[0080]
[0081] Table 3 contains abnormal continuous data segments
[0082] Step S2 is specifically as follows:
[0083] S21. Define a method for constructing a data anomaly detection model, define an encoder, a decoder, and a long short-term memory network, and initialize each parameter;
[0084] In the embodiment of the present application, the input feature dimension is defined as 200, the dimension of the latent space is defined as 16, the hidden layer size is defined as 256, and the number of long short-term memory network layers is defined as 1.
[0085] S22. Define an encoder encoding method, a decoder decoding method and a data forward propagation method.
[0086] Define the encoder method to encode the original data input sequence into the latent space. The input parameters include the packed input sequence pack_x and the initial hidden state of the encoder. The specific data processing flow is as follows:
[0087] Use the long short-term memory network to process the packed input sequence pack_x and update the encoder hidden state to obtain the packed output sequence out_pack;
[0088] Pass the updated hidden state to two fully connected layers to estimate the mean and logarithmic variance log_var of the latent space and calculate the standard deviation std;
[0089] The packed output sequence out_pack is converted back to the padded sequence to obtain the batch size batch_size;
[0090] Generate a Gaussian noise with the same size as the batch size batch_size and the size of the latent space. By multiplying the Gaussian noise with the standard deviation std and adding the mean mean, we can get the sample z of the latent space;
[0091] After processing, the sample z of the latent space, the mean mean of the latent space, the logarithmic variance log_var of the latent space, and the updated hidden state are returned.
[0092] Define the decoder method to convert the data from the latent space to the original space. The input parameters include the sample z generated from the latent space and the packed input sequence pack_x. The specific data processing flow is as follows:
[0093] Initialize the decoder’s hidden state using the sample z from the input latent space;
[0094] Pass the packed input sequence pack_x and the initial hidden state into the LSTM network to output the packed output sequence out_pack and the updated hidden state;
[0095] In order to improve the stability of model training, the packed output sequence out_pack is input into the linear layer to obtain the linear layer output out_linear, and then the activation function is used to convert the linear layer output out_linear into the logarithmic probability distribution log_x of the output sequence;
[0096] After processing the decoder input parameters, it returns the logarithmic probability distribution log_x of the output sequence.
[0097] Define the forward propagation method to implement the data flow and calculation logic of each network layer. The input parameters include the input sequence x and the hidden state. The specific data processing flow is as follows:
[0098] Get the embedded representation of the input sequence x in the mapping vocabulary, and obtain the packed embedded sequence pack_x after variable-length sequence processing and packing;
[0099] Input the packed embedding sequence pack_x and the hidden state into the decoder method defined in step B2 for processing, and obtain the sample z of the latent space, the mean mean of the latent space, the logarithmic variance log_var of the latent space, and the updated hidden state;
[0100] Input the sample z of the latent space and the packed embedding sequence pack_x into the decoder method defined in step B3 to obtain the output sequence x_hat.
[0101] After processing, the sequence x_hat, the mean of the latent space mean, the logarithmic variance of the latent space log_var, the sample z of the latent space, and the updated encoder hidden state are returned.
[0102] In the embodiments of the present application, the machine learning network in the data anomaly detection model can be implemented using any machine learning framework.
[0103] Step S3 is specifically as follows:
[0104] S31, define a loss function;
[0105] The loss function is:
[0106] loss elbo =loss KL (p∥q)+loss reconstruction (x,y),
[0107] In the formula, loss elbo is the total loss, loss KL (p∥q) is the KL divergence loss, loss reconstruction (x,y) is the reconstruction loss;
[0108] KL divergence loss calculation: According to the KL divergence formula: Among them, p(x) is the original distribution and q(x) is the predicted distribution. In actual calculation, KL divergence loss is first performed on each dimension of the latent space, then the divergence of the latent space dimension of each sample is summed, and finally the mean of the KL divergence loss of all samples is output.
[0109] Calculation of reconstruction loss: Negative log-likelihood loss is used to calculate the reconstruction loss. According to the negative log-likelihood loss formula: Among them, x is the reconstructed data, y is the original data, and the reconstruction loss is achieved by calculating the negative log-likelihood loss of input x and y.
[0110] The reconstruction loss method is called to calculate the difference between the reconstructed data and the original data, and the KL divergence loss method is called to calculate the difference between the distribution of the latent variable and the prior distribution.
[0111] S32. Define multiple time windows in the feature data set, and select matching continuous data segments according to the time windows for association, and train the data anomaly detection model multiple times.
[0112] Step S32 is specifically as follows:
[0113] A1. Set the model to training mode and initialize the model hidden state; A2. Load data in batches through the training data loader, obtain data and its labels, and transfer the data and labels to the device where the model is located; A3. The model receives input data and the corresponding sequence length, and then performs forward propagation to generate parameters for reconstructing the data, the mean and logarithmic variance of the latent variables. After each iteration, the hidden state is separated from the computational graph; A4. Call the loss function to calculate the total loss, perform backpropagation on the model parameters based on the total loss to calculate the gradient, and perform gradient clipping on the gradient; A5. The optimizer updates the model parameters according to the gradient. After each parameter update, the gradient information is cleared to prepare for the next iteration; A6. After the training iteration, the training loss array containing the loss value of each batch is returned.
[0114] In an embodiment of the present application, the unsupervised training process starts with instantiating a data anomaly detection model, selecting a continuous data segment at time t of a feature data set, loading it into a data loader, instantiating a loss function, instantiating an Adm optimizer, and calling a step loss function to start multiple rounds of unsupervised training. Through multiple rounds of training, the training data is input into each network layer of the data anomaly detection model in batches, and the model parameter weights are calculated and updated through the forward propagation method, and dynamically adjusted according to the loss function and optimizer. For example, the loss output of 40 rounds of training for the continuous data segment that does not contain anomalies in Table 2 is shown in the following table:
[0115]
[0116]
[0117] Table 4. Loss rate of 40 rounds of training for continuous data segments without abnormalities
[0118] The continuous data segments without anomalies and the continuous data segments containing anomalies associated with multiple time windows are the same as Table 2 and Table 3 in step S11, and the time windows and data segments are associated using a many-to-many relationship. For example, the time windows and data segments are associated as shown in the following table:
[0119]
[0120] Table 5 contains abnormal continuous data segments
[0121]
[0122]
[0123] Table 6 contains abnormal continuous data segments
[0124] Step S4 is specifically as follows:
[0125] S41, defining anomaly detection score function;
[0126] The anomaly detection score function is the prediction error between the reconstructed window prediction data sequence and the original data sequence, calculated by the loss function;
[0127] To facilitate the actual calculation of the score function, since the prediction error between the reconstructed window predicted data sequence and the original data sequence is linearly related to the KL divergence loss and the reconstruction loss, the prediction error is defined as KL divergence loss + reconstruction loss.
[0128] S42. Define multiple time windows in the feature data set, and select matching continuous data segments according to the time windows for association, evaluate the data anomaly detection model multiple times, and obtain model weights of multiple sequences.
[0129] By utilizing the model weights of multiple sequences, the characteristic changes of the test data in the multi-window time series are strengthened, thereby improving the accuracy of data anomaly detection.
[0130] Step S42 is specifically as follows:
[0131] B1. Set the model to evaluation mode and initialize the model hidden state; B2. Load data in batches through the data loader, obtain data and its labels, and transfer the data and labels to the device where the model is located; B3. The model receives input data and the corresponding sequence length, and then performs forward propagation to generate parameters for reconstructed data, the mean and logarithmic variance of latent variables; B4. Call the loss function to calculate the anomaly detection score function, and set the prediction error threshold based on the calculation result; B5. Extract multiple data segments at a certain moment for anomaly detection evaluation to obtain the optimal prediction error threshold at that moment; B6. After the evaluation iteration is completed, return the training loss array containing the loss value of each batch.
[0132] The data anomaly detection process in step S5 includes creating a new anomaly detection process, generating a time window, loading model weights, reading data to be tested, anomaly detection, and anomaly marking.
[0133] The automated process arrangement can be implemented by writing scripts or using any workflow arrangement. In one example of the present application, an automatic execution script is uploaded through the detection management module, and the scheduling detection process is implemented through the script.
[0134] Step S6 is specifically as follows:
[0135] Load the data anomaly detection model weights, read the data to be tested at a given time, call the loss function to calculate the anomaly detection score function; read the best prediction error threshold at a given time, determine whether the data is abnormal based on the calculation results, and mark the abnormal data.
[0136] In one example of an embodiment of the present application, anomaly detection is performed on the following test data sequence of device ZT8669235111012.
[0137] For example, the data sequence to be tested for the device is shown in the following table.
[0138]
[0139] Table 7, original continuous data segment
[0140] In the embodiment of the present application, if the time range of the Wt_1 window is defined as 10:00-12:00, a set of data D corresponding to the time range needs to be selected according to the Wt_1 window. t_1 , as the input parameter to call the custom window function (UDWF), start the anomaly detection process, and then pass the parameters to the anomaly detection score function. According to the anomaly detection data processing process, the model weights are automatically adapted and loaded according to the time window Wt_1, and then D t_1 Perform anomaly detection, output the check_losses array, and mark the abnormal data by the abnormal threshold. For example, if the error threshold θ is defined t_1 =0.7, by comparing the prediction error of the anomaly detection output check_losses, the output anomaly mark is shown in the following table.
[0141]
[0142]
[0143] Table 8. Anomaly markers for anomaly detection in a certain time window
[0144] Execute the above anomaly detection process to realize anomaly detection of the data sequence related to the selected time window Wt_1, and finally output the anomaly mark.
[0145] In the embodiment of the present application, the network scale of the defined model is relatively small, wherein the number of latent space linear layers (VAE) is 16 (i.e., the dimension of the latent space is 16), the number of long short-term correlation (LSTM) network layers is 1, and the number of hidden layers is 256. Using the training data in Table 2, under limited resources (CPU 4 core 8G memory hardware resources), 40 rounds of unsupervised training can be completed within 10 to 15 minutes, achieving the training effect in Table 4 and the application effect in Table 8. It can be seen from the scale of the training data set and the training time that the method used in the example of the present application has a faster convergence speed and consumes less computing resources than other similar machine learning methods.
[0146] See also Figure 2 As shown, this embodiment also provides a data anomaly detection system based on low-dimensional embedded time series features, including a data management module for generating and managing data sets required for data anomaly detection model training and evaluation.
[0147] To detect data from various data sources, the data management module supports seamless integration of multiple data sources. Whether it is structured data (such as table data in SQL database) or unstructured data (text), the system can quickly integrate. Through unified management and automatic synchronization of data sources, users can easily introduce different data sources and create comprehensive data sets for model training and anomaly detection.
[0148] Model management module, used to define data anomaly detection models and manage data anomaly detection model weights;
[0149] The model management module introduces a version control function, which allows users to manage each model release in a versioned manner, and users can select any historical version for testing or deployment. Version control ensures that when problems occur, they can quickly roll back to a stable version, thereby ensuring the reliability of the system.
[0150] Detection management module, used for abnormal detection process definition, process scheduling execution, status monitoring, execution log recording and abnormal process alarm management;
[0151] In actual applications, anomaly detection tasks often need to run concurrently in a multi-node, multi-threaded environment. The detection management module also includes real-time task scheduling and dynamic allocation functions, which can dynamically allocate detection tasks according to task priority, data volume, and system resource conditions. Users can define the execution plan of tasks through the interface, including task start time, detection frequency, number of parallel executions, and other parameters. The system will automatically generate task scheduling plans based on these settings and execute them efficiently in the background. Configure scheduling rules to issue real-time alerts for abnormal task scheduling.
[0152] The training and evaluation module is used to train and evaluate the data anomaly detection model and generate training and evaluation reports;
[0153] In order to provide timely feedback on the results of each training and model evaluation, the system also includes a training evaluation report generation function. After each training and evaluation, the system will automatically generate a detailed report covering parameter selection during the training process, model performance evaluation, abnormal data analysis, etc. The report is presented in the form of charts and text, allowing users to quickly understand the performance of the model.
[0154] Detection execution module, used for data anomaly detection task execution, logging, status monitoring and data uploading;
[0155] The anomaly analysis module is used to calculate the anomaly detection score function, compare the calculation result with the prediction error threshold, determine whether the data is abnormal, analyze the abnormal data returned by the detection task, and generate an abnormal data report based on a predefined format.
[0156] In actual applications, the next step in identifying abnormal data is to analyze and find the cause of the abnormality, so the abnormal analysis module also provides a multi-dimensional analysis function, which supports users to conduct in-depth query and analysis of abnormal data from multiple dimensions (such as time, space, attribute characteristics, etc.). For example, users can view the distribution of abnormal data over a period of time according to the time dimension, and can also perform detailed analysis according to data attribute characteristics (such as equipment, geographic location, etc.) to help users discover potential risk factors in the data. After the abnormal data analysis is completed, users can customize the report format according to their needs and select the analysis dimensions and information that need to be highlighted. Users can also predefine push templates and rules, and push abnormal reports to relevant responsible persons in real time through emails or message notifications to ensure timely communication and processing of abnormal information.
[0157] Technical features not described in the present invention can be achieved through or by adopting existing technologies and will not be described in detail here. Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by ordinary technicians in this technical field within the essential scope of the present invention should also fall within the scope of protection of the present invention.
Claims
1. A data anomaly detection method based on low-dimensional embedded time series features, characterized in that: The following steps are involved: S1, obtain feature data set; S2. Define data anomaly detection model; The data anomaly detection model includes a variational autoencoder and a long short-term memory network; S3, training data anomaly detection model; S4, performing anomaly detection evaluation on the data anomaly detection model; S5. Automatically orchestrate the data anomaly detection process; S6. Run the detection process and mark abnormal data.
2. The data anomaly detection method based on low-dimensional embedded time series features according to claim 1 is characterized in that: The step S1 is specifically as follows: S11, add data source, define basic information of data set and time window, and generate multi-sequence data segments; S12, extracting structural information of the data set to obtain basic features of the data set; S13: Pre-clean or filter the data set to generate a feature data set.
3. The data anomaly detection method based on low-dimensional embedded time series features according to claim 1 is characterized in that: The step S2 is specifically as follows: S21. Define a method for constructing a data anomaly detection model, define an encoder, a decoder, and a long short-term memory network, and initialize each parameter; S22. Define an encoder encoding method, a decoder decoding method and a data forward propagation method.
4. The data anomaly detection method based on low-dimensional embedded time series features according to claim 1 is characterized in that: The step S3 is specifically as follows: S31, define a loss function; The loss function is: loss elbo =loss KL (p∥q)+loss reconstruction (x,y), In the formula, loss elbo is the total loss, loss KL (p∥q) is the KL divergence loss, loss reconstruction (x,y) is the reconstruction loss; S32. Define multiple time windows in the feature data set, and select matching continuous data segments according to the time windows for association, and train the data anomaly detection model multiple times.
5. The data anomaly detection method based on low-dimensional embedded time series features according to claim 4 is characterized in that: The step S32 is specifically as follows: A1. Set the model to training mode and initialize the model hidden state; A2. Load data in batches through the training data loader, obtain data and its labels, and transmit the data and labels to the device where the model is located; A3. The model receives input data and the corresponding sequence length, and then performs forward propagation to generate parameters for reconstructed data, the mean and logarithmic variance of latent variables. After each iteration, the hidden state is separated from the computational graph; A4. Call the loss function to calculate the total loss, perform backpropagation on the model parameters based on the total loss to calculate the gradient, and perform gradient clipping on the gradient; A5. The optimizer updates the model parameters according to the gradient. After each parameter update, the gradient information is cleared to prepare for the next iteration. A6. After the training iteration is completed, the training loss array containing the loss value of each batch is returned.
6. The data anomaly detection method based on low-dimensional embedded time series features according to claim 1 is characterized in that: The step S4 is specifically as follows: S41, defining anomaly detection score function; The anomaly detection score function is the prediction error between the reconstructed window prediction data sequence and the original data sequence, calculated by the loss function; S42. Define multiple time windows in the feature data set, and select matching continuous data segments according to the time windows for association, and evaluate the data anomaly detection model multiple times.
7. The data anomaly detection method based on low-dimensional embedded time series features according to claim 6 is characterized in that: The step S42 is specifically as follows: B1. Set the model to evaluation mode and initialize the model hidden state; B2. Load data in batches through the data loader, obtain the data and its labels, and transfer the data and labels to the device where the model is located; B3. The model receives input data and the corresponding sequence length, and then performs forward propagation to generate parameters for reconstructed data, the mean and logarithmic variance of latent variables; B4. The loss function is called to calculate the anomaly detection score function, and the prediction error threshold is set according to the calculation result; B5. Multiple data segments at a certain moment are extracted for anomaly detection evaluation to obtain the optimal prediction error threshold at that moment; B6. After the evaluation iteration is completed, the training loss array containing the loss value of each batch is returned.
8. The data anomaly detection method based on low-dimensional embedded time series features according to claim 1 is characterized in that: The data anomaly detection process in step S5 includes creating a new anomaly detection process, generating a time window, loading model weights, reading data to be tested, anomaly detection and anomaly marking.
9. The data anomaly detection method based on low-dimensional embedded time series features according to claim 1 is characterized in that: The step S6 is specifically as follows: Load the data anomaly detection model weights, read the data to be tested at a given time, call the loss function to calculate the anomaly detection score function; read the best prediction error threshold at a given time, determine whether the data is abnormal based on the calculation results, and mark the abnormal data.
10. A data anomaly detection system based on low-dimensional embedded time series features, characterized in that: include: Data management module, used to generate and manage data sets required for training and evaluating data anomaly detection models; Model management module, used to define data anomaly detection models and manage data anomaly detection model weights; Detection management module, used for abnormal detection process definition, process scheduling execution, status monitoring, execution log recording and abnormal process alarm management; The training and evaluation module is used to train and evaluate the data anomaly detection model and generate training and evaluation reports; Detection execution module, used for data anomaly detection task execution, logging, status monitoring and data uploading; The anomaly analysis module is used to calculate the anomaly detection score function, compare the calculation result with the prediction error threshold, determine whether the data is abnormal, analyze the abnormal data returned by the detection task, and generate an abnormal data report based on a predefined format.
Citation Information
Cited By
Nor Flash access anomaly detection method and system based on LSTM auto-encoder
CN121071950A