Time series data anomaly detection method and apparatus, and non-volatile storage medium
By obtaining the detection task configuration information and training data sets, the abnormal detection model is trained, which solves the abnormal monitoring false alarm problem caused by threshold judgment in industrial scenarios, and realizes accurate abnormal detection without setting thresholds.
Patent Information
- Application Number
- PCT/CN2024/125749
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-10-18
- Publication Date
- 2025-07-03
AI Technical Summary
When the prior art faces a large number of monitoring indicators, it is difficult to set reasonable thresholds, resulting in frequent abnormal monitoring false positives, and it is impossible to effectively judge whether the data is abnormal, especially in industrial scenarios.
By obtaining detection task configuration information, establishing detection tasks, obtaining timing data and determining the training data set, and using the training data set to train the abnormal detection model to achieve abnormal detection without setting a threshold.
It realizes accurate judgment of whether the data is abnormal in scenarios with complex data types, reduces model detection tasks, prevents false alarms, and has universality and automated detection capabilities.
Smart Images

Figure CN2024125749_03072025_PF_FP_ABST
Abstract
Description
Time series data anomaly detection method, device and non-volatile storage medium
[0001] Related applications
[0002] This application claims priority to Chinese patent application number 2023118165084, filed on December 26, 2023, entitled “Time Series Data Anomaly Detection Method, Device and Non-volatile Storage Medium,” the entire text of which is hereby incorporated by reference. Technical Field
[0003] The present application relates to the field of data detection, and more specifically, to a method and device for detecting anomalies in time series data, and a non-volatile storage medium. Background Art
[0004] In related technologies, a common approach for detecting data anomalies is to use thresholds to determine whether the data being examined is anomalous. The problem with this approach is that in real industrial applications, the number of data indicators that need to be monitored is excessive, potentially tens of thousands or even millions. This makes it difficult for operations personnel to set reasonable thresholds for each indicator based on their business experience. Furthermore, some indicators have diverse data forms, making simple fixed thresholds incapable of determining whether such data is anomalous. Therefore, related technologies cannot effectively detect anomalous data in scenarios with a large number of indicators.
[0005] To address the above-mentioned problems, no effective solutions have been proposed so far.
[0006] Summary of the Invention
[0007] Embodiments of the present application provide a method, device, and non-volatile storage medium for detecting anomalies in time series data.
[0008] According to one aspect of an embodiment of the present application, a method for detecting anomalies in time series data is provided, including: obtaining detection task configuration information, and establishing a detection task based on the detection task configuration information, wherein the task configuration information includes model configuration information; when executing the detection task, obtaining time series data corresponding to the detection task, and determining a training data set based on the time series data; performing pre-detection on target data points to be detected based on the training data set; when the pre-detection result is abnormal, training an anomaly detection model based on the model configuration information and the training data set to obtain a trained anomaly detection model; and detecting the target data points to be detected based on the trained anomaly detection model.
[0009] In one embodiment, the target data point to be detected is a data point determined based on the time point to be detected.
[0010] In one embodiment, obtaining detection task configuration information and establishing a detection task based on the detection task configuration information includes: obtaining detection task configuration information with a status of effective; constructing a detection task based on the detection task configuration information with a status of effective, and setting the model configuration information to the task context of the detection task; delivering the detection task to a thread pool, and retrieving the detection task from the thread pool based on a scheduling period corresponding to the detection task.
[0011] In one embodiment, obtaining the detection task configuration information in the effective state includes: reading the detection task configuration information in the effective state from a database table, and determining the number of the detection task configuration information in the effective state.
[0012] In one embodiment, constructing the detection task according to the detection task configuration information in the effective state includes: constructing the detection task according to the detection task configuration information when the number of the detection task configuration information in the effective state is not zero.
[0013] In one embodiment, obtaining time series data corresponding to the detection task and determining a training data set based on the time series data includes: determining a data selection time interval; obtaining time series data corresponding to the data to be detected based on the data selection time interval; sampling the time series data to obtain a training data set.
[0014] In one embodiment, sampling time series data to obtain a training data set includes: sorting the time series data; deleting a preset proportion of data in the sorted time series data based on the sorting result to obtain an initial sampling result; and randomly sampling the initial sampling result a preset number of times to obtain a training data set.
[0015] In one embodiment, pre-detection of target data points to be detected based on a training data set includes: determining the mean and variance of the training data set; determining a first value interval and a second value interval based on the mean and variance of the training data set; pre-detection of the target data points to be detected based on the first value interval and the second value interval, wherein, when the target data points to be detected are within the first value interval, the detection result of the pre-detection is determined to be normal, and when the target data points to be detected are within the second value interval, the detection result is determined to be abnormal.
[0016] In one embodiment, training an anomaly detection model based on model configuration information and a training data set to obtain a trained anomaly detection model includes: constructing a request body based on task configuration information, wherein the request body includes a task identifier of the detection task, target data points to be detected, model configuration information, and a training data set; sending the request body to the anomaly detection model, and training the anomaly detection model based on the model configuration information and the training data set to obtain a trained anomaly detection model.
[0017] In one embodiment, a target data point to be detected is detected by a trained anomaly detection model, and a response body returned by the anomaly detection model is obtained, wherein the response body includes a task identifier of the detection task, the target data point to be detected, and the detection result.
[0018] In one embodiment, after the target data points to be detected are detected according to the trained anomaly detection model, the time series data anomaly detection method further includes: when the detection result is abnormal, obtaining a set of data points to be detected, wherein the data points to be detected in the set of data points to be detected are a preset number of data points collected continuously, and the latest data point to be detected corresponding to the time point in the set of data points to be detected is adjacent to the target data point to be detected, and the time point corresponding to any data point to be detected in the set of data points to be detected is earlier than the time point to be detected; determining the detection results of the data points to be detected in the set of data points to be detected; and when the detection results of the data points to be detected in the set of data points to be detected are all abnormal, generating and sending an alarm message according to a preset alarm template.
[0019] In one embodiment, the detection task configuration information includes whether it is a scheduled task, the task title, indicator information, data source, and data retrieval frequency.
[0020] According to another aspect of an embodiment of the present application, a time series data anomaly detection device is also provided, including: a first processing module, used to obtain detection task configuration information, and establish a detection task based on the detection task configuration information, wherein the task configuration information includes model configuration information; a second processing module, used to obtain time series data corresponding to the detection task when executing the detection task, and determine a training data set based on the time series data; a third processing module, used to pre-detect the target data points to be detected based on the training data set; a fourth processing module, used to train the anomaly detection model based on the model configuration information and the training data set when the pre-detection result is abnormal, to obtain the trained anomaly detection model; and detect the target data points to be detected based on the trained anomaly detection model.
[0021] According to another aspect of an embodiment of the present application, a non-volatile storage medium is provided, in which a program is stored. When the program is running, a device where the non-volatile storage medium is located is controlled to execute a method for detecting anomalies in time series data.
[0022] According to another aspect of an embodiment of the present application, an electronic device is provided, including: a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the method for detecting anomalies in time series data is executed when the program is run.
[0023] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.
[0025] FIG1 is a schematic structural diagram of a computer terminal (mobile terminal) provided according to an embodiment of the present application;
[0026] FIG2 is a flow chart of a method for detecting anomalies in time series data according to an embodiment of the present application;
[0027] FIG3 is a schematic diagram of a configuration flow of an anomaly detection task provided according to an embodiment of the present application;
[0028] FIG4 is a flow chart of a training data set acquisition process according to an embodiment of the present application;
[0029] FIG5 is a schematic diagram of a pre-detection process according to an embodiment of the present application;
[0030] FIG6 is a schematic diagram of a process for detecting target data points to be detected using a trained anomaly detection model according to an embodiment of the present application;
[0031] FIG7 is a schematic diagram of a process for determining whether a target data point to be detected is abnormal according to an embodiment of the present application;
[0032] FIG8 is a flow chart of a time series data anomaly detection process according to an embodiment of the present application;
[0033] FIG9 is a schematic structural diagram of a time series data anomaly detection device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0034] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0036] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:
[0037] 3σ test: In statistics, if a variable follows a normal distribution, and its mean is u and standard deviation is σ, then 99% of the data of the variable will fall within u±3σ, that is, the probability that the data is distributed in (u-3σ, u+3σ) is 99%. Therefore, when there is a data falling outside the mean (u) ± three times the standard deviation (3σ), it can be preliminarily regarded as abnormal data.
[0038] IForest (Isolation Forest): It is a fast anomaly detection algorithm based on ensemble learning, with linear time complexity and high accuracy. It is a state-of-the-art algorithm that meets the requirements of big data processing.
[0039] Indicator anomaly detection is widely used in the field of AIOps (intelligent operations and maintenance). Its purpose is to use algorithms to detect anomalies in the time series of KPIs (key performance indicators) and then notify operations and maintenance personnel of related risks through warnings. Indicator anomaly detection is also a prerequisite for other AIOps scenarios. Its detection results provide input information for subsequent scenarios such as alarm convergence, root cause location, and fault self-healing. Therefore, indicator anomaly detection is of great significance in the entire AIOps implementation scenario and the actual application of AIOps.
[0040] However, in actual industrial scenarios, many business systems usually need to monitor tens of thousands to millions of indicators, and general enterprises usually use fixed threshold methods to carry out abnormal monitoring. Due to the large amount of data, operation and maintenance personnel cannot set reasonable upper and lower thresholds for each indicator based on business experience. On the other hand, some monitoring indicators have rich forms from a data perspective (data types, such as stability, periodicity, trending, etc.), and simple fixed thresholds cannot determine whether such monitoring indicators are abnormal indicators. Therefore, in an industrial environment, the use of threshold judgment methods in related technologies will result in widespread abnormal monitoring false alarms, which will cause great trouble to operation and maintenance personnel and bring greater business risks to enterprise operations. Therefore, how to accurately and quickly issue abnormal alarms to the monitoring system is an urgent problem to be solved in the current AIOps field and many industrial enterprises.
[0041] In order to solve the above problems, relevant solutions are provided in the embodiments of the present application, which are described in detail below.
[0042] According to an embodiment of the present application, a method embodiment of a time series data anomaly detection method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0043] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a time series data anomaly detection method. As shown in Figure 1, the computer terminal 10 (or mobile device 10) may include one or more (102a, 102b, ..., 102n are used in the figure to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that the structure shown in Figure 1 is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may also include more or fewer components than shown in Figure 1, or have a configuration different from that shown in Figure 1.
[0044] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0045] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the time series data anomaly detection method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned time series data anomaly detection method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0046] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0047] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0048] In the above operating environment, an embodiment of the present application provides a method for detecting anomalies in time series data, as shown in FIG2 , which includes the following steps:
[0049] Step S202: acquiring detection task configuration information, and establishing a detection task according to the detection task configuration information, wherein the task configuration information includes model configuration information;
[0050] In the technical solution provided in step S202, the steps of obtaining detection task configuration information and establishing a detection task based on the detection task configuration information include: obtaining detection task configuration information with an effective status; constructing a detection task based on the detection task configuration information with an effective status, and setting the model configuration information to the task context of the detection task; delivering the detection task to the thread pool, and calling the detection task from the thread pool based on the scheduling cycle corresponding to the detection task.
[0051] Specifically, the detection task configuration information (JobConfig) of the detection task (Job) may include whether it is a scheduled task, task title, indicator information, data source, data retrieval frequency, etc. The model configuration information (ModelConfig) can be used for subsequent training of the anomaly detection model. The indicator information can be considered as data type information. The detection task configuration information in the effective state is the detection task configuration information that is in effect. The process of configuring anomaly detection tasks is shown in Figure 3, which includes the following steps:
[0052] Step S302: Read the detection task configuration information in the effective state from the database table, and determine the number of the detection task configuration information in the effective state;
[0053] Step S304: if the number of detection task configuration information in the effective state is not zero, construct a detection task according to the detection task configuration information, and set the model configuration information in the detection task configuration information into the task context;
[0054] Step S306: delivering the detection task to the thread pool, and retrieving the detection task from the thread pool according to the scheduling period corresponding to the detection task.
[0055] Step S204: when executing a detection task, obtaining time series data corresponding to the detection task, and determining a training data set based on the time series data;
[0056] In the technical solution provided in step S204, the process of obtaining the time series data corresponding to the detection task and determining the training data set based on the time series data is shown in FIG4 and includes the following steps:
[0057] Step S402, determining the data selection time interval;
[0058] In the technical solution provided in step S402, the data selection time interval can be determined based on the time point to be detected. For example, the time interval of the detection day can be determined to be [T-120min, T-1min], that is, the data point 120 minutes before the moment T, and the time interval of the previous days can be [T-90min, T+30min], that is, the data 90 minutes before the moment T and the data 30 minutes after the moment T, where T represents the time corresponding to the time point to be detected in each day. It is understandable that the number of days selected can be set by the user, and the duration of the time interval can also be set by the user.
[0059] Step S404, obtaining time series data corresponding to the data to be detected according to the data selection time interval;
[0060] Step S406: Sample the time series data to obtain a training data set.
[0061] As an optional implementation method, the steps of sampling time series data to obtain a training data set include: sorting the time series data; deleting a preset proportion of data in the sorted time series data based on the sorting result to obtain an initial sampling result; and randomly sampling the initial sampling result a preset number of times to obtain a training data set.
[0062] Specifically, after sorting time series data, the first M% of the data and the last N% of the data can be deleted to complete data cleaning. M and N can be set by the user. When sorting time series data, you can sort by data size or other preset sorting rules.
[0063] After the data cleaning is completed, a training data set can be obtained by randomly sampling the cleaned time series data a preset number of times.
[0064] Step S206, pre-detecting the target data points to be detected based on the training data set;
[0065] The target data point to be detected is a data point determined based on the time point to be detected.
[0066] In the technical solution provided in step S206, the process of pre-detecting the target data points to be detected based on the training data set is shown in FIG5 , and includes the following steps:
[0067] Step S502, determining the mean and variance of the training data set;
[0068] Step S504: determining a first value interval and a second value interval based on the mean and variance of the training data set;
[0069] Step S506, pre-detection is performed on the target data point to be detected based on the first value interval and the second value interval, wherein, when the target data point to be detected is located in the first value interval, the detection result of the pre-detection is determined to be normal, and when the target data point to be detected is located in the second value interval, the detection result is determined to be abnormal.
[0070] Specifically, 3σ detection can be performed on the target point CP based on the mean E and variance D of the training data. The first value interval mentioned above refers to [E-3σ, E+3σ], and the second value interval refers to the value interval in the training data set other than [E-3σ, E+3σ].
[0071] Step S208: When the pre-detection result is abnormal, the abnormality detection model is trained according to the model configuration information and the training data set to obtain a trained abnormality detection model.
[0072] Step S210: Detect the target data points to be detected based on the trained anomaly detection model.
[0073] Specifically, the above anomaly detection model may be an isolation forest detection model.
[0074] In the technical solution provided in steps S208-S210, the anomaly detection model is trained based on the model configuration information and the training data set to obtain a trained anomaly detection model; and the process of detecting the target data points to be detected based on the trained anomaly detection model is shown in FIG6 and includes the following steps:
[0075] Step S602: construct a request body based on the task configuration information, wherein the request body includes the task identifier of the detection task, the target data point to be detected, the model configuration information and the training data set;
[0076] Step S604: Send the request body to the anomaly detection model, and train the anomaly detection model based on the model configuration information and the training data set to obtain a trained anomaly detection model;
[0077] Step S606: Detect the target data point to be detected using the trained anomaly detection model, and obtain a response body returned by the anomaly detection model, wherein the response body includes the task identifier of the detection task, the target data point to be detected, and the detection result.
[0078] As an optional implementation, after the step of detecting the target data point to be detected according to the trained anomaly detection model, as shown in FIG7 , the target data point to be detected can also be further detected to ultimately determine whether the target data point to be detected is abnormal, including the following steps:
[0079] Step S702: If the detection result is abnormal, a set of data points to be detected is obtained, wherein the data points to be detected in the set are a preset number of data points collected continuously, and the latest data point to be detected in the set of data points to be detected is adjacent to the target data point to be detected, and the time point corresponding to any data point to be detected in the set of data points to be detected is earlier than the time point to be detected;
[0080] In the technical solution provided in step S702, the detection result can be determined based on the responding body. The above-mentioned preset number can be set by the user, for example, five.
[0081] Step S704, determining the detection results of the data points to be detected in the set of data points to be detected;
[0082] Step S706 : When the detection results of all the data points to be detected in the set of data points to be detected are abnormal, an alarm message is generated and sent according to a preset alarm template.
[0083] In summary, the complete execution flow of the time series data anomaly detection method provided in the embodiment of the present application is shown in FIG8 , and includes the following steps:
[0084] Step S802: Obtain task configuration information from a database, and construct a detection task according to the task configuration information;
[0085] Step S804: When executing the detection task, corresponding time series data is obtained according to the task configuration information, and the time series data is cleaned and sampled to obtain a training data set;
[0086] Step S806, using the training data set to pre-detect the target data points;
[0087] Step S808: If the pre-detection result is abnormal, the abnormality detection model is trained using the training data set and the model configuration information in the task configuration information, and the trained abnormality detection model is used to re-detect the data point to be detected;
[0088] Step S810, when the detection result is abnormal, continuously collect a preset number of data points to be detected from data points near the target data point to be detected, and determine that the target data point to be detected is abnormal when the detection results of the preset number of data points to be detected are all abnormal.
[0089] By acquiring detection task configuration information and establishing a detection task based on the detection task configuration information, wherein the task configuration information includes model configuration information; when executing the detection task, acquiring time series data corresponding to the detection task, and determining a training data set based on the time series data; performing pre-detection on the target data point to be detected based on the training data set; when the pre-detection result is abnormal, training the anomaly detection model based on the model configuration information and the training data set to obtain the trained anomaly detection model; and detecting the target data point to be detected based on the trained anomaly detection model, by detecting the target data point to be detected based on the training data set and the anomaly detection model, the purpose of determining whether the target data point to be detected is abnormal without setting a threshold is achieved, thereby achieving the technical effect of accurately determining whether various types of data are abnormal in scenarios with complex data types, thereby solving the technical problem that the related technology adopts a fixed threshold judgment method to determine whether the data to be detected is abnormal, resulting in the inability to accurately determine whether the data to be detected is abnormal when there are too many data types.
[0090] In addition, in the time series data anomaly detection method provided in the embodiment of the present application, by cleaning and sampling the time series data and using statistical methods for pre-detection, the detection task of the model is greatly reduced, and at the same time, the unsupervised machine learning model is trained online to perform quasi-real-time anomaly detection on the time series data. In addition, after the detection is completed, it is possible to further combine the abnormal state of the data attached to the point to be detected to determine whether to send an alarm, so as to prevent false alarms caused by data jitter. It can be seen that when the time series data anomaly detection method provided in the embodiment of the present application is used to determine whether the data point to be detected is abnormal, the entire anomaly detection process does not need to set an abnormal judgment threshold and label model training data, and the detection process can be fully automated and closed-loop. Moreover, the time series data anomaly detection method provided in the embodiment of the present application is universal and can be applied to any system that needs to perform anomaly detection on time series data.
[0091] The present application provides a time series data anomaly detection device, and FIG9 is a schematic diagram of the structure of the device. As shown in FIG9, the device includes: a first processing module 90 for obtaining detection task configuration information and establishing a detection task based on the detection task configuration information, wherein the task configuration information includes model configuration information; a second processing module 92 for obtaining time series data corresponding to the detection task when executing the detection task, and determining a training data set based on the time series data; a third processing module 94 for performing pre-detection of target data points to be detected based on the training data set; a fourth processing module 96 for training an anomaly detection model based on the model configuration information and the training data set when the pre-detection result is abnormal, to obtain the trained anomaly detection model; and detecting the target data points to be detected based on the trained anomaly detection model.
[0092] In some embodiments of the present application, the first processing module 90 is also used to obtain the detection task configuration information whose status is effective; construct the detection task based on the detection task configuration information whose status is effective, and set the model configuration information to the task context of the detection task; deliver the detection task to the thread pool, and call the detection task from the thread pool according to the scheduling period corresponding to the detection task.
[0093] In some embodiments of the present application, the first processing module 90 is further configured to read detection task configuration information in the effective state from a database table, and determine the amount of detection task configuration information in the effective state.
[0094] In some embodiments of the present application, the first processing module 90 is further configured to construct a detection task according to the detection task configuration information when the number of detection task configuration information in the effective state is not zero.
[0095] In some embodiments of the present application, the second processing module 92 is further used to determine a data selection time interval; obtain time series data corresponding to the data to be detected based on the data selection time interval; and sample the time series data to obtain a training data set.
[0096] In some embodiments of the present application, the second processing module 92 is also used to sort the time series data; delete a preset proportion of data in the sorted time series data based on the sorting result to obtain an initial sampling result; and perform random sampling on the initial sampling result a preset number of times to obtain a training data set.
[0097] In some embodiments of the present application, the third processing module 94 is also used to determine the mean and variance of the training data set; determine the first value interval and the second value interval based on the mean and variance of the training data set; perform pre-detection on the target data point to be detected based on the first value interval and the second value interval, wherein, when the target data point to be detected is located in the first value interval, the detection result of the pre-detection is determined to be normal, and when the target data point to be detected is located in the second value interval, the detection result is determined to be abnormal.
[0098] In some embodiments of the present application, the fourth processing module 96 is also used to: construct a request body based on the task configuration information, wherein the request body includes the task identifier of the detection task, the target data point to be detected, the model configuration information and the training data set; send the request body to the anomaly detection model, and train the anomaly detection model based on the model configuration information and the training data set to obtain a trained anomaly detection model.
[0099] In some embodiments of the present application, the fourth processing module 96 is also used to detect the target data points to be detected through the trained anomaly detection model, and obtain the response body returned by the anomaly detection model, wherein the response body includes the task identifier of the detection task, the target data points to be detected, and the detection results.
[0100] In some embodiments of the present application, after the target data points to be detected are detected according to the trained anomaly detection model, the fourth processing module 96 is also used to: when the detection result is abnormal, obtain a set of data points to be detected, wherein the data points to be detected in the set of data points to be detected are a preset number of data points collected continuously, and the latest data point to be detected corresponding to the time point in the set of data points to be detected is adjacent to the target data point to be detected, and the time point corresponding to any data point to be detected in the set of data points to be detected is earlier than the time point to be detected; determine the detection results of the data points to be detected in the set of data points to be detected; and when the detection results of the data points to be detected in the set of data points to be detected are all abnormal, generate and send an alarm message according to a preset alarm template.
[0101] It should be noted that the various modules in the above-mentioned time series data anomaly detection device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.
[0102] An embodiment of the present application provides a non-volatile storage medium, in which a program is stored. When the program is running, the device where the non-volatile storage medium is located is controlled to perform the following time series data anomaly detection method: obtaining detection task configuration information, and establishing a detection task based on the detection task configuration information, wherein the task configuration information includes model configuration information; when executing the detection task, obtaining time series data corresponding to the detection task, and determining a training data set based on the time series data; performing pre-detection on the target data point to be detected based on the training data set; when the pre-detection result is abnormal, training the anomaly detection model based on the model configuration information and the training data set to obtain the trained anomaly detection model; and detecting the target data point to be detected based on the trained anomaly detection model.
[0103] An embodiment of the present application provides an electronic device, including a memory and a processor, the processor being configured to run a program stored in the memory, wherein the program executes the following time series data anomaly detection method when running: obtaining detection task configuration information, and establishing a detection task based on the detection task configuration information, wherein the task configuration information includes model configuration information; when executing the detection task, obtaining time series data corresponding to the detection task, and determining a training data set based on the time series data; performing pre-detection on target data points to be detected based on the training data set; when the pre-detection result is abnormal, training an anomaly detection model based on the model configuration information and the training data set to obtain the trained anomaly detection model; and detecting the target data points to be detected based on the trained anomaly detection model.
[0104] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0105] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0106] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0107] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0108] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0109] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0110] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A method for detecting anomalies in time series data, comprising: Obtaining detection task configuration information and establishing a detection task based on the detection task configuration information, wherein the task configuration information includes model configuration information; When executing the detection task, obtaining time series data corresponding to the detection task and determining a training data set based on the time series data; Pre-detecting a target data point to be detected based on the training data set; In the case where the pre-detection result is abnormal, training an anomaly detection model based on the model configuration information and the training data set to obtain the trained anomaly detection model; and Detecting the target data point to be detected based on the trained anomaly detection model.
2. The abnormal detection method for time series data according to claim 1, wherein, The target data point to be detected is a data point determined based on the time point to be detected.
3. The abnormal detection method for time series data according to claim 1, wherein, The obtaining of the detection task configuration information and the establishment of the detection task based on the detection task configuration information include: Obtaining detection task configuration information with an effective status; Constructing the detection task based on the detection task configuration information with the effective status and setting the model configuration information into the task context of the detection task; and Submitting the detection task to a thread pool and retrieving the detection task from the thread pool according to the scheduling period corresponding to the detection task.
4. The method for detecting abnormal time-series data according to claim 3, wherein, The obtaining of the detection task configuration information with an effective status includes: Reading the detection task configuration information with the effective status from a database table and determining the number of the detection task configuration information with the effective status.
5. The method for detecting abnormal time-series data according to claim 4, wherein, The constructing of the detection task based on the detection task configuration information with the effective status includes: In the case where the number of the detection task configuration information with the effective status is not zero, constructing the detection task according to the detection task configuration information.
6. The abnormal detection method for time series data according to claim 1, wherein, The obtaining of the time series data corresponding to the detection task and the determination of the training data set based on the time series data include: Determining a data selection time interval; Obtaining the time series data corresponding to the data to be detected based on the data selection time interval; and Sampling the time series data to obtain the training data set.
7. The abnormal detection method for timing data according to claim 6, wherein, The sampling of the time series data to obtain the training data set includes: Sorting the time series data; Deleting a preset proportion of data from the sorted time series data according to the sorting result to obtain a primary sampling result; and Performing random sampling on the primary sampling result a preset number of times to obtain the training data set.
8. The abnormal detection method for time series data according to claim 1, wherein, The pre-detecting of the target data point to be detected based on the training data set includes: Determining the mean and variance of the training data set; Determining a first value interval and a second value interval based on the mean and variance of the training data set; and Pre-detecting the target data point to be detected based on the first value interval and the second value interval, wherein in the case where the target data point to be detected is within the first value interval, determining that the detection result of the pre-detection is normal, and in the case where the target data point to be detected is within the second value interval, determining that the detection result is abnormal.
9. The abnormal detection method for time series data according to claim 1, wherein, Training the anomaly detection model based on the model configuration information and the training data set to obtain the trained anomaly detection model includes: Constructing a request body according to the task configuration information, where the request body includes the task identifier of the detection task, the target data point to be detected, the model configuration information, and the training data set; and Sending the request body to the anomaly detection model, and training the anomaly detection model based on the model configuration information and the training data set to obtain the trained anomaly detection model.
10. The abnormal detection method for time series data according to claim 9, wherein, Detecting the target data point to be detected based on the trained anomaly detection model includes: Detecting the target data point to be detected through the trained anomaly detection model, and obtaining a response body returned by the anomaly detection model, where the response body includes the task identifier of the detection task, the target data point to be detected, and the detection result.
11. The abnormal detection method for time series data according to claim 1, wherein, After detecting the target data point to be detected based on the trained anomaly detection model, the method further includes: In the case where the detection result is abnormal, obtaining a set of data points to be detected, where the data points to be detected in the set of data points to be detected are a preset number of continuously collected data points, and the corresponding The data point to be detected with the latest time point in the set of data points to be detected is adjacent to the target data point to be detected, and the time point corresponding to any data point to be detected in the set of data points to be detected is earlier than the detection time point; Determining the detection results of the data points to be detected in the set of data points to be detected; and In the case where the detection results of the data points to be detected in the set of data points to be detected are all abnormal, generating and sending an alarm message according to a preset alarm template.
12. The abnormal detection method for time series data according to claim 1, wherein, The detection task configuration information includes whether it is a scheduled task, a task title, metric information, a data source, and a data retrieval frequency.
13. A time-series data anomaly detection device, including: A first processing module, configured to obtain detection task configuration information and establish a detection task according to the detection task configuration information, where the task configuration information includes model configuration information; A second processing module, configured to obtain time-series data corresponding to the detection task when executing the detection task, and determine a training data set according to the time-series data; A third processing module, configured to pre-detect a target data point to be detected according to the training data set; and A fourth processing module, configured to, in the case where the pre-detection result is abnormal, train an anomaly detection model according to the model configuration information and the training data set to obtain the trained anomaly detection model; and detect the target data point to be detected according to the trained anomaly detection model.
14. A non-volatile storage medium, in which a program is stored, wherein, Controlling the device where the non-volatile storage medium is located to execute the time-series data anomaly detection method according to any one of claims 1 to 12 during the running of the program.
15. An electronic device, comprising: A memory and a processor, the processor being configured to run a program stored in the memory, wherein, when the program runs, it executes the timing data anomaly detection method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Anomaly detection method, device and equipment and storage medium
CN112084056A
Inspection method of cloud platform, electronic equipment and nonvolatile storage medium
CN113313280A
Abnormal data processing method and device, computer equipment and storage medium
CN114564629A
Abnormal network flow detection method based on machine learning model optimization
CN116032526A
Time series data anomaly detection method and device and nonvolatile storage medium
CN117786575A
Cited By
Object detection method and device, computer equipment and readable storage medium
CN121434194A
Early warning method and device of meter and electronic equipment
CN122090603A