Monitoring data anomaly detection method and device and readable storage medium
By using an anomaly detection API based on a time series prediction model in monitoring data anomaly detection, the problem of relying on expert experience in the prior art is solved, and more flexible and accurate anomaly detection is achieved, improving the system's adaptability and detection reliability.
Patent Information
- Application Number
- CN202510173886.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-10
AI Technical Summary
The existing threshold-based monitoring data abnormality detection methods rely too much on expert experience, resulting in unreasonable settings and are prone to missed reports or frequent false alarms over time and business changes.
The anomaly detection API built on the time series prediction model is used to analyze the monitoring data, and determine whether the data is abnormal by comparing the difference between the predicted value and the real value, and avoid manually setting the threshold.
It effectively avoids excessive dependence on expert experience, improves the system's ability to adapt to time and business changes, significantly reduces the occurrence of missed and false alarms, and improves the reliability and effectiveness of monitoring data abnormal detection.
Smart Images

Figure CN120123129A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of monitoring technologies, and particularly to an abnormal detection method, device, and readable storage medium for monitoring data. Background Art
[0002] Currently, the monitoring systems of various services achieve real-time monitoring and management of the entire service system by collecting, analyzing, and processing system operation data. The monitoring systems generally adopt a threshold detection method to judge data anomalies, that is, corresponding thresholds and triggering rules are configured for monitoring indicators. When the collected indicator data is judged through the rules and the set thresholds, and does not conform to the rules, an alarm is generated to achieve comprehensive monitoring of various indicators of the system.
[0003] However, there are some problems with the existing threshold-based monitoring data anomaly detection method. Among them, the collection frequency, indicator thresholds, and alarm rules all depend on expert experience for manual setting. Therefore, the human factor is relatively high, and over time and with the development and change of the business, it may result in unreasonable data collection frequency, threshold setting, or alarm rules, or the settings may not be suitable for some applications. For example, if the collection frequency is too low and the threshold setting is too large, it may cause missed reports. If the threshold setting is too small, it may cause frequent false alarms. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an abnormal detection method, device, and readable storage medium for monitoring data in view of the above deficiencies of the prior art, so as to solve the problems that the existing threshold-based monitoring data anomaly detection method is too dependent on expert experience, resulting in unreasonable settings, and these settings are prone to cause missed reports or frequent false alarms over time and with business changes.
[0005] In a first aspect, the present invention provides an abnormal detection method for monitoring data, the
[0006] method includes:
[0007] Obtain a query requirement initiated by a user;
[0008] Perform intent recognition on the query requirement;
[0009] In response to the result of the intent recognition being abnormal detection, call an abnormal detection application programming interface (API) to analyze the monitoring data within the query range, and judge whether the monitoring data is abnormal by comparing the difference between the predicted value and the true value of the monitoring data. Among them, the abnormal detection API is constructed based on a time series prediction model with the ability to recognize the characteristics of monitoring data, and the time series prediction model is used to predict the predicted value of the monitoring data.
[0010] Further, the time series prediction model is a pre-trained model or a fine-tuned model. Before obtaining the query requirements initiated by the user, the method further includes:
[0011] Testing the pre-trained model with zero-shot prediction ability for the monitoring metric prediction task to verify whether the pre-trained model meets the requirements of the monitoring metric prediction task;
[0012] In response to the pre-trained model not meeting the requirements of the monitoring metric prediction task, fine-tuning the pre-trained model using the historical monitoring metric time series data corresponding to the monitoring metric prediction task to obtain the fine-tuned model;
[0013] In response to the pre-trained model meeting the requirements of the monitoring metric prediction task, using the pre-trained model as the time series prediction model;
[0014] Constructing the anomaly detection API based on the time series prediction model.
[0015] Further, the step of fine-tuning the pre-trained model using the historical monitoring metric time series data corresponding to the monitoring metric prediction task to obtain the fine-tuned model specifically includes:
[0016] Using the historical monitoring metric time series data corresponding to the monitoring metric prediction task and combining with the Prompt Tuning (P-Tuning) method to fine-tune the pre-trained model to obtain the fine-tuned model.
[0017] Further, the step of constructing the anomaly detection API based on the time series prediction model specifically includes:
[0018] Constructing a model prediction API based on the time series prediction model;
[0019] Packaging the anomaly detection API according to the model prediction API.
[0020] Further, the step of calling the anomaly detection application programming interface (API) to analyze the monitoring data within the query range and determining whether the monitoring data is abnormal by comparing the difference between the predicted value and the true value of the monitoring data specifically includes:
[0021] Invoke the anomaly detection API and perform the following steps through the anomaly detection API: Obtain the historical monitoring metric time series data corresponding to the anomaly detection task within the query range from at least one data source and perform preprocessing. Use the model prediction API to predict the preprocessed historical monitoring metric time series data to obtain the predicted values of the monitoring metrics within a preset future time period. Compare the predicted values of the monitoring metrics with the true values at the corresponding time points. If the difference is not within the preset range, it is determined that the monitoring metric data at the corresponding time point is abnormal.
[0022] Further, the pre-trained model is the Moirai time series model.
[0023] Further, the method further includes:
[0024] In response to the result of the intent recognition being the prediction of monitoring metric data, invoke the model prediction API to predict the historical monitoring metric time series data of the monitoring metric prediction task within the query range, and obtain the predicted values of the monitoring metrics within a preset future time period.
[0025] In a second aspect, the present invention provides an anomaly detection device for monitoring data, and the device includes:
[0026] A requirement acquisition module, configured to acquire the query requirements initiated by the user;
[0027] An intent recognition module, connected to the requirement acquisition module, and configured to perform intent recognition on the query requirements;
[0028] An anomaly detection module, connected to the intent recognition module, and configured to, in response to the result of the intent recognition being anomaly detection, invoke the anomaly detection application programming interface API to analyze the monitoring data within the query range, and determine whether the monitoring data is abnormal by comparing the difference between the predicted value and the true value of the monitoring data. Among them, the anomaly detection API is constructed based on a time series prediction model with the ability to identify the characteristics of the monitoring data, and the time series prediction model is used to predict the predicted value of the monitoring data.
[0029] In a third aspect, the present invention provides an anomaly detection device for monitoring data, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to implement the anomaly detection method for monitoring data described in the first aspect above.
[0030] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the anomaly detection method for monitoring data described in the first aspect above is implemented.
[0031] Anomaly detection method, device and readable storage medium for monitoring data provided by the present invention. First, obtain the query requirements initiated by the user; then perform intent recognition on the query requirements; and then, in response to the result of the intent recognition being anomaly detection, call the anomaly detection application programming interface API to analyze the monitoring data within the query range, and determine whether the monitoring data is abnormal by comparing the difference between the predicted value and the true value of the monitoring data. Among them, the anomaly detection API is constructed based on a time series prediction model with the ability to recognize the characteristics of monitoring data, and the time series prediction model is used to predict the predicted value of the monitoring data. By introducing the anomaly detection API constructed based on the time series prediction model, the present invention effectively avoids the excessive dependence on expert experience in traditional monitoring data anomaly detection methods and eliminates the irrationality of manually setting thresholds. This method uses the anomaly detection API to automatically analyze the monitoring data within the query range, enabling the system to autonomously evaluate the anomaly interval. This automated processing method significantly improves the system's adaptability to time and business changes, effectively reduces the occurrence of missed reports and false alarms, and thus significantly improves the reliability and effectiveness of monitoring data anomaly detection. It solves the problems that the existing threshold-based monitoring data anomaly detection methods rely too much on expert experience, resulting in unreasonable settings, and these settings are prone to cause missed reports or frequent false alarms as time and business change. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flowchart of an anomaly detection method for monitoring data according to Embodiment 1 of the present invention;
[0033] Figure 2 It is a schematic diagram of the implementation logic of the anomaly detection system for monitoring data according to an embodiment of the present invention;
[0034] Figure 3 It is an overall architecture diagram of the anomaly detection system for monitoring data according to an embodiment of the present invention;
[0035] Figure 4 It is a flowchart of the model construction module according to an embodiment of the present invention;
[0036] Figure 5 It is a flowchart of the anomaly detection application module according to an embodiment of the present invention;
[0037] Figure 6 It is a schematic diagram of the result of historical data anomaly detection according to an embodiment of the present invention;
[0038] Figure 7 It is a schematic diagram of the result of future data anomaly detection according to an embodiment of the present invention;
[0039] Figure 8 It is a schematic diagram of the structure of an anomaly detection device for monitoring data according to Embodiment 2 of the present invention;
[0040] Figure 9 This is a schematic structural diagram of an abnormal detection device for monitoring data according to Embodiment 3 of the present invention. Specific embodiments
[0041] To enable those skilled in the art to better understand the technical solutions of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0042] It can be understood that the specific embodiments and drawings described herein are only for explaining the present invention, rather than limiting the present invention.
[0043] It can be understood that, without conflict, the various embodiments in the present invention and the features in the embodiments can be combined with each other.
[0044] It can be understood that, for the convenience of description, only the parts related to the present invention are shown in the drawings of the present invention, and the parts unrelated to the present invention are not shown in the drawings.
[0045] It can be understood that each unit and module involved in the embodiments of the present invention may correspond to only one physical structure, or may be composed of multiple physical structures, or multiple units and modules may also be integrated into one physical structure.
[0046] It can be understood that the terms "first", "second", etc. in the embodiments of the present invention are used to distinguish different objects, or to distinguish different processes for the same object, rather than to describe a specific order of the objects.
[0047] It can be understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present invention may occur in an order different from that marked in the drawings.
[0048] It can be understood that in the flowcharts and block diagrams of the present invention, the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the embodiments of the present invention are shown. Among them, each block in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified function. Moreover, each block or combination of blocks in the block diagram and flowchart can be implemented by a hardware-based system for implementing the specified function, or by a combination of hardware and computer instructions.
[0049] It can be understood that the units and modules involved in the embodiments of the present invention can be implemented in software or in hardware. For example, the units and modules can be located in the processor.
[0050] Embodiment 1:
[0051] This embodiment provides an abnormal detection method for monitoring data, asFigure 1 As shown in Figure 1 , the method includes:
[0052] Step S101: Obtain the query requirements initiated by the user.
[0053] In this embodiment, the query requirements include two types of requirements: monitoring metric data prediction and anomaly detection.
[0054] Step S102: Perform intent recognition on the query requirements.
[0055] In this embodiment, the results of the intent recognition include anomaly detection and monitoring metric data prediction.
[0056] Step S103: In response to the result of the intent recognition being anomaly detection, call the anomaly detection API (Application Programming Interface) to analyze the monitoring data within the query range, and determine whether the monitoring data is abnormal by comparing the difference between the predicted value and the actual value of the monitoring data. Among them, the anomaly detection API is constructed based on a time series prediction model with the ability to identify the characteristics of monitoring data, and the time series prediction model is used to predict the predicted value of the monitoring data.
[0057] In this embodiment, the anomaly detection API has the ability to provide metric prediction and anomaly detection for the enterprise internal monitoring system. Compared with the existing threshold-based monitoring data anomaly detection method, the method of the present invention is implemented through a model and does not require manual threshold setting. It only needs to accumulate historical data, and the system automatically evaluates the anomaly range. This method can adapt to changes in time and business, avoid missed reports or frequent false alarms caused by unreasonable settings, and is more flexible and accurate, thereby improving the reliability and effectiveness of monitoring data anomaly detection.
[0058] Optionally, the time series prediction model is a pre-trained model or a fine-tuned model. Before obtaining the query requirements initiated by the user, the method further includes:
[0059] Test the pre-trained model with zero-shot prediction ability for the monitoring metric prediction task to verify whether the pre-trained model meets the requirements of the monitoring metric prediction task;
[0060] In response to the pre-trained model not meeting the requirements of the monitoring metric prediction task, fine-tune the pre-trained model with the historical monitoring metric time series data corresponding to the monitoring metric prediction task to obtain the fine-tuned model;
[0061] In response to the pre-trained model meeting the requirements of the monitoring metric prediction task, use the pre-trained model as the time series prediction model;
[0062] Construct the anomaly detection API based on the time series prediction model.
[0063] In this embodiment, the pre-trained model is a model with zero-shot prediction ability to ensure effective prediction when the model faces new data. When selecting a model, it can be tested for downstream tasks (such as monitoring metric prediction tasks) to verify the performance of the model for downstream tasks. If the pre-trained model cannot meet the requirements of the downstream task, fine-tuning can be performed on the historical dataset (i.e., historical monitoring metric time series data) for the downstream task. The fine-tuning process aims to enable the model to have the ability to handle downstream tasks.
[0064] Among them, the pre-trained model / fine-tuned model supports CPU (central processing unit) and GPU (graphics processing unit) computing power and has the ability to reside in memory or video memory. This design enables the model to run efficiently in different computing environments while ensuring the response speed and processing ability of the model.
[0065] Optionally, using the historical monitoring metric time series data corresponding to the monitoring metric prediction task to fine-tune the pre-trained model to obtain the fine-tuned model specifically includes:
[0066] Using the historical monitoring metric time series data corresponding to the monitoring metric prediction task, combined with the P-Tuning method, to fine-tune the pre-trained model to obtain the fine-tuned model.
[0067] In this embodiment, develop a corresponding fine-tuning script for the pre-trained model and use the P-Tuning method for fine-tuning. The fine-tuning parameters include sequence length, prediction length, model layer, learning rate, loss function, etc. Multiple fine-tuning and evaluations can be performed according to different fine-tuning parameters.
[0068] Optionally, constructing the anomaly detection API based on the time series prediction model specifically includes:
[0069] Construct a model prediction API based on the time series prediction model;
[0070] Package the anomaly detection API according to the model prediction API.
[0071] In this embodiment, for the pre-trained model / fine-tuned model, a model prediction API is developed for downstream tasks to call. The input of the interface is a time series data of monitoring metrics for a downstream task, which covers various business-related metrics in different time dimensions, including but not limited to business metrics, performance metrics, perception metrics, component metrics, and operation metrics, etc. The output is the predicted value of the monitoring metric data within a preset future time period. For example, the metric data for the next 30 days is predicted based on the daily data of the past 180 days.
[0072] In this embodiment, an anomaly detection API is encapsulated according to the model prediction API. For example, if it is necessary to monitor the anomaly metric data of the past 30 days, the metric data of the 180 days before the past 30 days is input.
[0073] Optionally, the anomaly detection application programming interface API is called to analyze the monitoring data within the query range, and it is determined whether the monitoring data is abnormal by comparing the difference between the predicted value and the real value of the monitoring data. Specifically, it includes:
[0074] Call the anomaly detection API, and perform the following steps through the anomaly detection API: obtain the historical monitoring metric time series data corresponding to the anomaly detection task within the query range from at least one data source and perform preprocessing, use the model prediction API to predict the preprocessed historical monitoring metric time series data to obtain the predicted value of the monitoring metric data within a preset future time period, compare the predicted value of the monitoring metric data with the real value at the corresponding time point, and if the difference is not within the preset range, it is determined that the monitoring metric data at the corresponding time point is abnormal.
[0075] In this embodiment, the data sources include data sources such as monitoring platforms and big data platforms. The steps for preprocessing the historical monitoring metric time series data include removing the noise signal in the data, and processing missing values, removing outliers, noise processing, data dimension processing, etc. according to the characteristics of the data and business requirements.
[0076] In this embodiment, taking the future data anomaly detection as an example, the monitoring metric data within a preset future time period can be predicted by means of a background scheduled task. When the real value occurs, anomaly detection judgment is performed. If the difference between the real value and the predicted value is not within a reasonable range, it is determined as abnormal, and an alarm is supported.
[0077] Optionally, the pre-trained model is the Moirai time series model.
[0078] In this embodiment, since the Moirai time series model has zero-shot prediction ability, can effectively handle the high heterogeneity of time series data, and provide a flexible prediction distribution, the pre-trained model is preferably the Moirai time series model. Based on the Moirai model, a model suitable for downstream tasks can be fine-tuned and the corresponding model prediction and anomaly detection APIs can be developed.
[0079] Optionally, the method further includes:
[0080] In response to the result of the intent recognition being the prediction of monitoring metric data, the model prediction API is called to predict the historical monitoring metric time series data of the monitoring metric prediction task within the query range, and the predicted value of the monitoring metric data within a preset future time period is obtained.
[0081] In this embodiment, assuming that the query range is from January to December this year, and the preset future time period is January next year, the model prediction API will analyze and learn using the historical monitoring metric time series data from January to December this year, and then generate the predicted value of the monitoring metric data for January next year.
[0082] It should be noted that the main objective of the anomaly detection method for monitoring data provided by the present invention is to construct a time series model suitable for enterprise business processes and fitting enterprise data characteristics based on a pre-trained time series model (i.e., a time series prediction model). On this basis, time series prediction is performed on the monitoring system metric data to discover anomalies in the monitoring system metric data. By using the method of the present invention to fine-tune the pre-trained time series model on the large-scale monitoring system historical data, the model is enabled to have the ability to recognize the data characteristics of the business system. The fine-tuned model is suitable for monitoring metric prediction and anomaly detection. We use the metric data as the model input to predict the data, compare the predicted value with the collected real value, and thus discover the difference between the predicted value and the real value. If the difference is not within a reasonable range, it is regarded as an anomaly.
[0083] It should be noted that time series prediction refers to predicting the future based on historical data. Specifically, time series prediction analyzes the trends, seasonality, and periodicity of time series data and establishes mathematical models. Through the fitting and prediction of these models, the trend changes and laws in the time series are described, and then the future changes are predicted. With the emergence of the LLM (Large Language Model), time series prediction models based on the transformer architecture have emerged. Time series prediction models trained on large-scale pre-trained datasets have good zero-shot prediction ability, and this ability brings the possibility of threshold-free anomaly detection in the monitoring field.
[0084] In a specific embodiment, the abnormal detection method for monitoring data is applied to an abnormal detection system for monitoring data, and the overall implementation logic is as Figure 2 shown, including "data -> model -> application". The following is a detailed explanation of its specific meaning:
[0085] 1) Data: It is time-series data for model input, that is, metric data collected from data source systems such as monitoring platforms and big data platforms, and after cleaning and processing, it is deposited as input data for model prediction and a dataset for fine-tuning.
[0086] 2) Model: It is a pre-trained or fine-tuned model for processing time-series data. Usually, a pre-trained model can be directly used for zero-shot prediction of the model input data, and a fine-tuned model can be used for prediction tasks for specific tasks.
[0087] 3) Application: It is to perform abnormal detection tasks for specific demand scenarios by calling the model prediction API and abnormal detection algorithms, such as abnormal detection of historical data and prediction of future data.
[0088] The overall architecture diagram of the abnormal detection system for monitoring data is as Figure 3 shown, and its implementation mainly consists of the following five parts: including a data collection module, a data processing and cleaning module, a model construction module, a model prediction API and an abnormal detection API module, and an abnormal detection application module.
[0089] 1. Data collection module
[0090] This data collection module includes a data source, a data collection unit, and a data storage unit.
[0091] (1) Data source: It includes data sources such as monitoring platforms and big data platforms, and publishes data query interfaces on the enterprise internal capability open platform so that other systems can conveniently obtain the required data;
[0092] (2) Data collection unit: Sets a timing task (such as executing a collection task every 5 minutes), collects data through API calls, and converts it into a format that can be processed within the system. Sensitive data is encrypted during the collection and transmission process, and means such as a retry mechanism, data backup, and recovery strategy are used to fully ensure the reliability, security, and accuracy of the data;
[0093] Among them, the collected data includes business metrics, performance metrics, perception metrics, component metrics, operation metrics, etc.
[0094] (3) Data storage unit: Stores the received data in a database for subsequent data processing and analysis. This unit uses a relational database such as mysql as the storage medium.
[0095] 2. Data Processing and Cleaning Module
[0096] This module is used to process and clean the original data. Here, the original data refers to the data collected from the data source and converted into an internally processable format. The steps of processing and cleaning include:
[0097] (1) According to the needs of business monitoring, remove the noise signals in the data, and process missing values, remove outliers, perform noise processing, data dimension processing, etc. according to the characteristics of the data and business requirements, and process the original collected data into data that can be used to build and train a fine-tuning model;
[0098] (a) Handling missing values: Missing values are a common problem in time series data. The handling methods may include interpolation (for example, filling in missing values with the average of the previous and subsequent observed values), default value replacement, etc.
[0099] (b) Removing outliers: Removing outliers means directly removing the data points identified as outliers from the data set. This strategy is usually applicable when the data set is large enough that deleting a few extreme values will not have a significant impact on the overall statistical characteristics. The pruning process can be divided into the following steps:
[0100] Defining a threshold: First, a threshold needs to be determined to define what constitutes an outlier. This can be achieved through statistical methods, such as using Z-score (standard score), IQR (interquartile range), or other statistical metrics.
[0101] Identifying outliers: Use the selected threshold to identify which data points are considered outliers.
[0102] Removing outliers: Once the outliers are identified, they can be removed from the data set. This usually involves modifying the data set so that it no longer contains these values.
[0103] (c) Noise processing: It means replacing the outliers with values within a certain threshold instead of completely deleting them. This method is applicable when the data points cannot be simply discarded because the data set may be small, or the number of outliers is large, and deleting them will significantly change the data distribution. The steps of noise processing include:
[0104] Defining the upper and lower limits: Similarly, a reasonable upper and lower limit needs to be determined, and all values above the upper limit or below the lower limit will be replaced.
[0105] Replacing outliers: Replace the identified outliers with the upper or lower limit values. For example, if the upper limit is 100, then all values greater than 100 will be set to 100; similarly, if the lower limit is 0, then all values less than 0 will be set to 0.
[0106] (2) According to the needs of business monitoring, the original collected data is processed into a multi-dimensional dataset. For example, the dataset is output according to the time dimension for data prediction in specific scenarios.
[0107] 3. Model Construction Module
[0108] The flowchart of this model construction module is as Figure 4 shown, where:
[0109] (1) Pre-trained model: The pre-trained model adopts a model with zero-shot prediction ability to ensure that the model can make effective predictions when facing new data. When selecting a model, it should be tested for downstream tasks to verify the performance of the model for downstream tasks. This system preferably adopts the Moirai time series model.
[0110] (2) Fine-tuning model: If the pre-trained model cannot meet the needs of downstream tasks, it can be fine-tuned on the historical dataset for downstream tasks. The fine-tuning process aims to enable the model to have the ability to handle downstream tasks.
[0111] In addition, the model supports CPU and GPU computing power and has the ability to reside in memory or video memory. This design enables the model to run efficiently in different computing environments while ensuring the response speed and processing ability of the model.
[0112] It should be noted that the training of time series models is usually carried out on open-source datasets. Currently, there is no operator business dataset (i.e., monitoring dataset). This means that the performance of time series models may not reach the best when making zero-shot predictions on operator data. Therefore, in order to improve the prediction accuracy, it is necessary to construct an operator business dataset and fine-tune on the operator business dataset. Among them, to construct an operator business dataset, it is necessary to collect indicator data for downstream tasks with typical characteristics (such as the call traffic data of billing services showing periodic fluctuations and seasonal factors, and the user development volume data showing a linear growth trend). After data collection, the data needs to be cleaned and processed, and then the model is fine-tuned. Specifically, the time series model fine-tuned on the operator's proprietary dataset has the ability to identify enterprise data characteristics. The overall steps of time series model fine-tuning can include:
[0113] (a) Dataset Preparation
[0114] The dataset for fine-tuning includes a training set, a validation set, and a test set to evaluate the model performance and adjust parameters. The training set is used to fit the model, the validation set is used to adjust hyperparameters and prevent overfitting, and the test set is used to evaluate the final model performance.
[0115] (b) Fine-tuning Script Preparation and Parameter Adjustment
[0116] Develop corresponding fine-tuning scripts for the model. In this embodiment, fine-tuning in the P-Tuning manner is supported. Parameter adjustments include sequence length, prediction length, model hierarchy, learning rate, loss function, etc. Multiple fine-tuning operations should be performed and evaluated for different parameters.
[0117] (c) Model fine-tuning
[0118] Performing model fine-tuning requires a large amount of GPU resources.
[0119] (d) Model evaluation
[0120] After the model fine-tuning is completed, the model should be tested and evaluated using the test set.
[0121] 4. Model prediction API and anomaly detection API modules
[0122] (1) Model prediction API: Develop application programming interfaces for the pre-trained model / fine-tuned model to be called by downstream tasks. The input of the interface is a segment of monitoring metric data for downstream tasks (including business metrics, performance metrics, perception metrics, component metrics, operation metrics, etc.). For example, predict the metric data for the next 30 days based on the daily data of the past 180 days.
[0123] (2) Anomaly detection API: Package the anomaly detection API based on the model prediction API. For example, if it is necessary to monitor the anomaly metric data of the past 30 days, then input the metric data of the 180 days before the past 30 days.
[0124] Among them, the model prediction API and the anomaly detection API have a certain concurrency ability to meet business needs.
[0125] It should be noted that both the pre-trained model and the fine-tuned model are used for prediction tasks. In the implementation process, we first use the pre-trained model for zero-shot prediction. If the prediction result meets the expectation, we will directly adopt this pre-trained model. If the result does not meet the expectation, we will fine-tune the model to improve the prediction accuracy. Among them, anomaly detection is carried out on the basis of prediction, that is, by comparing the predicted value with the true value, and if the deviation is greater than the confidence interval, it is judged as an anomaly.
[0126] It should be noted that the model prediction API and the anomaly detection API developed based on the pre-trained model / fine-tuned model can be deployed on the capability sharing platform for other business systems to call these time series model APIs to perform model prediction or anomaly detection tasks.
[0127] 5. Anomaly detection application module
[0128] The flowchart of this anomaly detection application module is as Figure 5Shown as follows: First, the user initiates a query requirement through natural language (such as anomaly detection), and then the system performs intent recognition: The system runs the anomaly detection API and obtains the system operation metrics (including anomaly prediction metric data), and finally outputs an inspection report through metric data comparison.
[0129] It should be noted that the query requirement includes two types of requirements: monitoring metric data prediction and anomaly detection. Each requirement calls one of the two APIs. Among them, anomaly detection includes the following two types:
[0130] (1) Historical data anomaly detection: Anomaly data query of historical metric data is performed in a user interaction mode or a background scheduled task mode. Custom metrics, query ranges, and output ranges are supported, and charts and data lists can be output. For example, the query range can be set from January to December, and the output range is from November to December, which means the system will detect the anomaly data from November to December in the historical data of these 12 months. Anomaly detection is essentially implemented based on a prediction method, so a period of historical data needs to be provided as input to support the model for accurate prediction and anomaly monitoring. For example, in a specific business scenario, if you want to detect the data anomaly of a certain month, the historical data of the past year needs to be provided as input to the model.
[0131] Figure 6 The schematic diagram of the result of historical data anomaly detection is shown. Two curves are shown in the figure: one represents the predicted value (Prediction), and the other represents the true value (GroundTruth). These two curves are used to compare the prediction result of the model with the actual observed value. The part with the gray background in the figure is used to highlight the detected anomaly area.
[0132] (2) Future data anomaly detection: The future metric data is predicted in a background scheduled task mode, and anomaly detection judgment is performed when the true value occurs. If the difference between the true value and the predicted value is not within a reasonable range, it is determined as an anomaly, and alarming is supported.
[0133] Figure 7 The schematic diagram of the result of future data anomaly detection is shown. Two curves are shown in the figure: one represents the predicted value (Prediction), and the other represents the true value (GroundTruth). The part with the gray background in the figure is used to highlight the detected anomaly area.
[0134] It should be noted that in practical applications, future data anomaly detection has been widely used because it can identify and predict potential anomalies in the future, thus discovering potential problems in advance, which is of great significance for risk management and timely response. In contrast, although historical data anomaly detection helps analyze and understand past data fluctuations and anomalies, since it mainly conducts retrospective analysis, its help for real-time decision-making and forward-looking management is relatively limited, so it is less used in practical applications.
[0135] It should be noted that the monitoring data anomaly detection method provided by the present invention relies on rich data sources such as the enterprise's internal monitoring platform and big data platform. This method efficiently collects data, precisely cleans data, and deeply processes data, and uses time series large model technology to generate time series data in different time dimensions. On this basis, a time series model that can identify the characteristics of enterprise data is fine-tuned, and model prediction APIs and anomaly detection APIs are developed. These APIs have the ability to provide metric prediction and anomaly detection for the enterprise's internal monitoring system. Compared with the existing threshold-based monitoring data anomaly detection method, the method of the present invention is implemented through a model and does not require manual threshold setting. It only needs to accumulate based on historical data, and the system automatically evaluates the anomaly range and issues early warnings. This method can adapt to changes in time and business, avoid missed reports or frequent false alarms caused by unreasonable settings, and is more flexible and accurate, thus improving the reliability and effectiveness of monitoring data anomaly detection.
[0136] The anomaly detection method for monitoring data provided by the embodiments of the present invention first obtains the query requirements initiated by the user; then performs intent recognition on the query requirements; and then, in response to the result of the intent recognition being anomaly detection, calls the anomaly detection application programming interface API to analyze the monitoring data within the query range, and determines whether the monitoring data is abnormal by comparing the difference between the predicted value and the actual value of the monitoring data. Among them, the anomaly detection API is constructed based on a time series prediction model with the ability to identify the characteristics of monitoring data, and the time series prediction model is used to predict the predicted value of the monitoring data. By introducing the anomaly detection API constructed based on the time series prediction model, the present invention effectively avoids the excessive dependence on expert experience in traditional monitoring data anomaly detection methods and eliminates the unreasonableness of manual threshold setting. This method uses the anomaly detection API to automatically analyze the monitoring data within the query range, enabling the system to autonomously evaluate the anomaly range. This automated processing method significantly improves the system's adaptability to changes in time and business, effectively reduces the occurrence of missed reports and false alarms, and thus significantly improves the reliability and effectiveness of monitoring data anomaly detection. It solves the problems that the existing threshold-based monitoring data anomaly detection methods rely too much on expert experience, resulting in unreasonable settings, and these settings are prone to cause missed reports or frequent false alarms as time and business change.
[0137] Embodiment 2:
[0138] As Figure 8 shown, this embodiment provides an abnormal detection device for monitoring data, which is used to execute the above-mentioned abnormal detection method for monitoring data, and includes:
[0139] A requirement acquisition module 11, which is used to acquire the query requirements initiated by the user;
[0140] An intention recognition module 12, which is connected to the requirement acquisition module 11 and is used to recognize the intention of the query requirements;
[0141] An abnormal detection module 13, which is connected to the intention recognition module 12 and is used to, in response to the result of the intention recognition being abnormal detection, call the abnormal detection application programming interface API to analyze the monitoring data within the query range, and determine whether the monitoring data is abnormal by comparing the difference between the predicted value and the true value of the monitoring data. Among them, the abnormal detection API is constructed based on a time series prediction model with the ability to recognize the characteristics of monitoring data, and the time series prediction model is used to predict the predicted value of the monitoring data.
[0142] Optionally, the device further includes:
[0143] A task testing module, which is used to test the monitoring index prediction task of a pre-trained model with zero-shot prediction ability to verify whether the pre-trained model meets the requirements of the monitoring index prediction task;
[0144] A model fine-tuning module, which is used to, in response to the pre-trained model not meeting the requirements of the monitoring index prediction task, fine-tune the pre-trained model with the historical monitoring index time series data corresponding to the monitoring index prediction task to obtain the fine-tuned model;
[0145] A model determination module, which is used to, in response to the pre-trained model meeting the requirements of the monitoring index prediction task, use the pre-trained model as the time series prediction model;
[0146] An abnormal detection interface construction module, which is used to construct the abnormal detection API based on the time series prediction model.
[0147] Optionally, the model fine-tuning module includes:
[0148] A fine-tuning unit, which is used to fine-tune the pre-trained model with the historical monitoring index time series data corresponding to the monitoring index prediction task in combination with the Prompt Tuning (P-Tuning) method to obtain the fine-tuned model.
[0149] Optionally, the abnormal detection interface construction module includes:
[0150] A first interface construction unit, configured to construct a model prediction API based on the time series prediction model;
[0151] A second interface construction unit, configured to encapsulate the anomaly detection API according to the model prediction API.
[0152] Optionally, the anomaly detection module 13 includes:
[0153] An interface call execution unit, configured to call the anomaly detection API and perform the following steps through the anomaly detection API: obtain historical monitoring metric time series data corresponding to the anomaly detection task within the query range from at least one data source and perform preprocessing, use the model prediction API to predict the preprocessed historical monitoring metric time series data to obtain predicted values of monitoring metric data within a preset future time period, compare the predicted values of the monitoring metric data with the true values at corresponding time points, and if the difference is not within the preset range, determine that the monitoring metric data at the corresponding time point is abnormal.
[0154] Optionally, the pre-trained model is a Moirai time series model.
[0155] Optionally, the device further includes:
[0156] A metric prediction module, configured to, in response to the result of the intent recognition being the prediction of monitoring metric data, call the model prediction API to predict the historical monitoring metric time series data of the monitoring metric prediction task within the query range to obtain predicted values of monitoring metric data within a preset future time period.
[0157] Embodiment 3:
[0158] Reference Figure 9 , this embodiment provides an anomaly detection device for monitoring data, including a memory 21 and a processor 22. A computer program is stored in the memory 21, and the processor 22 is configured to run the computer program to execute the anomaly detection method for monitoring data in Embodiment 1.
[0159] Wherein, the memory 21 is connected to the processor 22. The memory 21 can adopt flash memory or read-only memory or other memories, and the processor 22 can adopt a central processing unit or a single-chip microcomputer.
[0160] Embodiment 4:
[0161] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the anomaly detection method for monitoring data in the above Embodiment 1 is implemented.
[0162] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, computer program modules, or other data. The computer-readable storage medium includes, but is not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), digital versatile discs (DVDs) or other optical disc storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0163] In summary, for the anomaly detection method, device, and readable storage medium for monitoring data provided by the embodiments of the present invention, first, a query requirement initiated by a user is obtained; then, the intent of the query requirement is recognized; and then, in response to the result of the intent recognition being anomaly detection, an anomaly detection application programming interface (API) is called to analyze the monitoring data within the query range. By comparing the difference between the predicted value and the actual value of the monitoring data, it is determined whether the monitoring data is abnormal. Among them, the anomaly detection API is constructed based on a time series prediction model with the ability to identify the characteristics of the monitoring data, and the time series prediction model is used to predict the predicted value of the monitoring data. By introducing the anomaly detection API constructed based on the time series prediction model, the present invention effectively avoids the excessive dependence on expert experience in traditional monitoring data anomaly detection methods and eliminates the irrationality of manually setting thresholds. This method automatically analyzes the monitoring data within the query range using the anomaly detection API, enabling the system to autonomously evaluate the anomaly interval. This automated processing method significantly improves the system's adaptability to time and business changes, effectively reduces the occurrence of missed reports and false alarms, and thus significantly enhances the reliability and effectiveness of monitoring data anomaly detection. It solves the problems of the existing threshold-based monitoring data anomaly detection methods being overly dependent on expert experience, resulting in unreasonable settings, and these settings being prone to missed reports or frequent false alarms as time and business change.
[0164] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present invention, and the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered within the protection scope of the present invention.
Claims
1. A method for detecting anomalies in monitoring data, characterized in that: The method comprises: Get the query requirements initiated by the user; Performing intent recognition on the query demand; In response to the result of the intent recognition being an anomaly detection, an anomaly detection application programming interface API is called to analyze the monitoring data within the query range, and whether an anomaly occurs in the monitoring data is determined by comparing the difference between the predicted value and the actual value of the monitoring data, wherein the anomaly detection API is constructed based on a time series prediction model that has the ability to identify the characteristics of monitoring data, and the time series prediction model is used to predict the predicted value of the monitoring data.
2. The method according to claim 1, characterized in that The time series prediction model is a pre-trained model or a fine-tuned model. Before obtaining the query demand initiated by the user, the method further includes: Testing the monitoring indicator prediction task on the pre-trained model with zero-sample prediction capability to verify whether the pre-trained model meets the requirements of the monitoring indicator prediction task; In response to the pre-trained model not meeting the requirements of the monitoring indicator prediction task, fine-tuning the pre-trained model using the historical monitoring indicator time series data corresponding to the monitoring indicator prediction task to obtain the fine-tuned model; In response to the pre-trained model meeting the requirements of the monitoring indicator prediction task, using the pre-trained model as the time series prediction model; The anomaly detection API is constructed based on the time series prediction model.
3. The method according to claim 2, characterized in that The fine-tuning of the pre-trained model by using the historical monitoring indicator time series data corresponding to the monitoring indicator prediction task to obtain the fine-tuning model specifically includes: The historical monitoring indicator time series data corresponding to the monitoring indicator prediction task is used, combined with the prompt tuning P-Tuning method, to fine-tune the pre-trained model to obtain the fine-tuning model.
4. The method according to claim 2, characterized in that: The constructing of the anomaly detection API based on the time series prediction model specifically includes: Building a model prediction API based on the time series prediction model; The anomaly detection API is encapsulated according to the model prediction API.
5. The method according to claim 4, characterized in that The calling of the anomaly detection application programming interface API analyzes the monitoring data within the query range, and determines whether the monitoring data is abnormal by comparing the difference between the predicted value and the actual value of the monitoring data, specifically including: Call the anomaly detection API and perform the following steps through the anomaly detection API: obtain historical monitoring indicator time series data corresponding to the anomaly detection task within the query range from at least one data source and preprocess it, use the model prediction API to predict the preprocessed historical monitoring indicator time series data to obtain the predicted value of the monitoring indicator data within a future preset time period, compare the predicted value of the monitoring indicator data with the actual value at the corresponding time point, and if the difference is not within the preset range, it is determined that the monitoring indicator data at the corresponding time point is abnormal.
6. The method according to claim 2, characterized in that The pre-trained model is the Moirai time series model.
7. The method according to claim 4, characterized in that The method further comprises: In response to the result of the intent recognition being a monitoring indicator data prediction, the model prediction API is called to predict the historical monitoring indicator time series data of the monitoring indicator prediction tasks within the query scope to obtain the predicted value of the monitoring indicator data within a preset time period in the future.
8. A monitoring data anomaly detection device, characterized in that: The device comprises: Demand acquisition module, used to obtain query requirements initiated by users; An intention recognition module, connected to the demand acquisition module, for performing intention recognition on the query demand; An anomaly detection module is connected to the intention recognition module and is used to respond to an anomaly detection result of the intention recognition, call the anomaly detection application programming interface API to analyze the monitoring data within the query range, and judge whether an anomaly occurs in the monitoring data by comparing the difference between the predicted value and the actual value of the monitoring data, wherein the anomaly detection API is constructed based on a time series prediction model capable of identifying the characteristics of monitoring data, and the time series prediction model is used to predict the predicted value of the monitoring data.
9. A monitoring data anomaly detection device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the monitoring data anomaly detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for detecting anomalies in monitoring data according to any one of claims 1 to 7 is implemented.