Abnormality detection method based on time sequence database
Through the dynamic confidence interval detection method based on the Prophet model, the false alarm and missed response problems of the traditional fixed threshold method in business fluctuations and complex scenarios are solved, efficient and flexible abnormality detection is achieved, and the accuracy and stability of the monitoring system are improved.
Patent Information
- Application Number
- CN202510543242.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-19
AI Technical Summary
The existing abnormality detection methods based on fixed thresholds are difficult to adapt to dynamic changes caused by business volume fluctuations, resulting in false positives or missed reports, and are difficult to adapt to complex and changeable business scenarios, reducing the efficiency and reliability of the monitoring system.
The prediction model is trained using the Prophet model to generate dynamic confidence intervals, judge abnormalities by comparing them with the actual value, and combine real-time monitoring and alarm mechanisms to automatically handle abnormal detection.
It improves the accuracy and system adaptability of abnormal detection, reduces false alarm rates, reduces operation and maintenance costs, and enhances the flexibility and reliability of the system.
Abstract
Description
Technical Field
[0001] The present invention relates to the field of IT operation and maintenance, and in particular to an anomaly detection method based on a time series database. Background Art
[0002] In the IT operations and maintenance field, various performance indicator data generated by servers, databases, and middleware are typically stored in time series databases such as Prometheus, OpenTSDB, and InfluxDB. This data is crucial for monitoring system health, analyzing performance bottlenecks, and providing fault warnings.
[0003] Currently, mainstream anomaly detection methods mainly rely on fixed threshold strategies, that is, they periodically check whether the collected data exceeds the preset maximum or minimum thresholds to determine whether to trigger an alarm. However, this method has obvious limitations: it is difficult to adapt to the dynamic changes brought about by business volume fluctuations. For example, during business peak hours and normal hours, the normal operating indicators of the system (such as CPU utilization, response time, etc.) will be significantly different. Fixed thresholds cannot flexibly respond to such changes, resulting in a large number of false alarms during business peak hours, while some potential problems may be missed during non-peak hours. In addition, as the complexity and scale of the business increase, static threshold setting is not only time-consuming and labor-intensive, but also difficult to cover all scenarios, thereby reducing the overall effectiveness and reliability of the monitoring system. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention provides an anomaly detection method based on a time series database to detect outliers in time series data in real time, so as to improve the accuracy and efficiency of the monitoring system, ensure that potential problems can be discovered and handled in a timely manner, and ensure the stable operation of the business.
[0005] The technical solution of the present invention is:
[0006] An anomaly detection method based on a time series database includes the following steps:
[0007] Periodically pull historical data of specified indicators from the time series database and use the historical data to train the prediction model;
[0008] Generate the predicted value at the current time point based on the trained prediction model, including the maximum predicted value (yhat_upper) and the minimum predicted value (yhat_lower), and form a confidence interval;
[0009] Compare the actual collected current value with the confidence interval to determine whether it exceeds the upper and lower limits of the confidence interval, and mark it as abnormal if it exceeds the upper and lower limits;
[0010] Write the outliers, yhat_upper, and yhat_lower back to the time series database as new indicators, and display these indicators as lines of different colors on the user interface, where the part outside the confidence interval is highlighted in red;
[0011] When an anomaly is detected, the alarm mechanism is triggered and the alarm information is sent to the operation and maintenance personnel.
[0012] Further,
[0013] The time interval for periodically pulling historical data from the time series database is every fixed time period (such as every 2 hours), and the time window of the historical data is the data of the most recent few days (such as the most recent 3 days).
[0014] The prediction model is trained using the Prophet model or other models suitable for time series analysis, and each indicator corresponds to an independent prediction model.
[0015] The frequency of obtaining the predicted value is to obtain the predicted value of the current time point from the prediction model every period of time (such as every 2 minutes).
[0016] The standard for abnormal determination is that the current value exceeds yhat_upper or is lower than yhat_lower by a certain ratio (such as ±10%), and the specific ratio can be adjusted according to business needs.
[0017] Further,
[0018] Generate the yhat_upper value as a new indicator, name it as the original indicator name + _prophet + _upper, and write the value back to the time series database;
[0019] Generate the yhat_lower value as a new indicator, name it as the original indicator name + _prophet + _lower, and write the value back to the time series database;
[0020] The abnormal value is recorded as a new indicator named as the original indicator name + _prophet + _anomaly, with a value of 0 for abnormality or 1 for normality, and is written back to the time series database.
[0021] Further,
[0022] In the monitoring panel, the following indicators are displayed in the same chart: original indicator, that is, the original collected value; original indicator name + _prophet + _upper, original indicator name + _prophet + _lower, and original indicator name + _prophet + _anomaly.
[0023] Further,
[0024] Configure the percentage threshold for the current indicator value to exceed the confidence interval through the interface. When the actual value exceeds the threshold, an alarm is triggered and the alarm information is sent to the alarm platform.
[0025] The beneficial effects of the present invention are
[0026] This paper uses a dynamic prediction model to perform real-time anomaly detection on indicator data in a time series database. Compared with the traditional fixed threshold method, it has the following significant technical advantages and practical application value:
[0027] 1. Improve detection accuracy
[0028] Traditional fixed threshold methods are difficult to adapt to fluctuations between business peak periods and normal times, and are prone to false positives or missed positives.
[0029] The present invention uses the Prophet model to generate a dynamic normal range (confidence interval) based on historical data, which can flexibly respond to the periodic fluctuations and trend changes of indicator values and significantly improve the accuracy of anomaly detection.
[0030] 2. Enhance system adaptability
[0031] The dynamic prediction model can automatically learn the time series characteristics of different indicators (such as seasonality, trends, etc.), and is suitable for complex and changing business scenarios, especially in high-concurrency and high-frequency fluctuation environments.
[0032] It supports users to adjust the ratio threshold of abnormal judgment according to actual needs, further enhancing the flexibility and applicability of the system.
[0033] 3. Reduce operation and maintenance costs
[0034] High degree of automation: By regularly training models and detecting anomalies in real time, the need for manual intervention is reduced, reducing the workload of operations and maintenance personnel.
[0035] Reduce false alarm rate: The dynamic threshold mechanism effectively avoids the frequent false alarm problem caused by improper fixed threshold setting in traditional methods, thereby reducing unnecessary alarm processing costs. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are part of the embodiments of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0037] The purpose of the present invention is to provide an intelligent, efficient and flexible real-time anomaly detection method for identifying abnormal values in indicator data stored in a time series database. By dynamically predicting the normal range of each indicator (including the lowest lower limit and the highest upper limit), and comparing and analyzing with the actual collected data, abnormal values that exceed the normal range are discovered in a timely manner. Once an anomaly is detected, the system will mark these abnormal values in a significant manner (such as highlighting them in red on the monitoring interface), and trigger an alarm notification mechanism to push relevant information to the operation and maintenance personnel in real time. This enables the operation and maintenance team to quickly locate the problem and take appropriate treatment measures, thereby effectively improving the stability and reliability of the system and reducing the risk of business interruption or performance degradation due to abnormal events.
[0038] This paper proposes an anomaly detection method based on time series data. It generates the normal range of indicators through a dynamic prediction model and combines real-time monitoring and alarm mechanisms to achieve efficient anomaly detection. The specific steps are as follows:
[0039] Step 1: Train a predictive model using historical data
[0040] 1. Configure detection indicators
[0041] Configure the target indicator to be detected through the interface and replace the indicator label (such as adding host="192.168.12.66") to distinguish the detected object.
[0042] 2. Pull historical data
[0043] Periodically (e.g., every 2 hours) extract fixed window historical data (e.g., data from the last 3 days) of the target indicator from the time series database. This historical data is used to train the prediction model.
[0044] 3. Training the prediction model
[0045] The pulled historical data is fed into the Prophet forecasting model for training. Each metric (after label replacement) corresponds to a separate forecasting model. After training, the model is stored in a Map structure, where the key is the metric name + _prophet, and the value is the corresponding trained model.
[0046] Step 2: Obtain prediction data from the prediction model
[0047] 1. Generate predicted values
[0048] At regular intervals (e.g., 2 minutes), the prediction model stored in the Map is called to obtain the predicted value at the current time point, including the maximum value (yhat_upper) and the minimum value (yhat_lower). These two values define the normal range (i.e., confidence interval) of the indicator at the current time point.
[0049] 2. Abnormal judgment
[0050] Compare the current collected actual value (current_value) with the predicted maximum and minimum values. If current_value exceeds yhat_upper or is lower than yhat_lower by a certain percentage (the threshold can be set according to business needs, such as ±10%), it is determined to be an outlier.
[0051] Step 3: Generate prediction indicators and write them back to the time series database
[0052] 1. Generate predictive indicators
[0053] Generate the yhat_upper value as a new indicator, name it as the original indicator name + _prophet + _upper, and write the value back to the time series database;
[0054] Generate the yhat_lower value as a new indicator, name it as the original indicator name + _prophet + _lower, and write the value back to the time series database;
[0055] For detected outliers, a new indicator is generated, named the original indicator name + _prophet + _anomaly, with a value of 0 (abnormal) or 1 (normal), and written back to the time series database.
[0056] 2. Update the database
[0057] The newly generated indicators will be stored in the time series database as supplementary data for subsequent monitoring and analysis.
[0058] Step 4: Interface display
[0059] 1. Data Visualization
[0060] In monitoring panels such as Grafana, display the following metrics in the same chart: original metrics (original collected values);
[0061] Original indicator name + _prophet + _upper (predicted maximum value), original indicator name + _prophet + _lower (predicted minimum value), original indicator name + _prophet + _anomaly (abnormal marker)
[0062] 2. Confidence Interval Annotation
[0063] In the chart, the area between _prophet+_upper and _prophet+_lower is filled with a light gray background to represent the confidence interval. Any part of the original indicator that falls outside the confidence interval is highlighted in red, visually marking the abnormal part.
[0064] Step 5: Generate abnormal alarm
[0065] 1. Configure alarm rules
[0066] Configure the percentage threshold (such as ±15%) of the current indicator value exceeding the confidence interval through the interface. If the actual value exceeds the threshold, an alarm is triggered.
[0067] 2. Alarm notification
[0068] When an anomaly is detected, the system will generate an alarm message and send it to the alarm platform, notifying the operation and maintenance personnel via email, SMS or instant messaging tools for timely processing.
[0069] Technical advantages:
[0070] Dynamic adaptability: The Prophet model generates a dynamic normal range, solving the problem that traditional fixed threshold methods cannot adapt to business fluctuations.
[0071] Efficiency: The combination of regularly updated prediction models and real-time detection enables rapid identification of anomalies in large-scale time series data scenarios.
[0072] Visualization-friendly: Tools such as Grafana provide an intuitive visualization interface, making it easier for operations and maintenance personnel to quickly locate problems.
[0073] Flexibility: Supports user-defined exception thresholds and alarm rules to meet the needs of different business scenarios.
[0074] The above description is only a preferred embodiment of the present invention and is only used to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of protection of the present invention.
Claims
1. A method for anomaly detection based on a time series database, characterized in that: The following steps are involved: 1) Periodically pull historical data of specified indicators from the time series database and use the historical data to train the prediction model; 2) Generate the predicted value at the current time point based on the trained prediction model, including the maximum predicted value and the minimum predicted value, and form a confidence interval; 3) Compare the actual collected current value with the confidence interval to determine whether it exceeds the upper and lower limits of the confidence interval. If so, mark it as abnormal; 4) writing the outliers, maximum predicted values, and minimum predicted values back into the time series database as new indicators, and displaying these indicators as lines of different colors on the user interface, with the parts outside the confidence interval highlighted in red; 5) When an anomaly is detected, the alarm mechanism is triggered and the alarm information is sent.
2. The method according to claim 1, characterized in that The time interval for periodically pulling historical data from the time series database is every fixed time period, and the time window of the historical data is the data of the most recent few days.
3. The method according to claim 1, characterized in that The prediction model is trained using the Prophet model or other models suitable for time series analysis, and each indicator corresponds to an independent prediction model.
4. The method according to claim 3, characterized in that After training is completed, the model is stored in a Map structure, where the key is the indicator name + _prophet, and the value is the corresponding training model.
5. The method according to claim 1, wherein The frequency of obtaining the predicted value is to obtain the predicted value of the current time point from the prediction model at regular intervals.
6. The method according to claim 1, characterized in that The standard for abnormality determination is that if the current value exceeds the maximum predicted value or is lower than the minimum predicted value by a certain ratio, it is determined to be an abnormal value; the specific ratio can be adjusted according to business needs.
7. The method according to claim 1, characterized in that Generate the yhat_upper value as a new indicator, name it as the original indicator name + _prophet + _upper, and write the value back to the time series database; Generate the yhat_lower value as a new indicator, name it as the original indicator name + _prophet + _lower, and write the value back to the time series database; The abnormal value is recorded as a new indicator named as the original indicator name + _prophet + _anomaly, with a value of 0 for abnormality or 1 for normality, and is written back to the time series database.
8. The method according to claim 7, characterized in that In the monitoring panel, the following indicators are displayed in the same chart: original indicator, that is, the original collected value; original indicator name + _prophet + _upper, original indicator name + _prophet + _lower, and original indicator name + _prophet + _anomaly.
9. The method according to claim 1, characterized in that Configure the percentage threshold for the current indicator value to exceed the confidence interval through the interface. When the actual value exceeds the threshold, an alarm is triggered and the alarm information is sent to the alarm platform.