Method, device, equipment, and storage medium for determining abnormal thresholds of performance indicators

By establishing a relationship model between performance indicators and TPS, dynamically adjusting the abnormality threshold, the accuracy of performance indicator abnormality detection in intelligent operation and maintenance scenarios is solved, and efficient abnormality detection is achieved.

CN114358581BActive Publication Date: 2025-08-29BEIJING ZHONGTI JUN COLOR INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111667055.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-08-29
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

In intelligent operation and maintenance scenarios, how to accurately determine the abnormal threshold of application system performance indicators to improve the accuracy of abnormal detection of performance indicators and reduce false alarms and missed reports of abnormal alarms.

Method used

By establishing a relationship model between performance indicators and TPS, the training data set is used to determine multiple predicted values ​​of performance indicators, and the residual mean and distribution are calculated in combination with abnormal sensitivity parameters, and the abnormal threshold is dynamically adjusted to achieve accurate detection of performance indicators.

Benefits of technology

Improve the accuracy and efficiency of abnormal detection of performance indicators, and reduce false alarms and missed reports of abnormal alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114358581B_ABST
    Figure CN114358581B_ABST
Patent Text Reader

Abstract

The present application relates to the field of intelligent operation and maintenance technology, and in particular to a method and apparatus, device, and storage medium for determining an abnormal threshold value of a performance indicator. In one embodiment of the present application, the method for determining an abnormal threshold value of a performance indicator includes: using a training data set to determine a relationship model between a first performance indicator and TPS, the training data set including multiple historical values ​​of the first performance indicator and multiple historical values ​​of at least one TPS; using the relationship model and multiple historical values ​​to obtain multiple predicted values ​​of the first performance indicator; determining an abnormal threshold value of the first performance indicator based on the multiple historical values ​​of the first performance indicator, the multiple predicted values ​​of the first performance indicator, and a predetermined abnormal sensitivity parameter, the abnormal threshold value is used to detect whether the first performance indicator is abnormal. The present application can accurately determine the abnormal threshold value of the performance indicator in combination with the dynamic changes of the performance indicator, thereby improving the accuracy and efficiency of performance indicator anomaly detection and reducing false positives and missed positives of abnormal alarms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent operation and maintenance technology, and in particular to a method and apparatus, device, and storage medium for determining an abnormal threshold value of a performance indicator. Background Art

[0002] In the field of system operation and maintenance, anomaly detection of application system performance indicators is extremely important. The accuracy of anomaly detection is often directly determined by the anomaly threshold. Currently, setting the anomaly threshold has always been a challenge. If the anomaly threshold is set too low, it is easy to cause a large number of false detections, while if it is set too high, it is easy to miss detections.

[0003] In intelligent operation and maintenance scenarios, how to accurately determine the abnormal thresholds of application system performance indicators to effectively improve the accuracy of performance indicator anomaly detection and reduce false positives and missed alarms is a technical problem that needs to be solved urgently. Summary of the Invention

[0004] In view of the above problems in the prior art, the present application provides a method and apparatus, device, and storage medium for determining the abnormal threshold of a performance indicator, which can accurately determine the abnormal threshold of a performance indicator, thereby improving the accuracy of performance indicator anomaly detection and reducing false positives and missed alarms.

[0005] To achieve the above objectives, the first aspect of the present application provides a method for determining an abnormal threshold value of a performance indicator, comprising:

[0006] Determining a relationship model between the first performance indicator and the TPS using a training data set, the training data set including multiple historical values ​​of the first performance indicator and multiple historical values ​​of at least one TPS;

[0007] Obtaining a plurality of predicted values ​​of the first performance indicator using the relational model and a plurality of historical values ​​of at least one TPS;

[0008] An abnormality threshold of the first performance indicator is determined according to multiple historical values ​​of the first performance indicator, multiple predicted values ​​of the first performance indicator, and a predetermined abnormality sensitivity parameter. The abnormality threshold is used to detect whether the first performance indicator is abnormal.

[0009] Therefore, the abnormal threshold of the performance indicator can be accurately determined or adjusted in combination with the dynamic changes of the performance indicator, thereby improving the accuracy and efficiency of performance indicator anomaly detection and effectively reducing false positives and missed alarms of abnormal alarms.

[0010] In some embodiments, determining the abnormality threshold of the first performance indicator based on multiple historical values ​​of the first performance indicator, multiple predicted values ​​of the first performance indicator, and a predetermined abnormality sensitivity parameter includes:

[0011] Calculating residuals between a plurality of predicted values ​​of the first performance indicator and corresponding predicted values ​​to obtain a residual mean and a residual distribution of the first performance indicator;

[0012] An abnormality threshold of the first performance indicator is determined according to the abnormality sensitivity parameter, the residual mean and the residual distribution of the first performance indicator.

[0013] Therefore, the abnormal threshold can be accurately determined by combining the residual changes of the performance indicators.

[0014] In some embodiments, the residual distribution of the first performance indicator comprises a plurality of data points, each of the data points representing a residual between a historical value of the first performance indicator and a corresponding predicted value;

[0015] Determining the abnormality threshold of the first performance indicator according to the abnormality sensitivity parameter, the residual mean and the residual distribution of the first performance indicator includes:

[0016] Using the residual mean of the first performance indicator as the current value of the abnormal threshold of the first performance indicator;

[0017] Determining, based on a current value of the abnormality threshold of the first performance indicator, n consecutive data points in a residual distribution of the first performance indicator that satisfy a first condition, where n is an integer greater than 1, and the first condition is that the residuals corresponding to the data points are greater than or equal to the current value of the abnormality threshold;

[0018] When n is less than or equal to the abnormal sensitivity parameter, the abnormal threshold of the performance indicator is set to the current value of the abnormal threshold.

[0019] Therefore, the abnormal threshold of the performance indicator can be accurately determined through the residual distribution.

[0020] In some embodiments, determining the abnormality threshold of the first performance indicator based on the abnormality sensitivity parameter, the residual mean and the residual distribution of the first performance indicator further includes:

[0021] When n is greater than the abnormality sensitivity parameter, the current value of the abnormality threshold is iteratively updated according to a predetermined step size until n is less than or equal to the abnormality sensitivity parameter.

[0022] Therefore, the abnormal threshold of the performance indicator can be accurately determined by iterative updating using the residual distribution.

[0023] A second aspect of the present application provides a method for detecting anomalies in performance indicators, including:

[0024] Determining a relationship model between the first performance indicator and the TPS using a training data set, the training data set including multiple historical values ​​of the first performance indicator and multiple historical values ​​of at least one TPS;

[0025] Obtaining a plurality of predicted values ​​of the first performance indicator using the relational model and a plurality of historical values ​​of at least one TPS;

[0026] determining an abnormality threshold of the first performance indicator based on multiple historical values ​​of the first performance indicator, multiple predicted values ​​of the first performance indicator, and a predetermined abnormality sensitivity parameter, wherein the abnormality threshold is used to detect whether the first performance indicator is abnormal;

[0027] Whether the first performance indicator is abnormal is determined according to the current value of the TPS, the current value of the first performance indicator, the relationship model, and an abnormality threshold of the first performance indicator.

[0028] Therefore, the abnormal threshold of the performance indicator can be accurately determined or adjusted in combination with the dynamic changes of the performance indicator, and the accuracy and efficiency of the performance indicator anomaly detection can be improved, thereby effectively reducing the false positives and missed negatives of abnormal alarms.

[0029] In some embodiments, determining whether the first performance indicator is abnormal based on the current value of the TPS, the current value of the first performance indicator, the relationship model, and the abnormality threshold of the first performance indicator includes:

[0030] Determining a current predicted value of the first performance indicator according to the current value of the TPS and the relationship model;

[0031] calculating a residual between a current predicted value of the first performance indicator and a current value of the performance indicator;

[0032] When the residual is greater than an abnormal threshold of the first performance indicator, it is determined that the first performance indicator is abnormal.

[0033] Therefore, the residual of the performance indicator can be used to improve the accuracy and efficiency of performance indicator anomaly detection.

[0034] A third aspect of the present application provides a device for determining an abnormal threshold value of a performance indicator, comprising:

[0035] a model building module, configured to determine a relationship model between a first performance indicator and TPS using a training data set, wherein the training data set includes a plurality of historical values ​​of the first performance indicator and a plurality of historical values ​​of at least one TPS;

[0036] A threshold determination module is used to obtain multiple predicted values ​​of the first performance indicator using the relationship model and multiple historical values ​​of at least one TPS; and is used to determine an abnormal threshold of the first performance indicator based on the multiple historical values ​​of the first performance indicator, the multiple predicted values ​​of the first performance indicator and a predetermined abnormal sensitivity parameter, wherein the abnormal threshold is used to detect whether the first performance indicator is abnormal.

[0037] A fourth aspect of the present application provides a method for detecting anomalies in performance indicators, including:

[0038] a model building module, configured to determine a relationship model between a first performance indicator and TPS using a training data set, wherein the training data set includes a plurality of historical values ​​of the first performance indicator and a plurality of historical values ​​of at least one TPS;

[0039] a threshold determination module, configured to obtain a plurality of predicted values ​​of the first performance indicator using the relationship model and a plurality of historical values ​​of at least one TPS; and to determine an abnormality threshold of the first performance indicator based on the plurality of historical values ​​of the first performance indicator, the plurality of predicted values ​​of the first performance indicator, and a predetermined abnormality sensitivity parameter, wherein the abnormality threshold is used to detect whether the first performance indicator is abnormal;

[0040] An abnormality determination module is used to determine whether the first performance indicator is abnormal based on the current value of the TPS, the current value of the first performance indicator, the relationship model and the abnormality threshold of the first performance indicator.

[0041] In a fifth aspect, the present application provides a computing device comprising: a processor and a memory; wherein the memory is used to store program instructions, and when the program instructions are executed by the processor, the computing device implements the method of the first aspect or the method of the second aspect.

[0042] In a sixth aspect, the present application provides a computer-readable storage medium having program instructions stored thereon. When the program instructions are executed by a computer, the computer implements the method of the first aspect or the method of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The following further illustrates the various features of the present application and the relationships between the various features with reference to the accompanying drawings. The accompanying drawings are all exemplary, and some features are not shown in actual proportion. In addition, some drawings may omit features that are customary in the field to which the present application relates and are not necessary for the present application, or additional features that are not necessary for the present application may be shown. The combination of the various features shown in the accompanying drawings is not intended to limit the present application. In addition, throughout this specification, the same figure numbers refer to the same content. The specific description of the drawings is as follows:

[0044] Figure 1 is a flow chart of a method for determining an abnormal threshold value of a performance indicator of the present application;

[0045] Figure 2 is a flow chart of the performance indicator anomaly detection method of the present application;

[0046] Figure 3 It is a structural diagram of a device for determining an abnormal threshold value of a performance indicator of the present application;

[0047] Figure 4 It is a structural diagram of the performance index anomaly detection device of the present application;

[0048] Figure 5 It is a structural diagram of the computing device of this application. DETAILED DESCRIPTION

[0049] In order to accurately describe the technical content of this application and to accurately understand this application, the following explanations or definitions are given for the terms used in this specification before describing the specific implementation methods.

[0050] Performance indicators: refers to the resources used by the service when processing business, such as CPU utilization, disk read and write speed, memory usage, network access rate, etc.

[0051] In related technologies, there are two common solutions for setting anomaly thresholds in monitoring indicator anomaly detection:

[0052] 1) Fixed Threshold Method: A static threshold is set based on expert experience. The performance indicator's value is compared with the static threshold to determine whether the monitoring indicator is abnormal. This fixed threshold method is not suitable for scenarios with a large number of performance indicators of varying types and complex patterns. It is particularly unsuitable for distributed microservice architectures with a large number of services and frequent service changes.

[0053] 2) Data-driven dynamic threshold setting method: Use time series analysis algorithms such as ARIMA to analyze the historical time series data of monitoring indicators to obtain an indicator prediction model. Then use the monitoring data in the most recent window to predict the performance indicator values ​​at future moments. Calculate the percentage deviation between the predicted performance indicator values ​​and the actual performance indicator values ​​collected at that future moment. Determine whether the deviation percentage is within the preset abnormal tolerance error percentage. If so, it is considered normal. If not, the corresponding performance indicator is abnormal.

[0054] With the rapid development of microservices, the architecture of intelligent operations and maintenance systems has shifted from a stovepipe structure to a distributed microservice architecture. Under this microservices architecture, the number of services has increased dramatically, and service changes are frequent. Solution 2) above presents at least the following issues:

[0055] 1. Due to the instability of time series data, the indicator prediction model based on historical time series data has large errors, resulting in low accuracy of the final anomaly detection results;

[0056] 2. The anomaly tolerance error percentage needs to be set based on manual experience. First, for a large number of services, it is difficult to quickly and effectively use manual experience to give an accurate anomaly tolerance error percentage for each service. Second, it is impossible to automatically determine a reasonable range based on the fluctuations of different performance indicators. Moreover, it is impossible to effectively adjust the value range of the anomaly tolerance error to adapt to the sensitivity of different performance indicators to anomalies. This will inevitably significantly reduce the accuracy of anomaly detection results.

[0057] Under the microservice architecture, each service can comply with the single responsibility principle as much as possible according to the principle of service division. Under a single responsibility, the resource usage of each service is extremely correlated with the number of transactions processed per second (TPS, Transactions Per Second) carried by the service. In view of this, the embodiment of the present application provides a monitoring indicator anomaly detection method and apparatus, equipment, and storage medium, which establishes a relationship model between TPS and the performance index of the service based on the correlation between the resource usage of the service and the TPS carried by the service, and uses the relationship model to obtain a predicted value. The abnormal threshold of the performance index is calculated based on the predicted value, the historical value, and the pre-set abnormal sensitivity parameter of the performance index, and the abnormal threshold can be used to determine whether the corresponding performance index is abnormal. Thus, the embodiment of the present application can dynamically adjust the abnormal threshold in combination with the real-time changes of the performance index, accurately determine the abnormal threshold of the performance index, thereby improving the accuracy and efficiency of performance index anomaly detection, and effectively reducing the false positives and omissions of abnormal alarms.

[0058] The embodiments of the present application are applicable to various intelligent operation and maintenance scenarios, especially those using microservices, complex operation and maintenance, and those with massive amounts of monitoring data. For example, the embodiments of the present application are applicable to, but not limited to, the following scenarios: operation and maintenance scenarios with large-scale system software, operation and maintenance scenarios with high frequency of system software changes, operation and maintenance scenarios with complex system software call relationships, operation and maintenance scenarios with large amounts of monitoring data, and so on.

[0059] The "service" in the embodiments of the present application refers to a microservice, which refers to a single small service with business functions, and a microservice is only responsible for one type of business. A service can encapsulate business capabilities under the principle of single responsibility and can be deployed and run independently. In actual applications, a system or a large service can be split into multiple microservices with single functions according to actual needs. For example, the "service" in the embodiments of the present application can be but is not limited to a ticket sales service, a prize redemption service, etc. For a certain service, the service uses a variety of resources when processing business, and the multiple resources correspond to multiple performance indicators. For example, the performance indicators monitored by a certain service may include but are not limited to CPU utilization, disk read and write speed, memory occupancy, network access rate, etc.

[0060] The "first performance indicator" in this article refers to any performance indicator of a service.

[0061] Figure 1 The flowchart of the method for determining the abnormal threshold value of the performance indicator provided by the embodiment of the present application is shown. Figure 1 As shown, the method for determining the abnormal threshold value of the performance indicator provided in the embodiment of the present application may include:

[0062] Step S110, determining a relationship model between a first performance indicator and TPS using a training data set, wherein the training data set includes multiple historical values ​​of the first performance indicator and multiple historical values ​​of at least one TPS;

[0063] In some implementations, a correlation algorithm may be used to first determine the TPS associated with the first performance indicator, and then a relationship model between the first performance indicator and the TPS may be constructed using multiple historical values ​​of the TPS associated with the first performance indicator and multiple historical values ​​of the first performance indicator.

[0064] Here, a correlation algorithm is a method used to measure the correlation between two variables. In some examples, the correlation algorithm can be, but is not limited to, the Pearson correlation coefficient method, the Spearman rank correlation coefficient method, the Kendall rank correlation coefficient method, the Maximum Information Coefficient (MIC) method, etc.

[0065] Here, the relationship model may be, but is not limited to, a predefined function, a machine learning model, or other methods. Specifically, an exemplary process for constructing the relationship model may be: based on the predefined function or a predetermined machine learning algorithm, multiple historical values ​​of TPS associated with the first performance indicator and multiple historical values ​​of the first performance indicator are used to perform fitting to obtain a relationship model between the first performance indicator and TPS.

[0066] In some examples, the predefined function may be, but is not limited to, a custom function, a linear function as shown in the following formula (1), or a logarithmic function as shown in the following formula (2).

[0067] y=k*x+b (1)

[0068] y=a*log2b*x+c (2)

[0069] Wherein, y represents the first performance index, x represents TPS, k, a, b, and c represent fitting parameters, respectively, and k, a, b, and c can be determined by fitting.

[0070] In some examples, the machine learning model can be, but is not limited to, support vector machines (SVM), rank-preserving regression, polynomial regression, neural networks, decision trees, etc.

[0071] In some embodiments, the training data set may include E samples (E is an integer greater than 1, for example, E can be 20, 50, 100, or other numerical values). Each sample may include the historical value of at least one performance indicator and at least one TPS historical value. The historical value in each sample is consistent with the historical value time information (for example, the collection time or the statistical time). For the i-th sample in the training sample set (i is an integer greater than or equal to 1 and less than or equal to E), its time information is 12:00:40. The i-th sample includes the historical CPU utilization value and the historical TPS value at the collection time of 12:00:40.

[0072] Here, the historical value can be a value collected at a certain past moment. The historical values ​​of the performance indicator and the historical TPS values ​​in the sample can be obtained through aggregation. Typically, TPS is counted once per second, and the first performance indicator is collected every 30 seconds. The aggregation process can be: calculate the average of the TPS values ​​collected within a predetermined time period (for example, 1 minute) from the collection time of the first performance indicator as the end time, and use this average as the collection value of TPS at that collection time, i.e., the historical value of TPS at that moment. For example, the collection time of the first performance indicator can be: 12:00:10, 12:00:40, 12:01:10. Taking the collection time of 12:00:40 as an example, the average of all TPS values ​​between the time 11:59:40 and the time 12:00:40 is taken as the historical value of TPS at the time 12:00:40, and the average of all performance indicator values ​​collected between the time 11:59:40 and the time 12:00:40 (that is, the performance indicator value of 12:00:10 and the performance indicator value of 12:00:40) is taken as the historical value of the first performance indicator at the time 12:00:40.

[0073] Here, the training data set can be obtained by performing data preprocessing on historical data of performance indicators and historical data of TPS. The data preprocessing may include: data cleaning, data type conversion, data standardization and / or data alignment, etc.

[0074] Step S120, using the relational model and the multiple historical values, obtaining multiple predicted values ​​of the first performance indicator;

[0075] In specific applications, the relational model can be run on some or all samples in the training dataset to obtain multiple corresponding prediction values. For example, by running the relational model based on the TPS history associated with the first performance indicator in the i-th sample in the training dataset, the i-th predicted value for the first performance indicator can be obtained.

[0076] Step S130 , determining an abnormality threshold of the first performance indicator according to multiple historical values ​​of the first performance indicator, multiple predicted values ​​of the first performance indicator, and a predetermined abnormality sensitivity parameter, where the abnormality threshold is used to detect whether the first performance indicator is abnormal.

[0077] Here, the anomaly sensitivity parameter represents the number n of consecutive anomaly data points that can be tolerated without being reported. The larger the value of n, the smaller the anomaly threshold.

[0078] In some embodiments, step S130 may include: step a1, calculating the residuals between multiple predicted values ​​of the first performance indicator and the corresponding predicted values ​​to obtain the residual mean and residual distribution of the first performance indicator; step a2, determining the abnormal threshold of the first performance indicator based on the abnormal sensitivity parameter, the residual mean and residual distribution of the first performance indicator.

[0079] In some embodiments, the residual distribution of the first performance indicator includes multiple data points, each data point represents the residual between a historical value of the first performance indicator and a corresponding predicted value, and the residual satisfies a Gaussian distribution.

[0080] In some embodiments, step a2 may include:

[0081] Step a21, using the residual mean of the first performance indicator as the current value of the abnormal threshold of the first performance indicator;

[0082] Step a22: determining, based on the current value of the abnormality threshold of the first performance indicator, n consecutive data points in the residual distribution of the first performance indicator that meet a first condition, where n is an integer greater than 1, and the first condition is that the residual corresponding to the data point is greater than or equal to the current value of the abnormality threshold;

[0083] Step a23: When n is less than or equal to the abnormal sensitivity parameter, the abnormal threshold of the first performance indicator is set to the current value of the abnormal threshold.

[0084] In some embodiments, step a2 may further include: step a24, when n is greater than the abnormality sensitivity parameter, iteratively updating the current value of the abnormality threshold according to a predetermined step size, and repeating step a22 until n is less than or equal to the abnormality sensitivity parameter. Specifically, the current value of the abnormality threshold thresh ab Iterative update is performed according to the following formula:

[0085] thresh ab =thresh ab *(1+η)

[0086] The parameter η is a predetermined step size for iterative updating, and its specific value can be set as needed. In some examples, the value of the parameter η (i.e., the predetermined step size) is less than 0.01.

[0087] Then, step a22 is repeated to determine the updated abnormality threshold n. If n is less than or equal to the abnormality sensitivity parameter, step a23 is executed, and the iteration ends. If n is greater than the abnormality sensitivity parameter, step a24 is continued. This iteration continues until n is less than or equal to the abnormality sensitivity parameter.

[0088] In some examples, the residual can be calculated using the following formula (3):

[0089]

[0090] Among them, resdual i represents the residual of the first performance indicator of the i-th sample in the training data set, y pred is the predicted value of the first performance indicator of the i-th sample, y pred Obtained through the TPS history value and relationship model of the i-th sample; y true is the historical value of the first performance indicator of the i-th sample.

[0091] In some examples, the residual mean can be calculated as follows:

[0092]

[0093] Among them, resdual_mean represents the residual mean, and E is the number of samples in the training data set.

[0094] In practical applications, any category of performance indicators can be calculated using formula (3) to calculate the residual and formula (4) to calculate the residual mean.

[0095] It should be noted that the anomaly threshold determination method in the embodiments of the present application can be applied to performance indicators that can indicate the consumption of various resources of a service, such as CPU, disk, network, and memory. That is, the anomaly threshold determination method in the embodiments of the present application can be applied to various performance indicators, including but not limited to CPU utilization, disk read and write speeds, memory occupancy, and network access rates. The principles of the anomaly threshold determination methods for different performance indicators are the same, for example, the values ​​of the anomaly sensitivity parameter and the predetermined step size can be different.

[0096] Figure 2 The flowchart of the method for detecting anomalies of performance indicators provided in the embodiment of the present application is shown. Figure 2 As shown, the performance indicator anomaly detection method provided in the embodiment of the present application may include: steps S110 to S130 mentioned above and the following step S140: determining whether the first performance indicator is abnormal based on the current value of TPS, the current value of the first performance indicator, the relationship model and the anomaly threshold of the first performance indicator.

[0097] In some embodiments, step S140 may include:

[0098] Step b1, determining a current predicted value of the first performance indicator based on the current value of TPS and the relationship model;

[0099] Step b2, calculating the residual between the current predicted value of the first performance indicator and the current value of the performance indicator;

[0100] Step b3: When the residual is greater than the abnormal threshold of the first performance indicator, determine that the first performance indicator is abnormal.

[0101] Step b4: When the residual is less than or equal to the abnormal threshold of the first performance indicator, determine that the first performance indicator is normal.

[0102] Here, the current value can be the current value collected at the moment. The current value of the performance indicator and the current value of TPS can be obtained through aggregation. The aggregation method is the same as the historical value mentioned above and will not be repeated here.

[0103] Figure 3 FIG2 shows a schematic diagram of the structure of a device for determining an abnormal threshold value of a performance indicator provided by an embodiment of the present application. Figure 3 As shown, the apparatus for determining an abnormal threshold value of a performance indicator provided in an embodiment of the present application may include:

[0104] A model building module 31 is configured to determine a relationship model between a first performance indicator and TPS using a training data set, wherein the training data set includes a plurality of historical values ​​of the first performance indicator and a plurality of historical values ​​of at least one TPS;

[0105] The threshold determination module 32 is used to obtain multiple predicted values ​​of the first performance indicator using a relational model and multiple historical values ​​of at least one TPS; and is used to determine an abnormal threshold of the first performance indicator based on the multiple historical values ​​of the first performance indicator, the multiple predicted values ​​of the first performance indicator and a predetermined abnormal sensitivity parameter, wherein the abnormal threshold is used to detect whether the first performance indicator is abnormal.

[0106] In some embodiments, the threshold determination module 32 is specifically used to: calculate the residuals between multiple predicted values ​​of the first performance indicator and the corresponding predicted values ​​to obtain the residual mean and residual distribution of the first performance indicator; determine the abnormal threshold of the first performance indicator based on the abnormal sensitivity parameter, the residual mean and residual distribution of the first performance indicator.

[0107] In some embodiments, the residual distribution of the first performance indicator includes multiple data points, each of which represents the residual between a historical value of the first performance indicator and the corresponding predicted value; the threshold determination module 32 is specifically used to: use the residual mean of the first performance indicator as the current value of the abnormality threshold of the first performance indicator; based on the current value of the abnormality threshold of the first performance indicator, determine n consecutive data points in the residual distribution of the first performance indicator that meet a first condition, where n is an integer greater than 1, and the first condition is: the residual corresponding to the data point is greater than or equal to the current value of the abnormality threshold; when n is less than or equal to the abnormal sensitivity parameter, set the abnormality threshold of the performance indicator to the current value of the abnormality threshold.

[0108] In some embodiments, the threshold determination module 32 is further configured to: when n is greater than an abnormality sensitivity parameter, iteratively update the current value of the abnormality threshold according to a predetermined step size until n is less than or equal to the abnormality sensitivity parameter.

[0109] Figure 4 The schematic diagram of the structure of the abnormality detection device of the performance index provided in the embodiment of the present application is shown. Figure 4 As shown, the performance indicator anomaly detection device provided in this embodiment of the present application may include: a model construction module 31, a threshold determination module 32, and an anomaly determination module 33. The functions of the model construction module 31 and the threshold determination module 32 are described above and are not further described. The anomaly determination module 33 is configured to determine whether the first performance indicator is abnormal based on the current value of the TPS, the current value of the first performance indicator, the relationship model, and the anomaly threshold of the first performance indicator.

[0110] In some embodiments, the abnormality determination module 33 is specifically used to: determine the current predicted value of the first performance indicator based on the current value of the TPS and the relationship model; calculate the residual between the current predicted value of the first performance indicator and the current value of the performance indicator; and determine that the first performance indicator is abnormal when the residual is greater than the abnormality threshold of the first performance indicator.

[0111] In some embodiments, the abnormality determination module 33 is further configured to: determine that the first performance indicator is normal when the residual is less than or equal to an abnormality threshold of the first performance indicator.

[0112] Figure 5 5 is a schematic structural diagram of a computing device 500 provided in an embodiment of the present application. The computing device 500 includes: a processor 510 and a memory 520.

[0113] The processor 510 may be connected to a memory 520. The memory 520 may be used to store the program code and data. Therefore, the memory 520 may be a storage unit within the processor 510, an external storage unit independent of the processor 510, or a component including both a storage unit within the processor 510 and an external storage unit independent of the processor 510.

[0114] Optionally, the computing device 500 may further include: a communication interface 530. It should be understood that Figure 5 The communication interface 530 in the computing device 500 shown can be used to communicate with other devices.

[0115] Optionally, the computing device 500 may further include a bus 540. The memory 520 and the communication interface 530 may be connected to the processor 510 via the bus 540. The bus 540 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus 540 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 The fact that only one line is used does not mean that there is only one bus or one type of bus.

[0116] It should be understood that in the embodiment of the present application, the processor 510 can adopt a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Alternatively, the processor 510 adopts one or more integrated circuits to execute relevant programs to implement the technical solutions provided in the embodiment of the present application.

[0117] The memory 520 may include a read-only memory and a random access memory, and provides instructions and data to the processor 510. A portion of the processor 510 may also include a non-volatile random access memory. For example, the processor 510 may also store information about the device type.

[0118] When the computing device 500 is running, the processor 510 executes the computer execution instructions in the memory 520 to perform the operating steps of the above-mentioned method for determining an abnormality threshold of a performance indicator or the method for detecting an abnormality of a performance indicator.

[0119] It should be understood that the computing device 500 according to the embodiment of the present application can correspond to the corresponding subject in executing the method according to each embodiment of the present application, and the above-mentioned and other operations and / or functions of each module in the computing device 500 are respectively for implementing the corresponding processes of each method of the present embodiment. For the sake of brevity, they will not be repeated here.

[0120] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0121] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is used to execute the above-mentioned method for determining an abnormality threshold of a performance indicator or the method for detecting an abnormality of a performance indicator.

[0122] The computer storage medium of the embodiment of the present application can adopt any combination of one or more computer-readable media.Computer-readable media can be computer-readable signal media or computer-readable storage media.Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof.More specific examples (non-exhaustive list) of computer-readable storage media include: electrical connection with one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination thereof.In this document, computer-readable storage media can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.

[0123] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0124] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0125] The computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0126] The present application also provides a computer program product, comprising a computer program that, when executed by a processor, causes the processor to execute the above-described method for determining an abnormality threshold of a performance indicator or the method for detecting an abnormality of a performance indicator. The computer program product may be written in one or more programming languages, including but not limited to object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C.

[0127] Note that the above are only preferred embodiments of the present application and the technical principles employed. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of protection of the present application, all of which fall within the scope of protection of the present application.

Claims

1. A method for determining an abnormal threshold value of a performance indicator, characterized in that: include: Determine a relationship model between a first performance indicator and TPS using a training data set, the training data set including multiple historical values ​​of the first performance indicator and multiple historical values ​​of at least one TPS; wherein the first performance indicator includes: CPU utilization, disk read / write speed, memory occupancy, and network access rate; and the TPS is the number of transactions processed per second; Obtaining a plurality of predicted values ​​of the first performance indicator using the relational model and a plurality of historical values ​​of at least one TPS; Determining an abnormality threshold of the first performance indicator based on multiple historical values ​​of the first performance indicator, multiple predicted values ​​of the first performance indicator, and a predetermined abnormality sensitivity parameter, wherein the abnormality threshold is used to detect whether the first performance indicator is abnormal; the residual distribution of the first performance indicator includes multiple data points, each of which represents a residual between a historical value of the first performance indicator and a corresponding predicted value; determining the abnormality threshold includes: Calculate the residuals between multiple historical values ​​of the first performance indicator and the corresponding predicted values ​​to obtain the residual mean and residual distribution of the first performance indicator. Specifically, calculate the residuals using the following formula: Among them, resdual i represents the residual of the first performance indicator of the i-th sample in the training dataset, is the predicted value of the first performance indicator of the i-th sample, Obtained through the TPS history value and relationship model of the i-th sample; is the historical value of the first performance indicator of the i-th sample; Using the residual mean of the first performance indicator as the current value of the abnormal threshold of the first performance indicator; Determining, based on a current value of the abnormality threshold of the first performance indicator, n consecutive data points in a residual distribution of the first performance indicator that satisfy a first condition, where n is an integer greater than 1, and the first condition is that the residuals corresponding to the data points are greater than or equal to the current value of the abnormality threshold; When n is less than or equal to the abnormal sensitivity parameter, setting the abnormal threshold of the performance indicator to the current value of the abnormal threshold; When n is greater than the abnormal sensitivity parameter, the current value of the abnormal threshold is iteratively updated according to a predetermined step size until n is less than or equal to the abnormal sensitivity parameter; specifically, the current value of the abnormal threshold is iteratively updated according to the following formula: Among them, thresh ab is the current value of the abnormal threshold, η is the predetermined step size, and its value is less than 0.

01.

2. A method for detecting anomalies in performance indicators, characterized in that: include: Determining a relationship model between the first performance indicator and the TPS using a training data set, the training data set including multiple historical values ​​of the first performance indicator and multiple historical values ​​of at least one TPS; Obtaining a plurality of predicted values ​​of the first performance indicator using the relational model and a plurality of historical values ​​of at least one TPS; Using the method for determining an abnormality threshold of a performance indicator according to claim 1, determining an abnormality threshold of the first performance indicator based on multiple historical values ​​of the first performance indicator, multiple predicted values ​​of the first performance indicator, and a predetermined abnormality sensitivity parameter, the abnormality threshold being used to detect whether the first performance indicator is abnormal; Whether the first performance indicator is abnormal is determined according to the current value of the TPS, the current value of the first performance indicator, the relationship model, and an abnormality threshold of the first performance indicator.

3. The method for detecting anomalies in performance indicators according to claim 2, characterized in that: The determining whether the first performance indicator is abnormal according to the current value of the TPS, the current value of the first performance indicator, the relationship model, and the abnormality threshold of the first performance indicator includes: Determining a current predicted value of the first performance indicator according to the current value of the TPS and the relationship model; calculating a residual between a current predicted value of the first performance indicator and a current value of the performance indicator; When the residual is greater than an abnormal threshold of the first performance indicator, it is determined that the first performance indicator is abnormal.

4. A device for determining an abnormal threshold value of a performance indicator, characterized in that: include: a model building module, configured to determine a relationship model between a first performance indicator and TPS using a training data set, the training data set comprising multiple historical values ​​of the first performance indicator and multiple historical values ​​of at least one TPS; wherein the first performance indicator comprises: CPU utilization, disk read / write speed, memory occupancy, and network access rate; and the TPS is the number of transactions processed per second; A threshold determination module is configured to obtain multiple predicted values ​​of the first performance indicator using the relationship model and multiple historical values ​​of at least one TPS; and to determine an abnormality threshold of the first performance indicator based on the multiple historical values ​​of the first performance indicator, the multiple predicted values ​​of the first performance indicator, and a predetermined abnormality sensitivity parameter, wherein the abnormality threshold is used to detect whether the first performance indicator is abnormal; the residual distribution of the first performance indicator includes multiple data points, each of which represents a residual between a historical value of the first performance indicator and a corresponding predicted value; and determining the abnormality threshold includes: Calculate the residuals between multiple historical values ​​of the first performance indicator and the corresponding predicted values ​​to obtain the residual mean and residual distribution of the first performance indicator. Specifically, calculate the residuals using the following formula: Among them, resdual i represents the residual of the first performance indicator of the i-th sample in the training dataset, is the predicted value of the first performance indicator of the i-th sample, Obtained through the TPS history value and relationship model of the i-th sample; is the historical value of the first performance indicator of the i-th sample; Using the residual mean of the first performance indicator as the current value of the abnormal threshold of the first performance indicator; Determining, based on a current value of the abnormality threshold of the first performance indicator, n consecutive data points in a residual distribution of the first performance indicator that satisfy a first condition, where n is an integer greater than 1, and the first condition is that the residuals corresponding to the data points are greater than or equal to the current value of the abnormality threshold; When n is less than or equal to the abnormal sensitivity parameter, setting the abnormal threshold of the performance indicator to the current value of the abnormal threshold; When n is greater than the abnormal sensitivity parameter, the current value of the abnormal threshold is iteratively updated according to a predetermined step size until n is less than or equal to the abnormal sensitivity parameter; specifically, the current value of the abnormal threshold is iteratively updated according to the following formula: Among them, thresh ab is the current value of the abnormal threshold, η is the predetermined step size, and its value is less than 0.

01.

5. A device for detecting abnormality of performance indicators, characterized in that: include: a model building module, configured to determine a relationship model between a first performance indicator and TPS using a training data set, the training data set comprising multiple historical values ​​of the first performance indicator and multiple historical values ​​of at least one TPS; wherein the first performance indicator comprises: CPU utilization, disk read / write speed, memory occupancy, and network access rate; and the TPS is the number of transactions processed per second; A threshold determination module is configured to obtain multiple predicted values ​​of the first performance indicator using the relationship model and multiple historical values ​​of at least one TPS; and to determine an abnormality threshold of the first performance indicator based on the multiple historical values ​​of the first performance indicator, the multiple predicted values ​​of the first performance indicator, and a predetermined abnormality sensitivity parameter, wherein the abnormality threshold is used to detect whether the first performance indicator is abnormal; the residual distribution of the first performance indicator includes multiple data points, each of which represents a residual between a historical value of the first performance indicator and a corresponding predicted value; and determining the abnormality threshold includes: Calculate the residuals between multiple historical values ​​of the first performance indicator and the corresponding predicted values ​​to obtain the residual mean and residual distribution of the first performance indicator. Specifically, calculate the residuals using the following formula: Among them, resdual i represents the residual of the first performance indicator of the i-th sample in the training dataset, is the predicted value of the first performance indicator of the i-th sample, Obtained through the TPS history value and relationship model of the i-th sample; is the historical value of the first performance indicator of the i-th sample; Using the residual mean of the first performance indicator as the current value of the abnormal threshold of the first performance indicator; Determining, based on a current value of the abnormality threshold of the first performance indicator, n consecutive data points in a residual distribution of the first performance indicator that satisfy a first condition, where n is an integer greater than 1, and the first condition is that the residuals corresponding to the data points are greater than or equal to the current value of the abnormality threshold; When n is less than or equal to the abnormal sensitivity parameter, setting the abnormal threshold of the performance indicator to the current value of the abnormal threshold; When n is greater than the abnormal sensitivity parameter, the current value of the abnormal threshold is iteratively updated according to a predetermined step size until n is less than or equal to the abnormal sensitivity parameter; specifically, the current value of the abnormal threshold is iteratively updated according to the following formula: Among them, thresh ab is the current value of the abnormal threshold, η is the predetermined step size, and is less than 0.01; An abnormality determination module is used to determine whether the first performance indicator is abnormal based on the current value of the TPS, the current value of the first performance indicator, the relationship model and the abnormality threshold of the first performance indicator.

6. A computing device, characterized in that include: processor and memory; The memory is used to store program instructions, and when the program instructions are executed by the processor, the computing device implements the method of claim 1 or any one of claims 2 to 3.

7. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a computer, the computer implements the method of claim 1 or the method of any one of claims 2 to 3.

Citation Information

Patent Citations

  • Distributed system CPU abnormity detection method and device and storage medium

    CN109766244A

  • Method and device for detecting network performance abnormity, equipment and storage medium

    CN110278121A