Dynamic Threshold Calculation Method, Device, Equipment and Medium

Through frequency graph analysis and clustering technology of time series data, the dynamic threshold of the system is calculated, which solves the problem of high frequency of false alarms and missed alarms in traditional technologies, and achieves more accurate alarm processing.

CN115145781BActive Publication Date: 2025-06-20CHINA MOBILE GROUP ANHUI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110336827.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-29
Publication Date
2025-06-20
Estimated Expiration
2041-03-29

AI Technical Summary

Technical Problem

In the prior art, the judgment methods and processing methods of system alarms are still in traditional ways, resulting in the frequency of false alarms and missed alarms being too high and cannot be accurately distinguished.

Method used

By obtaining the timing data and its frequency, a frequency graph is generated to determine the busy time threshold, cluster the busy time period to determine the busy time interval, and calculate the dynamic threshold, including the false alarm threshold and the missed alarm threshold.

Benefits of technology

The accurate distinction between false alarms and missed alarms is achieved, the frequency of false alarms and missed alarms is reduced, and the accuracy of alarms is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115145781B_ABST
    Figure CN115145781B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a dynamic threshold calculation method, device, equipment, and medium. The method includes: obtaining time series data and the frequency of the time series data within a first preset time period; generating a frequency diagram based on the frequency of the time series data, where the frequency diagram includes single hump feature information or multi-hump feature information; determining a busy-idle threshold for dividing the time series data based on the frequency diagram; clustering the busy time periods of the time series data according to a second preset time period to determine the busy time interval of the time series data, where the busy time period is the time period obtained when the time series data is greater than the busy-idle threshold, and the second preset time period includes the first preset time period; in the case where the maximum value of the time series data in the busy time interval does not meet the preset condition, calculating the dynamic threshold of the time series data according to the time series data, where the dynamic threshold includes a false alarm threshold and a missed alarm threshold. According to the embodiment of the present application, the dynamic threshold for dividing false alarms and missed alarms can be accurately obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of operation management, and particularly relates to a dynamic threshold calculation method, device, equipment and medium. Background Art

[0002] At present, the operation and maintenance mode of Internet Technology (IT) has evolved from traditional manual operation and maintenance, to script operation and maintenance, to platform operation and maintenance. Most customers have also achieved batch regular operation and maintenance operations and fault handling. However, the judgment method and handling means for system alarms still adopt traditional methods. Most traditional methods make some analysis and judgments based on alarm indicators, and the set threshold is too broad and is a static value, resulting in too high frequencies of false alarms and missed alarms, which cannot solve the fundamental problem and cannot achieve the accuracy required in production.

[0003] Therefore, how to accurately obtain a dynamic threshold for distinguishing false alarms and missed alarms is an urgent problem to be solved. Summary of the Invention

[0004] The embodiments of this application provide a dynamic threshold calculation method, device, equipment and medium, which can accurately obtain a dynamic threshold for distinguishing false alarms and missed alarms.

[0005] In a first aspect, the embodiments of this application provide a dynamic threshold calculation method, and the method includes: obtaining time series data and the frequency of the time series data within a first preset time period; generating a frequency diagram based on the frequency of the time series data, where the frequency diagram includes single hump feature information or multi-hump feature information; determining a busy / idle threshold for dividing the time series data based on the frequency diagram; clustering the busy time periods of the time series data according to a second preset time period to determine the busy time interval of the time series data, where the busy time period is the time period obtained when the time series data is greater than the busy / idle threshold, and the second preset time period includes the first preset time period; in the case that the maximum value of the time series data in the busy time interval does not meet the preset condition, calculating a dynamic threshold of the time series data according to the time series data, where the dynamic threshold includes a false alarm threshold and a missed alarm threshold.

[0006] In some embodiments of the first aspect, determining a busy / idle threshold for dividing the time series data based on the frequency diagram includes: in the case that the frequency diagram presents single hump feature information, determining the time series data corresponding to a preset range length as the busy / idle threshold; in the case that the frequency diagram presents multi-hump feature information, determining the number of troughs in the frequency diagram; in the case that the number of troughs is 1, determining the time series data corresponding to the trough as the busy / idle threshold; in the case that the number of troughs is greater than 1, determining the average value of the time series data corresponding to multiple troughs as the busy / idle threshold.

[0007] In some embodiments of the first aspect, clustering the busy time periods of the time series data according to a second preset time period to determine the busy time intervals of the time series data includes: obtaining a plurality of time series data sets within the second preset time period; counting the occurrence frequencies of the busy time periods in the plurality of time series data sets; and determining the busy time intervals of the time series data according to the occurrence frequencies of the busy time periods.

[0008] In some embodiments of the first aspect, after clustering the busy time periods of the time series data according to the second preset time period to determine the busy time intervals of the time series data, the method further includes: determining a dynamic threshold of the time series data when the maximum value of the time series data in the busy time intervals meets a preset condition, where the false alarm threshold is the first interval threshold and the missed alarm threshold is the second interval threshold.

[0009] In some embodiments of the first aspect, when the maximum value of the time series data in the busy time intervals does not meet the preset condition, calculating the dynamic threshold of the time series data according to the time series data includes: calculating the variance and kurtosis coefficient of the time series data according to the time series data in the busy time intervals; when the number of times the time series data is greater than a preset high load value is less than a first number threshold and the variance is less than a first variance threshold, calculating the dynamic threshold of the time series data by using an outlier judgment method; when the number of times the time series data is greater than a preset high load value is less than a first number threshold and the variance is greater than a first variance threshold, calculating the dynamic threshold of the time series data by using a normal distribution formula; when the number of times the time series data is greater than a preset high load value is greater than a first number threshold and the kurtosis coefficient is within a first kurtosis range, calculating the dynamic threshold of the time series data by using a normal distribution formula to perform a normal distribution judgment on the time series data greater than the preset high load value; when the number of times the time series data is greater than a preset high load value is greater than a first number threshold and the kurtosis coefficient is within a second kurtosis range, calculating the dynamic threshold of the time series data by using a normal distribution formula for the time series data after removing the time series data within a first preset range by using a peak shaving method and a skewness coefficient; when the number of times the time series data is greater than a preset high load value is greater than a first number threshold and the kurtosis coefficient is within a third kurtosis range, calculating the dynamic threshold of the time series data by using Chebyshev's inequality for the time series data after removing the time series data within a second preset range by using a peak shaving method and a skewness coefficient.

[0010] In some embodiments of the first aspect, the outlier judgment method includes: calculating an outlier upper bound and an outlier lower bound according to the time series data, where the outlier upper bound is the false alarm threshold and the outlier lower bound is the missed alarm threshold.

[0011] Second aspect, an embodiment of the present application provides a dynamic threshold calculation device, which includes: an acquisition module for acquiring time series data and the frequency of the time series data within a first preset time period; a generation module for generating a frequency diagram according to the frequency of the time series data, the frequency diagram including single hump feature information or multi-hump feature information; a determination module for determining the busy and idle time threshold based on the frequency diagram; a clustering module for clustering the busy time periods of the time series data according to a second preset time period to determine the busy time interval of the time series data, where the busy time period is the time period obtained when the time series data is greater than the busy and idle time threshold, and the second preset time period includes the first preset time period; a calculation module for calculating the dynamic threshold of the time series data according to the time series data when the maximum value of the time series data in the busy time interval does not meet the preset condition, where the dynamic threshold includes a false alarm threshold and a missed alarm threshold.

[0012] In some embodiments of the second aspect, the determination module specifically includes: a first threshold determination unit for determining the time series data corresponding to a preset range length as the busy and idle time threshold when the frequency diagram presents single hump feature information; a first trough determination unit for determining the number of troughs in the frequency diagram when the frequency diagram presents multi-hump feature information; a second threshold determination unit for determining the time series data corresponding to the trough as the busy and idle time threshold when the number of troughs is 1; a third threshold determination unit for determining the average value of the time series data corresponding to multiple troughs as the busy and idle time threshold when the number of troughs is greater than 1.

[0013] Third aspect, a dynamic threshold calculation device is provided, which includes: a memory for storing computer program instructions; a processor for reading and running the computer program instructions stored in the memory to execute the dynamic threshold calculation method provided in the first aspect or any optional implementation manner of the first aspect.

[0014] Fourth aspect, a computer storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the dynamic threshold calculation method provided in the first aspect or any optional implementation manner of the first aspect is implemented.

[0015] The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects:

[0016] In an embodiment of the present application, a frequency diagram of time-series data is generated by obtaining the time-series data within a first preset time period and the frequency of the time-series data. After generating the frequency diagram of the time-series data, the busy and idle time thresholds for dividing the time-series data are determined according to the frequency diagram, and then the busy time periods of the divided time-series data are clustered according to a second preset time period to obtain busy time intervals. When the largest time-series data in the busy time intervals does not meet the preset conditions, the dynamic threshold of the time-series data is calculated. Since the busy time periods are further confirmed to be busy time periods most of the time, the accurate division false alarm threshold and the dynamic threshold of missed alarms can be further calculated accordingly. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a flowchart of a method for calculating a dynamic threshold provided by an embodiment of the present application;

[0019] Figure 2 is a flowchart of another method for calculating a dynamic threshold provided in an embodiment of the present application;

[0020] Figure 3 is a flowchart of yet another method for calculating a dynamic threshold provided in an embodiment of the present application;

[0021] Figure 4 is a schematic structural diagram of a device for calculating a dynamic threshold provided by an embodiment of the present application;

[0022] Figure 5 is a schematic structural diagram of a device for calculating a dynamic threshold provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The features and exemplary embodiments of various aspects of the present application will be described in detail below. To make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application and not to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application.

[0024] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0025] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0026] In the existing threshold design scheme, in addition to the low efficiency easily caused by more manual interventions, there are also the following problems:

[0027] First, due to the lack of a scientific evaluation standard, business personnel can only judge based on the experience of the business support scenario, and cannot give a more specific busy time period, failing to achieve the purpose of alarm convergence, so that the range of the busy time period is too wide and false alarms are likely to occur. Second, the traditional thresholds are all set as static values. However, as time goes by, the usage of the service will definitely change, and the life cycle of the services hosted on the host is also changing. The usage frequency of the services in the old generation is relatively low, and the busy time period will also change. Fixed values cannot be used for evaluation. Therefore, the threshold settings in the existing technologies are not accurate enough, easily leading to too high frequencies of false alarms and missed alarms.

[0028] In summary, to solve the problem of too high frequencies of false alarms and missed alarms caused by inaccurate threshold settings in the existing dynamic threshold calculation methods, the embodiments of the present application provide a dynamic threshold calculation method, device, equipment and medium.

[0029] The technical solutions provided by the embodiments of the present application will be described below with reference to the accompanying drawings.

[0030] Figure 1 It is a schematic flow chart of a dynamic threshold calculation method provided by an embodiment of the present application.

[0031] As Figure 1 shown, this method can be implemented based on the above-mentioned dynamic threshold calculation device or part of the modules of the dynamic threshold calculation device. The dynamic threshold calculation method can include the following steps:

[0032] First, in step 110, obtain the time series data and the frequency of the time series data within the first preset time period.

[0033] Secondly, in step 120, generate a frequency diagram based on the frequency of the time series data.

[0034] Thirdly, in step 130, determine the busy and idle time thresholds for dividing the time series data based on the frequency diagram.

[0035] Then, in step 140, cluster the busy time periods of the time series data according to the second preset time period to determine the busy time intervals of the time series data.

[0036] Next, in step 150, when the maximum value of the time series data in the busy time interval does not meet the preset condition, calculate the dynamic threshold of the time series data.

[0037] Thus, by obtaining the time series data and the frequency of the time series data within the first preset time period, a frequency diagram of the time series data is generated. After generating the frequency diagram of the time series data, determine the busy and idle time thresholds for dividing the time series data based on the frequency diagram, and then cluster the busy time periods of the divided time series data according to the second preset time period to obtain the busy time intervals. When the maximum time series data in the busy time interval does not meet the preset condition, calculate the dynamic threshold of the time series data. Since the busy time periods are further confirmed to be busy time periods most of the time, the accurate division false alarm threshold and the dynamic threshold of missed alarms can be further calculated.

[0038] The above steps are described in detail below, as follows:

[0039] First, regarding step 110, in the embodiment of the present application, the first preset time period is a self-defined time period based on production needs, which can be one day, three months, or one year. The time series data can be index data with obvious fluctuations reflecting the performance of the system or device, such as resource occupancy data and resource capacity data. Exemplarily, it can be the utilization rate of the central processing unit (CPU) of the system or device. The obtained time series data may include the date and time when the time series data is collected and the value of the time series data. Specifically, as shown in Table 1 Time Series Data Table below, it should be noted that for the purpose of helping understanding, the time series data mentioned in this method simply represents a value, and the frequency of the time series data refers to the number of times the time series data appears after removing the date and time, that is, the number of times the value of the time series data appears.

[0040] Table 1 Time Series Data Table

[0041] Parameter Name Required Meaning Remarks Date_time Yes Date and time during time-series data acquisition (Y-M-D h:m) Date_value Yes Value of time-series data Actual data

[0042] In a specific example, it is necessary to obtain the CPU utilization rate of a certain device from February 3, 2021 to March 3, 2021. Among them, among the multiple CPU utilization rates, there is CPU utilization rate A (at 13:12 on February 14, 2021, a), and the occurrence frequency of the values of the multiple CPU utilization rates from February 3, 2021 to March 3, 2021.

[0043] Secondly, it involves step 120 of generating a frequency diagram of the time-series data within the first preset time period according to the frequency of the time-series data obtained within the first preset time period. The frequency diagram will have single-peak characteristic information or multi-peak characteristic information. In short, it is to determine whether there is one peak or multiple peaks in the frequency diagram. Exemplarily, the frequency diagram of the CPU utilization rate from February 3, 2021 to March 3, 2021 has one peak and is a frequency diagram with single-peak characteristic information.

[0044] Thirdly, it involves step 130 of determining the busy and idle time thresholds for dividing the time-series data according to the single-peak characteristic information or multi-peak characteristic information in the frequency diagram. Among them, the busy and idle time threshold refers to the threshold for dividing the time-series data into a busy time period and an idle time period. The time period corresponding to the time-series data greater than the busy and idle time threshold is the busy time period, and the time period corresponding to the time-series data less than the busy and idle time threshold is the idle time period.

[0045] Exemplarily, in the case where the frequency diagram has single-peak characteristic information, the time-series data corresponding to the preset range length is determined as the busy and idle time threshold. In the case where the frequency diagram has multi-peak characteristics, it is necessary to determine the number of troughs in the frequency diagram. In the case where there is only one trough in the frequency diagram, the time-series data corresponding to the trough is determined as the busy and idle time threshold. In the case where the frequency diagram has multiple troughs, the average value of the sum of the time-series data corresponding to each trough is determined as the busy and idle time threshold. Among them, the preset range is also a range that needs to be customized.

[0046] In a specific example, the frequency diagram of the CPU utilization rate from February 3, 2021 to March 3, 2021 has one peak, that is, the CPU utilization rate of this time period has single-peak characteristic information. Assuming that four data A, B, C, and D are collected in this time period and the preset range length is set to 75%, then multiply the total number of the collected CPU utilization rates by 75% to get the number 3. Then the third CPU utilization rate is the busy and idle time threshold. It is also possible to directly set the preset range length to the number 3, which is not limited here.

[0047] In another specific example, the frequency diagram of the CPU utilization rate from February 3, 2021 to March 3, 2021 has two peaks, that is, the frequency diagram has one trough. Assuming that the CPU utilization rate corresponding to the trough is data Z, then Z is determined as the busy and idle time threshold.

[0048] In yet another specific example, the frequency diagram of the CPU utilization rate from February 3, 2021 to March 3, 2021 has multiple peaks, that is, the frequency diagram has multiple valleys. Assuming there are three valleys, and the corresponding time series data are X1, X2, and X3 respectively, then the mean of the sum of the data X1, X2, and X3 is taken as the busy and idle threshold, that is, the busy and idle threshold is

[0049] Then, regarding step 140, after using the busy and idle threshold to divide the time series data to obtain the busy time period, cluster the obtained busy time period according to the second preset time period, and then obtain the busy time interval. Among them, the second preset time period is also a range defined according to needs, and the second preset time period includes the first preset time period. The busy time interval is further determined based on the principle of the minority obeying the majority for the busy time period to ensure that it is a busy time period for most of the time, so as to ensure the accuracy of subsequent calculation of the dynamic threshold.

[0050] Next, regarding step 150, after determining the busy time interval, it is judged whether the maximum value of the time series data in the busy time interval meets the preset condition. In the case where it does not meet the preset condition, calculate the dynamic threshold for the time series data according to the time series data within the busy time interval. Among them, the dynamic threshold includes a false alarm threshold and a missed alarm threshold. The false alarm threshold is to avoid a large number of invalid alarms in the static threshold, that is, to reduce the occurrence of invalid alarms and improve the work efficiency of personnel; the missed alarm threshold is to identify abnormal data with too low index data values. The preset condition is a condition preset based on experience, which can be a numerical value. Exemplarily, the preset condition can be set as the maximum value of the time series data is less than the numerical value A.

[0051] In addition, in the case where the maximum value of the time series data in the busy time interval meets the preset condition, directly determine the false alarm threshold in the dynamic threshold as the first interval threshold, and determine the missed alarm threshold in the dynamic threshold as the second interval threshold. Among them, the first interval threshold and the second interval threshold are thresholds for distinguishing false alarms or missed alarms preset based on production experience. Exemplarily, assuming the maximum value of the CPU utilization rate value range is 100%, the algorithm can set the limit false alarm threshold to 95%, that is, when the false alarm threshold calculated by the algorithm exceeds 95%, directly set the false alarm threshold to 95%, then the false alarm threshold is between 95% and 100%, and the CPU utilization rate data falling in this interval has the risk of false alarms; if the missed alarm threshold calculated by the algorithm is less than 0, take the current data minimum value as the missed alarm threshold. Assuming that at this time, the minimum value of the CPU utilization rate is 10%, then the missed alarm threshold is obtained as 0 to 10%, and the CPU utilization rate falling in this interval has the risk of missed alarms.

[0052] Based on this, step 140 involved above, as Figure 2 shown, may specifically include: step 1401 to step 1403.

[0053] Step 1401, obtain multiple sets of time series data in a second preset time period.

[0054] Here, it means obtaining multiple sets of time series data in units of the first preset time period in the second preset time period, where the set of time series data is a set containing the time series data of the lock brother obtained within the first preset time period. Exemplarily, assuming that the second preset time period is set to one month and the first preset time period is set to one day, then 30 sets of time series data are obtained in one month, and each set of time series data contains multiple time series data obtained on the corresponding day in one month.

[0055] Step 1402, count the occurrence frequency of the busy time periods in the multiple sets of time series data.

[0056] Here, after determining the busy time periods in the first preset time period according to the busy and idle time threshold, count the occurrence frequency of this busy time period in the multiple sets of time series data in the second preset time period in units of the first preset time period. Exemplarily, assuming that the busy time period determined in one day is 8:00 - 9:00, and in one month, that is, within 30 days, the time period 8:00 - 9:00 is a busy time period on 17 days, then the obtained occurrence frequency of the busy time period is 17.

[0057] Step 1403, determine the busy time interval of the time series data according to the occurrence frequency of the busy time period.

[0058] Specifically, based on the principle of the minority obeying the majority, or set a threshold for the occurrence frequency of the busy time period, and based on this, further determine the busy time period of the time series data. The method of determining the busy time period of the time series data is not limited to this and is not defined here.

[0059] In a specific example, when choosing to determine the busy time interval of the time series data based on the principle of the minority obeying the majority, if the time period 8:00 - 9:00 is a busy time period on 17 days, then this time period is an idle time period on 13 days, and at this time, this time period can be determined as the busy time interval.

[0060] In addition to the above steps, in another embodiment, step 150 above, as Figure 3 shown, may specifically include: step 1501 to step 1508.

[0061] Step 1501, calculate the variance and kurtosis coefficient of the time series data according to the time series data in the busy time interval.

[0062] Specifically, when the time-series data in the busy-hour time interval does not meet the preset conditions, calculate the variance and kurtosis coefficient of the time-series data based on the time-series data in the busy-hour time interval. Among them, the variance is used to describe the degree of deviation of the time-series data in the busy-hour time interval, and the kurtosis coefficient is a variable that measures whether the distribution of the time-series data in the busy-hour time interval is a normal distribution. The kurtosis coefficient of a normal distribution is 0.

[0063] The kurtosis coefficient can be calculated using the following formula.

[0064]

[0065] Among them, K here represents the kurtosis coefficient, n is a positive integer, indicating that there are n time-series data in the busy-hour time interval, and x i represents the i-th time-series data, where i is also a positive integer. represents the average value of the time-series data within the busy-hour time interval. represents the variance of the time-series data in the busy-hour time interval.

[0066] Step 1502, determine whether the number of times the time-series data is greater than the high-load value is less than the first number threshold. If so, execute Step 1503; if not, execute Step 1506, Step 1507, or Step 1508.

[0067] Among them, the high-load value represents the data when the system or device has entered the high-load state. Among them, the time-series data is corresponding to the high-load state, and the first number threshold is a threshold set based on requirements. Exemplarily, when the high-load value is set to 80%, and the first number threshold is set to 5, when the number of times the CPU utilization rate is greater than 80% is 7 times, which is greater than the first number threshold, then execute Step 1503. When the CPU utilization rate is always greater than 80%, then select the corresponding execution step.

[0068] Step 1503, determine whether the variance is less than the first variance threshold. If so, execute Step 1504; if not, execute Step 1505.

[0069] Among them, the first variance threshold is a threshold set in advance for judging the variance. Exemplarily, the first variance threshold can be set to 15. Specifically, when the time-series data within the busy-hour time interval is less than 15, execute Step 1504; otherwise, execute Step 1505.

[0070] Step 1504, calculate the dynamic threshold of the time-series data using the outlier judgment method.

[0071] Here, it is necessary to calculate the upper bound of the outlier and the lower bound of the outlier using the outlier judgment method. Among them, the upper bound of the outlier is the threshold for distinguishing false alarms, and the lower bound of the outlier is the threshold for distinguishing missed alarms.

[0072] The calculation formula for the outlier judgment method in statistics is as follows:

[0073] Upper bound of outlier = Upper quartile point + (Upper quartile - Lower quartile) * K (K = 3)

[0074] When certain conditions are met, substitute K = 3 into the above formula, and the calculated value is used as the false alarm threshold;

[0075] The calculation method for the lower bound of outliers in statistics is as follows:

[0076] Lower bound of outlier = Lower quartile + (Upper quartile - Lower quartile) * K (K = -3),

[0077] When certain conditions are met, substitute K = -3 into the above formula, and the calculated value is used as the missed alarm threshold.

[0078] Specifically, it is necessary to first sort the time series data within the busy hour time interval in a certain sorting order, and then divide the sorted time series data into four equal parts by three points, where each part contains 25% of the time series data. The value at the 25% position is called the lower quartile, and the value at the 75% position is called the upper quartile. Among them, there is no limit on the sorting method of the time series data.

[0079] Exemplarily, sort the time series data within the busy hour time interval from small to large, and obtain the time series data a and b corresponding to the upper quartile and the lower quartile respectively. Then, the false alarm threshold = b + (b - a) * 3; the missed alarm threshold = a + (b - a) * (-3).

[0080] Step 1505, calculate the dynamic threshold of the time series data using the normal distribution formula.

[0081] Here, when the amount of data is large, the data must conform to the normal distribution. However, when the amount of data is small, the algorithm determines whether the distribution is normal based on the kurtosis coefficient, and performs simple processing on the data using the skewness coefficient before subsequent calculations.

[0082] The normal distribution formula is as follows:

[0083]

[0084] Among them, μ represents the expected value, and σ represents the variance.

[0085] The specific calculation steps are as follows:

[0086] First, use the preset confidence level 1 - α to query the value in the standard normal distribution table

[0087] Secondly, calculate the mean value m and variance n of the timing data based on the timing data in the busy-hour time interval;

[0088] Then, substitute the mean value m and variance s into where n refers to the number of the timing data.

[0089] Finally, the false alarm threshold and missed alarm threshold are calculated according to the formula.

[0090] Step 1506, when the kurtosis coefficient is within the first kurtosis range, calculate the dynamic threshold of the timing data by using the normal distribution formula to judge the normal distribution of the timing data greater than the preset high load value.

[0091] Among them, the first kurtosis range is a preset range. When the calculated kurtosis coefficient is within the first kurtosis range, it is necessary to use the normal distribution formula to calculate the timing data greater than the high load value to obtain the timing data in this busy-hour time interval. Exemplarily, when the first kurtosis coefficient range can be set within the interval [-1, 1], which is not limited here. When the calculated kurtosis coefficient is within this interval, the normal distribution formula is used to calculate the timing data greater than the high load value. For the specific calculation process, please refer to step 1505, which will not be elaborated here.

[0092] Step 1507, when the kurtosis coefficient is within the second kurtosis range, calculate the dynamic threshold of the timing data by using the normal distribution formula for the timing data after removing the timing data within the first preset range by using the peak clipping method and skewness coefficient.

[0093] Here, the second kurtosis range is a preset range. When the calculated kurtosis coefficient is within the second kurtosis range, calculate the dynamic threshold of the timing data by using the normal distribution formula for the timing data obtained after removing the timing data within the first preset range by using the peak clipping method and skewness coefficient. Among them, the second preset range is a range set based on experience. Exemplarily, the second preset range can be set within the interval (-2, 1) or (1, 2), which is not limited here. Assume that the second preset range is 10% at this time. The peak clipping method means that after sorting the timing data in the busy-hour time interval from small to large, if the skewness coefficient is greater than 0 at this time, the largest 10% of the numbers are removed, otherwise the smallest 10% are removed.

[0094] Specifically, the skewness coefficient is used to measure the deviation of the data average value relative to the overall data, and the skewness coefficient of the normal distribution is equal to 0; if the skewness coefficient is greater than 0, it means that the average value of the data is larger than the median of the data, and vice versa, the average value is smaller than the median. The specific formula is as follows:

[0095]

[0096] Wherein, S represents the skewness coefficient, n is a positive integer, indicating that there are n time series data in this busy hour time interval, and x i represents the i-th time series data, where i is also a positive integer. represents the average value of the time series data within this busy hour time interval. represents the variance of the time series data in this busy hour time interval.

[0097] Step 1508, when the kurtosis coefficient is within the third kurtosis range, calculate the dynamic threshold of the time series data by using the Chebyshev inequality for the time series data after removing the time series data in the second preset range by using the peak shaving method and the skewness coefficient.

[0098] Here, the third kurtosis range is a preset range. When the calculated kurtosis coefficient is within the third kurtosis range, the dynamic threshold of the time series data can be calculated by using the Chebyshev inequality for the time series data after removing according to the second preset range by using the peak shaving method and the skewness coefficient. Among them, the second preset range is greater than the first preset range. Exemplarily, the third kurtosis coefficient range is set to the kurtosis coefficient less than or equal to -2 or greater than or equal to 2. There is no excessive limitation here.

[0099] The formula of the Chebyshev inequality is as follows:

[0100]

[0101] Wherein, ε can be any value, μ is the expected value of the time series data, and σ is the standard deviation of the time series data.

[0102] This formula means that in any time series data set, the proportion (or part) located within m standard deviations of its average value is always at least where m is any positive number greater than 1. This inequality is used to process non-normal distribution data. Exemplarily, it is assumed that according to the formula: among all time series data, at least 24 / 25 (or 96%) of the time series data is located within 5 standard deviations of the average value. So we take the right boundary of the 5 standard deviation range as the false alarm threshold; and take the left boundary of the 5 standard deviation range as the missed alarm threshold.

[0103] In this way, through the organic integration of multiple algorithms, the automation and intelligence of threshold analysis are realized, the manual participation is reduced, the work efficiency of operation and maintenance is improved, and the busy and idle time zones of the service can be intelligently distinguished and based on this, dynamic threshold analysis is carried out, the false alarm frequency is reduced, and the frequency of missed alarm discovery is increased.

[0104] Based on the same inventive concept, the embodiment of the present application also provides a prediction device for the virtual machine hosting state. Specifically combinedFigure 4 will be described.

[0105] Figure 4 It is a schematic structural diagram of a dynamic threshold calculation device provided by an embodiment of the present application.

[0106] As Figure 4 shown, the dynamic threshold calculation device may include: an acquisition module 410, a generation module 420, a determination module 430, a clustering module 440, and a calculation module 450.

[0107] Among them, the acquisition module 410 is used to acquire the time series data and the frequency of the time series data within the first preset time period;

[0108] The generation module 420 is used to generate a frequency diagram according to the frequency of the time series data, and the frequency diagram includes single hump feature information or multi-hump feature information;

[0109] The determination module 430 is used to determine the busy and idle time threshold for dividing the time series data based on the frequency diagram;

[0110] The clustering module 440 is used to cluster the busy time periods of the time series data according to the second preset time period, and determine the busy time interval of the time series data, where the busy time period is the time period obtained when the time series data is greater than the busy and idle time threshold, and the second preset time period includes the first preset time period;

[0111] The calculation module 450 is further used to calculate the dynamic threshold of the time series data according to the time series data when the maximum value of the time series data in the busy time interval does not meet the preset condition, where the dynamic threshold includes a false alarm threshold and a missed alarm threshold.

[0112] In some embodiments, the determination unit 430 specifically includes:

[0113] The first determination threshold unit is used to determine the time series data corresponding to the preset range length as the busy and idle time threshold when the frequency diagram presents single hump feature information;

[0114] The first determination trough unit is used to determine the number of troughs in the frequency diagram when the frequency diagram presents multi-hump feature information;

[0115] The second determination threshold unit is used to determine the time series data corresponding to the trough as the busy and idle time threshold when the number of troughs is 1;

[0116] The third determination threshold unit is used to determine the average value of the time series data corresponding to multiple troughs as the busy and idle time threshold when the number of troughs is greater than 1.

[0117] In some embodiments, the clustering module 440 specifically includes:

[0118] An acquisition unit, configured to acquire a plurality of time-series data sets in a second preset time period;

[0119] A statistics unit, configured to count the occurrence frequency of busy time periods in a plurality of time-series data sets;

[0120] A determination time interval unit, configured to determine the busy time interval of the time-series data according to the occurrence frequency of the busy time periods.

[0121] In some embodiments, the determination module is further configured to determine a dynamic threshold of the time-series data when the maximum value of the time-series data in the busy time interval meets a preset condition, where the false alarm threshold is a first interval threshold and the missed alarm threshold is a second interval threshold.

[0122] In some embodiments, the calculation module 450 specifically includes:

[0123] A first calculation unit, configured to calculate the variance and kurtosis coefficient of the time-series data according to the time-series data in the busy time interval;

[0124] A second calculation unit, configured to calculate the dynamic threshold of the time-series data by using an outlier judgment method when the number of times the time-series data is greater than a preset high load value is less than a first number threshold and the variance is less than a first variance threshold;

[0125] A third calculation unit, configured to calculate the dynamic threshold of the time-series data by using a normal distribution formula when the number of times the time-series data is greater than a preset high load value is less than a first number threshold and the variance is greater than a first variance threshold;

[0126] A fourth calculation unit, configured to calculate the dynamic threshold of the time-series data by performing a normal distribution judgment on the time-series data greater than the preset high load value by using a normal distribution formula when the number of times the time-series data is greater than a preset high load value is greater than a first number threshold and the kurtosis coefficient is within a first kurtosis range;

[0127] A fifth calculation unit, configured to calculate the dynamic threshold of the time-series data by using a normal distribution formula for the time-series data after removing the time-series data within a first preset range by using a peak shaving method and a skewness coefficient when the number of times the time-series data is greater than a preset high load value is greater than a first number threshold and the kurtosis coefficient is within a second kurtosis range;

[0128] A sixth calculation unit, configured to calculate the dynamic threshold of the time-series data by using Chebyshev's inequality for the time-series data after removing the time-series data within a second preset range by using a peak shaving method and a skewness coefficient when the number of times the time-series data is greater than a preset high load value is greater than a first number threshold and the kurtosis coefficient is within a third kurtosis range.

[0129] In some embodiments, the second calculation unit specifically includes:

[0130] The first calculation subunit calculates the upper bound and lower bound of outliers based on the timing data, where the upper bound of outliers is the false alarm threshold, and the lower bound of outliers is the missed alarm threshold.

[0131] In the embodiment of the present application, a frequency diagram of the timing data is generated by obtaining the timing data within the first preset time period and the frequency of the timing data. After generating the frequency diagram of the timing data, the busy and idle time thresholds for dividing the timing data are determined according to the frequency diagram, and then the busy time periods of the divided timing data are clustered according to the second preset time period to obtain the busy time intervals. When the largest timing data in the busy time intervals does not meet the preset conditions, the dynamic threshold of the timing data is calculated. Since the busy time periods are further confirmed to be busy time periods most of the time, the accurate false alarm threshold and the dynamic threshold of missed alarms can be further calculated.

[0132] Figure 5 It is a schematic structural diagram of an updated test object device provided by the embodiment of the present application.

[0133] As Figure 5 shown, the dynamic threshold calculation device 500 in this embodiment includes an input device 501, an input interface 502, a central processing unit 503, a memory 504, an output interface 505, and an output device 506. Among them, the input interface 502, the central processing unit 503, the memory 504, and the output interface 505 are connected to each other through a bus 510. The input device 501 and the output device 506 are respectively connected to the bus 510 through the input interface 502 and the output interface 505, and then connected to other components of the information acquisition device 500.

[0134] Specifically, the input device 501 receives external input information and transmits the input information to the central processing unit 503 through the input interface 502; the central processing unit 503 processes the input information based on the computer-executable instructions stored in the memory 504 to generate output information, temporarily or permanently stores the output information in the memory 504, and then transmits the output information to the output device 506 through the output interface 505; the output device 506 outputs the output information to the outside of the information acquisition device 500 for user use.

[0135] In one embodiment, Figure 5 the dynamic threshold calculation device 500 shown includes: a memory 504 for storing programs; a processor 503 for running the programs stored in the memory to execute the Figures 1 - 3 method of any of the embodiments shown in the present application.

[0136] The embodiments of the present application further provide a computer-readable storage medium, on which computer program instructions are stored; when the computer program instructions are executed by a processor, the methods provided by the embodiments of the present application are implemented. Figures 1 - 3 The method of any of the illustrated embodiments.

[0137] It should be clear that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0138] The functional blocks shown in the above structure block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memory (ROM), flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0139] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0140] As described above, this is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.

Claims

1. A dynamic threshold calculation method, characterized in that, Including: Obtain the time-series data within a first preset time period and the frequency of the time-series data. The time-series data includes at least one of resource occupancy data and resource capacity data. The resource occupancy data and the resource capacity data are index data used to reflect the performance of a system or device, and the frequency of the time-series data is the number of times the value of the time-series data appears. Generate a frequency diagram based on the frequency of the time-series data. The frequency diagram includes single-hump feature information or multi-hump feature information. Determine the busy-idle threshold for dividing the time-series data based on the frequency diagram. Cluster the busy time periods of the time-series data according to a second preset time period to determine the busy time interval of the time-series data. Here, the busy time period is the time period obtained when the time-series data is greater than the busy-idle threshold, and the second preset time period includes the first preset time period. In the case where the maximum value of the time-series data in the busy time interval does not meet the preset condition, calculate the dynamic threshold of the time-series data according to the time-series data, where the dynamic threshold includes a false alarm threshold and a missed alarm threshold. The calculating the dynamic threshold of the time-series data according to the time-series data in the case where the maximum value of the time-series data in the busy time interval does not meet the preset condition includes: Calculate the variance and kurtosis coefficient of the time-series data according to the time-series data in the busy time interval. In the case where the number of times the time-series data is greater than a preset high-load value is less than a first number threshold and the variance is less than a first variance threshold, calculate the dynamic threshold of the time-series data using an outlier judgment method. In the case where the number of times the time-series data is greater than the preset high-load value is less than the first number threshold and the variance is greater than the first variance threshold, calculate the dynamic threshold of the time-series data using a normal distribution formula. In the case where the number of times the time-series data is greater than the preset high-load value is greater than the first number threshold and the kurtosis coefficient is within a first kurtosis range, calculate the dynamic threshold of the time-series data by performing a normal distribution judgment on the time-series data greater than the preset high-load value using a normal distribution formula. In the case where the number of times the time-series data is greater than the preset high-load value is greater than the first number threshold and the kurtosis coefficient is within a second kurtosis range, calculate the dynamic threshold of the time-series data according to the time-series data after removing the time-series data within a first preset range using a peak clipping method and a skewness coefficient using a normal distribution formula. In the case where the number of times the time-series data is greater than the preset high-load value is greater than the first number threshold and the kurtosis coefficient is within a third kurtosis range, calculate the dynamic threshold of the time-series data according to the time-series data after removing the time-series data within a second preset range using a peak clipping method and the skewness coefficient using Chebyshev's inequality.

2. The method according to claim 1, characterized in that, The determining the busy-idle threshold for dividing the time-series data based on the frequency diagram includes: In the case where the frequency diagram presents single-hump feature information, determine the time-series data corresponding to a preset range length as the busy-idle threshold. When the frequency diagram presents multi-hump characteristic information, determine the number of troughs in the frequency diagram; When the number of troughs is 1, determine the time-series data corresponding to the trough as the busy-idle threshold; When the number of troughs is greater than 1, determine the mean of the time-series data corresponding to multiple troughs as the busy-idle threshold.

3. The method according to claim 1, characterized in that, The clustering of the busy time periods of the time-series data according to a second preset time period to determine the busy time interval of the time-series data includes: Obtain multiple time-series data sets in the second preset time period; Count the occurrence frequency of the busy time periods in the multiple time-series data sets; Determine the busy time interval of the time-series data according to the occurrence frequency of the busy time periods.

4. The method according to claim 1, characterized in that, After clustering the busy time periods of the time-series data according to a second preset time period to determine the busy time interval of the time-series data, the method further includes: When the maximum value of the time-series data in the busy time interval meets a preset condition, determine the dynamic threshold of the time-series data, where the false alarm threshold is the first interval threshold and the missed alarm threshold is the second interval threshold.

5. The method according to claim 1, characterized in that, The outlier judgment method includes: Calculate the upper bound of the outlier and the lower bound of the outlier according to the time-series data, where the upper bound of the outlier is the false alarm threshold and the lower bound of the outlier is the missed alarm threshold.

6. A dynamic threshold calculation device, characterized in that, The device includes: An acquisition module, configured to acquire time-series data and the frequency of the time-series data within a first preset time period, where the time-series data includes at least one of resource occupancy data and resource capacity data, the resource occupancy data and the resource capacity data are index data for reflecting the performance of a system or a device, and the frequency of the time-series data is the number of times the value of the time-series data appears; A generation module, configured to generate a frequency diagram according to the frequency of the time-series data, where the frequency diagram includes single-hump characteristic information or multi-hump characteristic information; A determination module, configured to determine a busy-idle threshold for dividing the time-series data based on the frequency diagram; A clustering module, configured to cluster the busy time periods of the time-series data according to a second preset time period to determine the busy time interval of the time-series data, where the busy time period is a time period obtained when the time-series data is greater than the busy-idle threshold, and the second preset time period includes the first preset time period; A calculation module is further configured to calculate the dynamic threshold of the time-series data according to the time-series data when the maximum value of the time-series data in the busy time interval does not meet the preset condition, where the dynamic threshold includes a false alarm threshold and a missed alarm threshold; The calculation module specifically includes: A first calculation unit, configured to calculate the variance and kurtosis coefficient of the time-series data according to the time-series data in the busy time interval; A second calculation unit, configured to calculate the dynamic threshold of the time-series data by using an outlier judgment method when the number of times the time-series data is greater than a preset high load value is less than a first number threshold and the variance is less than a first variance threshold; A third calculation unit, configured to calculate a dynamic threshold of the time series data by using a normal distribution formula when the number of times the time series data is greater than the preset high-load value is less than the first number threshold and the variance is greater than the first variance threshold; A fourth calculation unit, configured to calculate a dynamic threshold of the time series data by performing a normal distribution judgment on the time series data greater than the preset high-load value by using a normal distribution formula when the number of times the time series data is greater than the preset high-load value is greater than the first number threshold and the kurtosis coefficient is within a first kurtosis range; A fifth calculation unit, configured to calculate a dynamic threshold of the time series data by using a normal distribution formula for the time series data after removing the time series data within a first preset range by using a peak clipping method and a skewness coefficient when the number of times the time series data is greater than the preset high-load value is greater than the first number threshold and the kurtosis coefficient is within a second kurtosis range; A sixth calculation unit, configured to calculate a dynamic threshold of the time series data by using the Chebyshev inequality for the time series data after removing the time series data within a second preset range by using a peak clipping method and the skewness coefficient when the number of times the time series data is greater than the preset high-load value is greater than the first number threshold and the kurtosis coefficient is within a third kurtosis range.

7. The device according to claim 6, wherein, The determining module specifically includes: A first determining threshold unit, configured to determine the time series data corresponding to a preset range length as the busy-idle time threshold when the frequency diagram presents single-hump feature information; A first determining valley unit, configured to determine the number of valleys in the frequency diagram when the frequency diagram presents multi-hump feature information; A second determining threshold unit, configured to determine the time series data corresponding to the valley as the busy-idle time threshold when the number of valleys is 1; A third determining threshold unit, configured to determine the mean value of the time series data corresponding to multiple valleys as the busy-idle time threshold when the number of valleys is greater than 1.

8. A dynamic threshold calculation device, wherein, The device includes: a processor and a memory storing computer program instructions; The processor reads and executes the computer program instructions to implement the dynamic threshold calculation method according to any one of claims 1-5.

9. A computer storage medium, wherein, Computer program instructions are stored on the computer storage medium, and when the computer program instructions are executed by the processor, the dynamic threshold calculation method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Equipment control method and related product

    CN108509804A

  • Cloud service alarm method and device

    CN110727560A