Manage network service level thresholds
By introducing automatic service level threshold adjustment technology based on statistical analysis and machine learning in the service monitoring system, the problems of complexity and high maintenance costs of service measurement threshold management in the prior art are solved, and more efficient and reliable system performance management is achieved.
Patent Information
- Application Number
- CN202410459799.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-29
- Filing Date
- 2024-04-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-04-17
AI Technical Summary
The prior art is difficult to effectively manage and maintain service level thresholds corresponding to a large number of service metrics, especially in large and complex environments, resulting in high maintenance costs and the thresholds may become invalid.
Using a system based on statistical analysis and machine learning, service level thresholds are automatically adjusted by automatically adjusting service level thresholds and using statistical models and machine learning models to analyze service metric data to automatically set and adjust service level thresholds.
Reduces the burden of manually setting and maintaining service level thresholds by SRE or other users, improves system performance, provides traceability of service level threshold changes, and allows dynamic adjustment of thresholds during specific tests.
Smart Images

Figure CN119011424B_ABST
Abstract
Description
Background Art
[0001] Some computing systems have been designed to provide one or more services over a computer network. For example, these services can include providing hardware resources (e.g., processing resources, storage resources, network resources, etc.), application resources, management resources (e.g., resource management resources), and / or any suitable combination of other types of services provided over a computer network. A service provider can provide some or all of these services to multiple customers. The computing systems implemented or otherwise used to provide these services can include distributed computing systems, such as cloud systems, multi-cloud systems, hybrid cloud systems, or any other suitable type of distributed computing system. Brief Description of the Drawings
[0002] To more fully understand the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which:
[0003] Figure 1 illustrates an example system for managing service level thresholds in accordance with certain embodiments;
[0004] Figure 2 illustrates details of an example service monitoring system in accordance with certain embodiments;
[0005] Figure 3 illustrates an example method for managing service level thresholds in accordance with certain embodiments;
[0006] Figure 4 illustrates an example method for managing service level thresholds in accordance with certain embodiments;
[0007] Figure 5 illustrates an example method for freezing and thawing service level thresholds in accordance with certain embodiments;
[0008] Figure 6 illustrates an example method for managing service level thresholds in accordance with certain embodiments; and
[0009] Figure 7 illustrates a block diagram of an example computing device in accordance with certain embodiments. Detailed Description
[0010] A service provider that provides computerized services over a communication network can monitor the services to attempt to ensure that the services are provided to customers at a satisfactory level. To this end, the system can track one or more service metrics related to the performance of the services. At least some of those service metrics may somehow reflect the customers' service experience.
[0011] A Site Reliability Engineer (SRE) can track service metrics related to a customer's interaction with a service (which can be referred to as Service Level Indicators (SLIs)), and create values for service level thresholds (which can be referred to as Service Level Objectives (SLOs)) to distinguish acceptable values from unacceptable service metric values for these service metrics. Service level thresholds can be used to identify whether a customer is likely to be satisfied with the service. A service that breaches (e.g., exceeds) a service level threshold can be considered to have failed to provide an acceptable service to the customer, and actions may be appropriate. Typically, when a service metric value for a service metric approaches or crosses the value of the service level threshold for that service metric, an alert can be generated to the service owner / operator, potentially prompting actions to fix the problem.
[0012] Service level thresholds can be set manually by an SRE or other users based on an understanding of the tracked service metrics. However, an SRE may be encouraged to measure a large number of service level metrics - that is, identify a large number of service metrics - so defining and maintaining service level thresholds at appropriate values for each service metric can be challenging, especially for large complex environments. That is, as the number of service metrics increases, maintaining appropriate values for the corresponding service level thresholds can involve a significant amount of effort.
[0013] Additionally or alternatively, systems change over time, and as the underlying system changes, existing service level threshold definitions may become invalid or at least not optimal. For example, the system may be changed to support new features or improve scalability. Changes may also result from mandatory changes to third - party / open - source components. At least in part due to these system changes, it may be appropriate to modify service level thresholds, but it is not always clear which service level thresholds should be changed or how they should be changed.
[0014] Certain embodiments provide improved techniques for maintaining service level thresholds. For example, certain embodiments provide techniques for automatic service level threshold adjustment. Certain embodiments provide systems based on statistical analysis that use concepts from data engineering and machine learning (sometimes abbreviated as "ML" throughout this disclosure and the figures) to derive and modify service level thresholds based on observed behavior and knowledge of when the system is considered to be misbehaving (e.g., support calls, other service level thresholds being breached, infrastructure failure events, etc.).
[0015] Some embodiments may manage service level thresholds in one or more phases. For example, some embodiments may use a service level threshold auto - tuning phase to manage service level thresholds, during which the values for the service level thresholds are automatically set (e.g., adjusted) based on an analysis performed using one or more statistical models for managing service level thresholds. In some embodiments, the results of setting the values for the service level thresholds for a service metric may be stored, including potential changes to the values of the service level thresholds, the timing of changes to the service level thresholds, the circumstances leading to changes in the service level thresholds, and / or any other suitable information. In some embodiments, the results of setting the values for the service level thresholds for a service metric may be stored as time - series data.
[0016] As another example, some embodiments may use a machine - learning analysis phase to manage service level thresholds, during which one or more machine - learning models are used to analyze changes to the values of the service level thresholds to identify patterns in the changes to the values of the service level thresholds. In some embodiments, one or more machine - learning models are used to evaluate time - series data to identify patterns in the changes to the values of the service level thresholds.
[0017] Some embodiments may provide improved values for service level thresholds, which can improve overall system performance. Some embodiments are capable of automatically tuning service level thresholds based on statistical models (which provide more reliable and up - to - date information than possible with manual user intervention), rather than relying solely on sporadic manual intervention by SREs or other users to modify service level thresholds. In some embodiments, by automatically tuning the values for one or more service level thresholds, the burden on SREs or other users of manually setting the service level thresholds (e.g., SLOs) for service metrics (e.g., SLI) can be reduced or eliminated. This may be particularly useful for service metrics for which the impact on the end - customer experience is currently uncertain. Some embodiments can provide traceability of changes to service level thresholds, such as by maintaining a log of changes to the values of service level thresholds.
[0018] Some embodiments may allow the values of one or more service level thresholds to “float” during specific tests designed to induce poor service behavior and allow the values of those one or more service level thresholds to be set thereby via system testing (e.g., via an auto - tuning mechanism). Some embodiments may allow SRE engineers or other users to focus on a specific subset of important service level thresholds for manual intervention and let the system set and adjust (via an automated auto - tuning mechanism) other service level thresholds. If desired, some embodiments may allow service level thresholds to be manually frozen at any time. Some embodiments may provide a time - series analysis of the values of service level thresholds adjusted using an auto - tuning mechanism, which can provide insights and / or predictions.
[0019] Figure 1 FIG. illustrates an example system 100 for managing service level thresholds in accordance with certain embodiments. In the illustrated example, system 100 includes a client system 102, a service provider system 104, a communication network 106, and a service monitoring system 108. Although system 100 is illustrated and described as including specific components, system 100 may include any suitable components according to specific needs.
[0020] Client system 102 may include any suitable type and number of electronic processing devices, including a single processing device, multiple processing devices, multiple processing devices communicating via a computer network, an enterprise network, or any other suitable type of processing device(s) in any suitable arrangement, some of which may overlap in type. In certain embodiments, client system 102 may be one of multiple client systems 102 that interact with service provider system 104.
[0021] Service provider system 104 may include any suitable type and number of electronic processing devices, including a single processing device, multiple processing devices, multiple processing devices communicating via a computer network, an enterprise network, or any other suitable type of processing device(s) in any suitable arrangement, some of which may overlap in type.
[0022] Service provider system 104 may provide one or more computerized services 110 (referred to as "services 110" etc. for simplicity in the remainder of this disclosure) to client system 102 via communication network 106. Services 110 provided by service provider system 104 may include any type of electronic service that may be provided to client system 102 via communication network 106. For example, services 110 may include one or more of cloud services, multi-cloud services, hybrid cloud services, web services, web hosting services, data storage services, data processing services, high-performance computing services, or any other suitable type of electronic service(s), some of which may overlap in type.
[0023] The communication network 106 facilitates wireless and / or wired communication. The communication network 106 can transfer, for example, IP packets, Frame Relay frames, Asynchronous Transfer Mode (ATM) cells, voice, video, data, and other suitable information between network addresses. The communication network 106 can include any suitable combination of the following: one or more local area networks (LANs), radio access networks (RANs), metropolitan area networks (MANs), wide area networks (WANs), mobile networks (e.g., using WiMax (802.16), WiFi (802.11), 3G, 4G, 5G, or any other suitable wireless technology in any suitable combination), all or a portion of the global computer network known as the Internet, and / or one or more any other communication systems at one or more locations, where any one can be any suitable combination of wireless and wired.
[0024] The system 100 includes a service monitoring system 108, which can monitor the performance of the service 110. The service monitoring system 108 can obtain service metric data 112 related to the service 110. The service monitoring system 108 can obtain the service metric data 112 from one or more agents deployed throughout the system 100 (e.g., at the client system 102, the service provider system 104, and / or the network 106) and / or in any other suitable manner. For example, the service monitoring system 108 can poll the agents for the service metric data 112 at regular or irregular time intervals. Additionally or alternatively, the agents can automatically transmit the service metric data 112 to the service monitoring system 108 at regular or irregular time intervals. Such agents can be implemented using any suitable combination of hardware, firmware, and software. Additionally, although described as "agents", the present disclosure contemplates the service monitoring system 108 obtaining the service metric data 112 in any suitable way or combination of ways.
[0025] The service metric data 112 can include one or more of the following: service metric values for one or more service metrics, timestamp information, and / or any other suitable information. A service metric can be a parameter that captures an aspect of the performance of the service 110. For example, a service metric can be considered a service level indicator (SLI). In some embodiments, the service metric measures aspects of the interaction of a user (e.g., a user of the client system 102) with the service 110 to determine whether the user may be experiencing or about to experience a problem with the service 110.
[0026] Some specific examples of service metrics can include the number of times an application programming interface (API) call results in a 500 error, latency or response time (e.g., how long it takes for a user to receive a response from the service provider system 104, such as the latency to load a user interface (e.g., a web page), the latency of an API call, the amount of time it takes for a virtual machine (VM) to start up, or the amount of time it takes to create a Kubernetes cluster), error rate or quality, uptime, availability, and / or the number of times a running virtual machine stops and is restarted due to a maintenance operation or other service interruption. Although specific examples have been described, service metrics can include any suitable service metrics.
[0027] In some embodiments, the service metric is measurable as a numerical value, which means that in such embodiments, the service metric value is a number. The service metric can have any unit of a type suitable for the measurement being made of the service metric. For example, the unit of the service metric can be time, percentage, or any other suitable type of measurement.
[0028] To the extent that the service metric data 112 includes timestamp information, the timestamp information can indicate the time at which the measurement was made (e.g., at which the associated one or more service metric values were generated), the time at which the one or more service metric values were received by the service monitoring system 108 (e.g., generated by the associated measurement), and / or any other suitable time information.
[0029] The service monitoring system 108 can maintain corresponding service level thresholds for some or all of the service metrics. The service level threshold for a corresponding service metric can establish a goal for the corresponding service metric. For example, the service level threshold can be considered a service level objective (SLO). In some embodiments, the service monitoring system maintains a separate service level threshold for each service metric, but the values of the different service level thresholds can be the same or different.
[0030] The service monitoring system 108 can compare the service metric value for a service metric with the service level threshold for that service metric to determine whether the service metric value for that service metric breaches the service level threshold for that service metric.
[0031] The service monitoring system 108 can determine, in any suitable manner, whether a service metric value for a service metric breaches a service level threshold for that service metric. For example, the service monitoring system 108 can compare an individual service metric value, multiple service metric values (e.g., over a particular time period and / or a particular number of service metric values), or a value of multiple service metric values (e.g., an average value) with the service level threshold for that service metric to determine whether the service metric value for that service metric breaches the service level threshold for that service metric. As another example, one or more service metric values for a service metric breaching the service level threshold for that service metric can include one or more service metric values exceeding (e.g., being greater than, or greater than or equal to, depending on the implementation) the current value of the service level threshold, one or more service metric values being below (e.g., being less than, or less than or equal to, depending on the implementation) the current value of the service level threshold, or any other suitable implementation.
[0032] In some embodiments, if the service monitoring system 108 determines that the value of a service level metric breaches the service level threshold for that service level metric, then it can be considered that the service provider system 104 has a delivery failure, or there may be a risk of being unable to deliver the service 110 to the client system 102 at the desired level, at least for the aspect of the service 110 measured by that service metric.
[0033] As described above, it may be appropriate to adjust the value of the service threshold over time. Such adjustments may be due to changes in the service provider system 104, changes in the client system 102 (or changes to other client systems 102, such as when the service provider system 104 serves multiple client systems 102), changes to the communication network 106, the time at which the service provider system 104 is providing the service 110 (e.g., time of day, time of month, time of year), or for any other suitable reason or combination of reasons. These changes may affect what values may be considered "normal" for the service metric values of the service metric, and thus may warrant a change in the appropriate value of the service level threshold for that service metric. Of course, multiple service metrics and service level thresholds may be affected.
[0034] The service monitoring system 108 can include a service level threshold analysis engine 114, which can be configured to periodically evaluate the service level threshold to determine how to set the corresponding value of the service level threshold. For example, the service level threshold analysis engine 114 can continuously evaluate the service level threshold when receiving a new service metric value for a service metric corresponding to a particular service level threshold (e.g., each new service metric value, a particular number of service metric values, etc.), or at any other suitable interval.
[0035] The composition for setting a service level threshold can vary in different scenarios. For example, if no service level threshold has been established for a specific service metric, in some scenarios, setting the service level threshold can include establishing a service level threshold (e.g., defining variables for the service level threshold) and setting the service level threshold to an initial value. As another example, if a service level threshold has been established for a specific service metric, setting the service level threshold can include keeping the current value of the service level threshold unchanged, adjusting the current value of the service level threshold by a predefined adjustment amount, or taking other appropriate actions regarding the service level threshold.
[0036] The service level threshold analysis engine 114 can manage the service level threshold in one or more phases. These phases can run simultaneously or at different times, depending on the situation.
[0037] For example, the service level threshold analysis engine 114 can use the service level threshold automatic adjustment phase to manage the service level threshold, during which the service level threshold analysis engine 114 can automatically set (e.g., adjust) the value of the service level threshold based on the analysis performed using a statistical model for managing the service level threshold. In some embodiments, as part of determining how to set the service level threshold, the service level threshold analysis engine 114 can analyze service metric values for a service metric according to one or more statistical models to identify abnormal service metric values. The statistical model can include a predicted distribution of service metric values for the service metric and a range of normal values within the predicted distribution of service metric values for the service metric. An outlier can be a service metric value that falls outside the range of normal values. Refer to Figures 2 - 5 for additional details.
[0038] In some embodiments, the service level threshold analysis engine 114 can store the results of setting the service level threshold for a service metric. For example, the service level threshold analysis engine 114 can store historical values of the service level threshold, which can include changes to the value of the service level threshold, the timing of changes to the service level threshold, the situations that led to changes in the service level threshold, and / or any other appropriate information. In some embodiments, the service level threshold analysis engine 114 can store the results of setting the value of the service level threshold for a service metric as time series data.
[0039] As another example, the service level threshold analysis engine 114 can use a machine learning analysis phase to manage service level thresholds, during which the service level threshold analysis engine 114 can use one or more machine learning models to analyze changes to the value of the service level threshold to identify patterns in the changes to the value of the service level threshold. In certain embodiments, the service level threshold analysis engine 114 can use one or more machine learning models to evaluate time series data to identify patterns in the changes to the value of the service level threshold. Refer to Figures 2 - 3 and Figure 6 describe additional details.
[0040] The service monitoring system 108 and / or the service level threshold analysis engine 114 can be implemented using any suitable combination of hardware, firmware, and software. Although illustrated separately from the client system 102 and the service provider system 104, the service monitoring system 108 can be implemented as part of the client system 102 and / or the service provider system 104 or separate from the client system 102 and / or the service provider system 104.
[0041] Figure 2 Details of an example service monitoring system 108 according to certain embodiments are illustrated. Although a particular implementation of the service monitoring system 108 is illustrated and described, the present disclosure contemplates any suitable implementation of the service monitoring system 108.
[0042] The service monitoring system 108 can include a processor 200, a network interface 202, and a memory 204. Although described in the singular for convenience, the service monitoring system 108 can include one or more processors 200, one or more network interfaces 202, and one or more memories 204.
[0043] The processor 200 can be any component or collection of components suitable for performing computing and / or other processing-related tasks. The processor 200 can be, for example, a microprocessor, a microcontroller, a control circuit, a digital signal processor, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SoC), a graphics processing unit (GPU), or a combination thereof. The processor 200 can include any suitable number of processors, or multiple processors can together form a single processor 200. The processor 200 can operate alone or with other components of the system 100 (see Figure 1 ) to provide some or all of the functionality of the service monitoring system 108 described herein.
[0044] The network interface 202 represents that information can be received from a communication network, transmitted through the communication network, perform appropriate processing of the information, and communicate with (e.g., Figure 1Any suitable computer element that communicates with other components of the system 100, or any combination of the above. The network interface 202 represents any port or connection, real or virtual, including any suitable combination of hardware, firmware, and software, including protocol conversion and data processing capabilities, to communicate over a LAN, WAN, or other communication system that permits information exchange with the devices of the system 100. The network interface 202 can facilitate wireless and / or wired communication.
[0045] The memory 204 can include any suitable combination of volatile memory, non-volatile memory, and / or its virtualization. For example, the memory can include any suitable combination of magnetic media, optical media, random access memory (RAM), read-only memory (ROM), removable media, and / or any other suitable memory components. The memory 204 can include data structures used to organize and store all or a portion of the stored data. Generally, the memory 204 can store any data used or accessible by the service monitoring system 108.
[0046] The memory 204 can include storage devices 206. The memory 204 can include any suitable number of storage devices 206, and if appropriate, the contents of the storage devices 206 can be spread across multiple storage devices 206. Although the storage devices 206 are shown as part of the memory 204, the storage devices 206 can be part of or separate from the memory 204. Additionally, one or more of the storage devices 206 can be separate and potentially remote from one or more processing devices on which the service monitoring system 108 operates. Further, the memory 204 and / or the storage devices 206 can be further subdivided or combined in any suitable manner.
[0047] In the illustrated example, the memory 204 stores a service level threshold analysis engine 114, which can include logic for managing network service level thresholds. In certain embodiments, the service level threshold analysis engine 114 includes service level threshold adjustment logic 208 and time series data analysis logic 210. Additionally, in the illustrated example, the storage device 206 stores service metric data 112, service level thresholds 212, statistical models 214, statistical model data 216, service level threshold adjustment logs 218, training data 220, untrained ML models 222, trained ML models 224, and ML analysis results 226.
[0048] Turning to the service level threshold analysis engine 114, the service level threshold analysis engine 114 can be configured to periodically evaluate service level thresholds to determine corresponding values for how to set the service level thresholds. As described above, the service level threshold analysis engine 114 can manage service level thresholds in one or more phases. InFigure 2 In the illustrated example, the service level threshold analysis engine 114 includes service level threshold adjustment logic 208 and time series data analysis logic 210, each of which is configured to perform a respective stage of managing the service level threshold. For example, the service level adjustment logic 208 may implement an automatic service level threshold adjustment stage that includes automatically setting (e.g., adjusting) the value of the service level threshold 212 based on an analysis performed using a usage statistical model 214. As another example, the time series data analysis logic 210 may include using one or more ML models (e.g., a trained ML model 224) to analyze changes to the value of the service level threshold 212 to identify patterns in the changes to the value of the service level threshold 212. These stages may be allowed to run simultaneously or at different times as needed.
[0049] The service level threshold adjustment logic 208 may be configured to manage the service level threshold 212 using the automatic service level threshold adjustment stage, during which the service level threshold adjustment logic 208 may automatically set the value of the service level threshold 212 based on an analysis performed using a statistical model 214 for managing the service level threshold 212.
[0050] In some embodiments, as part of determining how to set the service level threshold, the service level threshold analysis engine 114 may analyze service metric values for a service metric according to a statistical model 214 to identify anomalous service metric values. The statistical model 214 may be configured using statistical model data 216. The statistical model 214 may use a predicted distribution of service metric values for the service metric and a normal value range within the predicted distribution of service metric values for the service metric. An outlier may be a service metric value that falls outside the normal value range.
[0051] Examples of the statistical model 214 and the statistical model data 216 are described below. It should be understood that this is merely an example, and the statistical model 214 and / or the statistical model data 216 may be implemented in other suitable ways to identify anomalous service metric values. Any specific values of the statistical model 214 and / or the statistical model data 216 identified as part of this description are for example purposes only. For the purposes of this example, it will be assumed that each service level metric uniquely maps to one service level threshold 212; however, this is for example purposes only.
[0052] A set of data analyzed using a statistical model may be referred to as a population. In some embodiments, the data being analyzed may be, for example, service metric values for the service metric in this example. Service metric values for a particular service metric collected over time may represent the population for that service metric.
[0053] The configuration statistical model 214 can include, for each service metric for which an outlier's service metric value is being evaluated, determining a distribution (e.g., normal distribution, Poisson distribution, etc.) to be used to evaluate the service metric value; determining a confidence interval and a significance level for the service metric; defining a hypothesis based on the distribution, confidence interval, and significance level; and evaluating the hypothesis to determine whether the null hypothesis can be rejected. Thus, in some embodiments, for each service metric, the statistical model 214 and / or the statistical model data 216 can include a distribution, a confidence interval, a significance level, a null hypothesis, and an alternative hypothesis. Each of these is now described in more detail.
[0054] For each service metric, the statistical model 214 and / or the statistical model data 216 can include a predicted distribution of the values of the service metric and a normal value range within the predicted distribution of the values of the service metric. The predicted distribution of the values of the service metric can be a mathematical function that describes the probabilities of the different possible values of a variable (e.g., the service metric value of the service metric under discussion). The predicted distribution of the values of the service metric can be, for example, a so-called normal distribution (N(μ,σ), Poisson distribution (λ), or other suitable distribution. The predicted distribution can reflect the expectation of how the service metric value of the service metric can vary. The appropriate distribution can be different for different service metrics and / or can be transformed at different times for a given service metric.
[0055] As part of the statistical model 214 and the statistical model data 216, a normal value range can be defined. The normal value range can be part of the predicted distribution. The normal value range can be defined based on a hypothesis, a significance level, and a confidence interval.
[0056] To this end, the statistical model 214 and the statistical model data 216 can include a definition of a hypothesis for each service metric. The purpose of the hypothesis is to obtain a single conclusion that is statistically significant or not statistically significant at the alpha (α) level. Hypothesis testing is used to obtain a comparison with a pre-specified hypothesis, which includes whether the service metric value of the service metric is less than or equal to the absolute value of a threshold X, where X defines the normal value range of the service metric value of the service metric at the significance level α (e.g., the service metric value is in the range [-X, X]). For example, for a normal distribution, the normal value range can be -2 to 2.
[0057] A confidence interval or confidence level can be a range of reasonable values for a population. For example, the confidence level can be defined in part by the significance level and can correspond to a range of normal values within the predicted distribution of values of a service metric. The confidence level can be used to describe the magnitude of the effect (e.g., the mean difference) of a particular value (e.g., a particular service metric value). For example, for a 95% confidence interval, 95% of the values (e.g., service metric values for a service metric) fall within the specified interval. The confidence level can have any suitable value, and may be expressed as a percentage, but some example values include 90%, 95%, and 99%.
[0058] Taking a 95% confidence level as an example, the confidence level means that there is a 95% confidence interval, which is a range with an upper number and a lower number calculated based on the sample or in some other suitable way. Since the true population mean is unknown (e.g., it may change over time as new data is reported), this range describes the possible values of the mean. If multiple samples are drawn from the same population and the 95% confidence interval is calculated for each sample, the system would expect to find the population mean within 95% of these confidence intervals. That is, 95% of the sample means for a specified sample size will be within 1.96 standard deviations of the hypothesized population mean. Similarly, for a 99% confidence interval, 99% of the sample means will be within 2.5 standard deviations of the population mean.
[0059] In some embodiments, the confidence interval can be calculated by adding and subtracting a margin of error from a point estimate. The general form of the confidence interval estimate for the population mean can be X bar + / - the margin of error. In some embodiments, to calculate the interval estimate, the following formula can be used: where x bar is the sample mean, z is the number of standard deviations from the sample mean, s is the standard deviation in the sample, and n is the size of the sample.
[0060] The significance level (α) of a hypothesis test can be expressed as a percentage (e.g., α = 5%) and can represent the threshold for rejecting the null hypothesis. In some embodiments, the significance level is 5%, but the significance level can be 1% (or 0.01), 5% (or 0.05), 10% (or 0.1), or any other suitable value. The significance level can represent the probability / possibility that the true population parameter (e.g., the value of a service metric) lies outside the confidence interval. The confidence interval can be expressed as 1 - α.
[0061] Configuring the statistical model 214 can include defining hypotheses for each service metric. The hypotheses can formulate the service metric as hypotheses of a statistical model about the population (service metric values of the service metric). In some embodiments, the hypotheses can be a null hypothesis (H0) and an alternative hypothesis (HA), which can be expressed as follows:
[0062] H0: |Service metric value| > |X|, meaning that the service metric value for the service metric is expected to be outside the range [-X, X]; and
[0063] HA: |Service metric value| ≤ |X|, meaning that the service metric value for the service metric is expected to be within the range [-X, X].
[0064] The null hypothesis can be an assumption that an event will not occur (e.g., no effect on a population). The null hypothesis can reflect a true assertion unless there is sufficient statistical evidence to conclude otherwise. If the sample (e.g., one or more service metric values of a service metric) provides sufficient evidence against the claim of no effect on the population (e.g., p ≤ α), then the null hypothesis can be rejected. Otherwise, the analysis cannot reject the null hypothesis. The alternative hypothesis may be logically opposite to the null hypothesis. The alternative hypothesis is accepted after the null hypothesis is rejected.
[0065] Additionally, for each service level threshold, the statistical model data 216 can include a formula for calculating a predetermined initial value of the service level threshold, which can be referred to as the threshold_edge value or the SLO_edge value. In some embodiments, the formula for calculating the threshold_edge value can be X standard deviation values from the mean in the currently set direction. Additionally or alternatively, the threshold_edge value can be manually set by the SRE or other appropriate user.
[0066] Additionally, for each service level threshold, the statistical model data 216 can include a formula for calculating a predetermined adjustment amount of the service level threshold, which can be referred to as the threshold_step value or the SLO_step value. When appropriate, the threshold_step value can indicate the amount by which the service level threshold is adjusted. In some embodiments, the formula for calculating the threshold_step value can be a step size associated with the difference between the service metric value for the service metric and the mean of the service metric values for the service metric (normalized by the standard deviation). Additionally or alternatively, the threshold_step value can be manually set by the SRE or other appropriate user.
[0067] Additionally, the statistical model data 216 can include one or more variables that can be initially set and then adjusted when performing auto-tuning. For example, the threshold_breached variable can be a binary variable indicating whether a service level threshold has been breached. The service level threshold can be a service level threshold corresponding to the service metric for which an anomaly is being detected or any other suitable service level threshold. As another example, the threshold_manually_set variable can be a binary value indicating whether the service level threshold has been manually set (e.g., by an SRE or other suitable user).
[0068] An outlier can be a value of a service metric that falls outside the normal value range. In other words, if the service level threshold adjustment logic 208 fails to reject the null hypothesis when evaluating a service metric value for a service metric, then it can be determined that the service metric value (or values) is / are anomalous.
[0069] The service level threshold adjustment logic 208 can monitor the service metric values of a service metric and use the statistical model 214 to evaluate whether one or more service metric values are anomalous. For example, for a given service metric, the service level threshold adjustment logic 208 can determine whether one or more service metric values are anomalous by: failing to reject the null hypothesis (H0), which means the service metric value is not anomalous and within the expected normal range, or accepting the alternative hypothesis (HA), which means the service metric value is anomalous and outside the expected normal range. In some embodiments, in this way, the service level threshold adjustment logic 208 can identify one or more anomalous service metric values of a service metric.
[0070] In cases where the service level threshold adjustment logic 208 is evaluating the service metric values of multiple service metrics, the service level threshold adjustment logic 208 can use the same or different statistical models and / or statistical model data 216 to evaluate whether the service metric values of the service metrics are anomalous.
[0071] The service level threshold adjustment logic 208 may store, in the service level threshold adjustment log 218, information associated with the value of the service level threshold 212 and associated values of the service metrics. That is, the service level threshold adjustment logic 208 may store the result of setting the value of the service level threshold 212 for the service metrics. For example, the service level threshold adjustment logic 208 may store historical values of the service level threshold 212, which may include changes to the value of the service level threshold 212, the timing of changes to the service level threshold 212, the circumstances that caused the changes to the service level threshold 212, and / or any other suitable information. In some embodiments, modifications to the service level threshold are tracked via the service level threshold adjustment log 218, which, if appropriate, may allow the system to roll back to a previous value of the service level threshold 212.
[0072] In some embodiments, the service level threshold analysis engine 114 may store the result of setting the value of the service level threshold for the service metrics as time series data. For example, the service level threshold adjustment log 218 may include time series data of the value of the service level threshold 212 over time.
[0073] The time series data analysis logic 210 may be configured to manage the service level threshold 212 using a machine learning analysis phase, during which the time series data analysis logic 210 may use one or more ML models (e.g., the trained ML model 224) to analyze changes to the value of the service level threshold 212 to identify patterns in the changes to the value of the service level threshold 212. In conjunction with the machine learning analysis phase for managing the service level threshold 212, the storage module 206 may store training data 220, untrained ML models 222, trained ML models 224, and ML analysis results 226. Although described in the singular for simplicity, the present disclosure contemplates using one or more untrained ML models 222 and one or more trained ML models 224, either alone or in combination, as may be suitable for the type of service level threshold 212 and / or service metric data 112 to be analyzed using the trained ML model 224 and / or the type of desired ML analysis results 226.
[0074] The training data 220 can include data that is used to train an untrained ML model 222 to generate a trained ML model 224 and / or retrain a trained ML model 224. The training data 220 can include test service metric values and / or values of service level thresholds as time series data. The training data 220 can provide test service metric values and / or service level thresholds generated under different environments (such as during various operating conditions, various system configurations, different times of the day / month / year, etc.) to train an appropriate ML model to identify or otherwise predict patterns in the actual service metric data 112 and / or service level thresholds 212 for the service 110. In some embodiments, some or all of these test service metric values and / or service level thresholds can be used for such service metrics and / or service level thresholds: which correspond to the service metrics measured for the service 110, and for which the service metric data 112 includes service metric values and for which service level thresholds can be defined. The test service metric values and / or service level thresholds can be generated from an actual test environment or a simulated test environment.
[0075] The time series data analysis logic 210 can include training logic for using the training data 220 to perform a training phase to train an untrained ML model 222 and / or retrain a trained ML model 224 to generate a trained ML model 224.
[0076] The trained ML model 224 is a version of the ML model after being trained using the training data 220. The trained ML model 224 is ready for deployment for use during the actual operation of the service 110 and the service monitoring system 108. The untrained ML model 222 and the trained ML model 224 can include any suitable type (or combination of types) of ML model. For example, the untrained ML model 222 and the trained ML model 224 can use any suitable combination of deep learning, linear regression, logistic regression, general regression analysis, clustering analysis, neural networks, statistical classification, support vector machines, K-means clustering, supervised learning, autoregressive integrated moving average (ARIMA), Prophet, long short-term memory, convolutional neural networks, seasonal decomposition, DeepAR, and / or any other suitable type of machine learning model (some of which may overlap in type) to analyze the time series data of the service level threshold adjustment logs.
[0077] The time series data analysis logic 210 can use the trained ML model 224 to analyze the actual service level threshold adjustment log 218, including, for example, time series data associated with the value of the service level threshold 212 and the associated service metric data 112 to generate the ML analysis result 226. For example, the time series data analysis logic 210 can input the time series data associated with the value of the service level threshold 212 and the associated service metric data 112 into the trained ML model 224. The trained ML model 224 can process this information and generate one or more outputs that reflect one or more patterns associated with the analysis values of the time series data for the service level threshold. This output from the ML model 224 can include one or more predictions regarding how the detected patterns continue into a future time frame.
[0078] The time series data analysis logic 210 can use one or more machine learning models (e.g., the trained ML model 224) to analyze the time series data. One or more machine learning models (e.g., the trained ML model 224) can be trained to identify one or more patterns for the value of the service level threshold 212. Such patterns can include trends in the value of the service level threshold 212, seasonal characteristics in the value of the service level threshold 212, cyclic patterns in the value of the service level threshold 212, residual patterns in the value of the service level threshold, and / or any other suitable patterns.
[0079] Trends can include long - running patterns in the time series data for the service level threshold 212. Trends can be upward trends or downward trends. Seasonal characteristics in the value of the service level threshold 212 can include repeating patterns at certain times (e.g., certain times of the day / week / month / year). Cyclic patterns in the value of the service level threshold 212 can include fluctuations that occur at any time during the year in the time series data (e.g., the value of the service level threshold 212 over time). Residual patterns in the value of the service level threshold can represent the irregular component of the data, such as the data remaining after removing trends, seasonal characteristics, and cyclic patterns from the time series data (e.g., the value of the service level threshold 212 over time).
[0080] The time series data analysis logic 210 can output the ML analysis result 226, which can be stored in the storage device 206 or another suitable location. The ML analysis result 226 can provide the ability to display visualizations of the machine learning analysis results, including potentially the ability to view spreadsheets, charts, graphs, or other suitable visualizations to facilitate identifying patterns in the value of the service level threshold 212.
[0081] The analysis performed by the time series data analysis logic 210 can include a descriptive / explanatory component and / or a forecasting component. The descriptive / explanatory component can provide an analysis of one or more service level thresholds 212 to understand the relevance of the service level thresholds 212, any interrelationships between the service level thresholds 212, whether some of the service level thresholds 212 (via an automatic adjustment phase implemented by the service level threshold adjustment logic 208) are adjusted more than others, and so on. The forecasting component can include a prediction of the values of the service level thresholds 212 based on historical trends.
[0082] If appropriate, the trained ML model 224 can be retrained. For example, it may be desirable to periodically retrain the ML model 224, which can help maintain and / or improve the performance of the ML model 224 in providing relatively accurate prediction patterns. The trained ML model 224 can be retrained using entirely new training data 220, modifications to the existing training data 220, modifications to aspects of the ML model 224 (e.g., one or more layers of the ML model 224), or any other appropriate information.
[0083] The service monitoring system 108 can receive instructions 228 related to managing the service level thresholds 212. Depending on the nature of the instructions 228, the instructions 228 can be received by the service level threshold analysis engine 114, the service level threshold adjustment logic 208, the time series data analysis logic 210, or another appropriate component of the service monitoring system. The instructions 228 can be received from any appropriate source, including, for example, from a user. To this end, in some embodiments, some or all of the instructions 228 can reflect manual intervention by the user. The user can include, for example, any appropriate combination of data scientists, support engineers (e.g., SREs), business developers, managers, or any other appropriate type of user.
[0084] In some embodiments, the instructions 228 can include one or more of the following: a statistical model 214, statistical model data 216 (e.g., parameters for configuring the statistical model 214), instructions to freeze the values of the service level thresholds 212, instructions to unfreeze the values of the service level thresholds 212, training data 220, an untrained ML model 222, a trained ML model 224, and / or any other appropriate instructions.
[0085] In the operation of an example embodiment of the service monitoring system 108, the service level threshold analysis engine 114 (e.g., the service level threshold adjustment logic 208) can perform the service metric threshold automatic adjustment phase in the manner described in connection with Figure 4 Additional or alternatively, the service level threshold adjustment logic 208 can allow instructions in the manner described in connection with Figure 5Freeze / thaw the service level threshold 212 in the manner described. In the operation of an example embodiment of the service monitoring system 108, the service level threshold analysis engine 114 (e.g., the time series data analysis logic 210) may perform a machine learning analysis phase in the manner described in conjunction with Figure 6 as described. Figures 3 - 6 The description of which is incorporated by reference into Figure 2 the description of.
[0086] Although the functionality and data are shown as grouped in a particular manner in Figure 2 , the functionality and / or data may be separated or combined differently, which may be suitable for a particular implementation. As an example, in some embodiments, the service monitoring system 108 may receive a trained ML model 224 from another computer system that processes the training of the trained ML model 224 using training data 220 and an untrained ML model 222. As another example, although shown and described separately, the service level threshold adjustment logic 208 and the time series data analysis logic 210 may be combined if appropriate. The service monitoring system 108, the service level threshold analysis engine 114, the service level threshold adjustment logic 208, and the time series data analysis logic may be implemented using any suitable combination of hardware, firmware, and software. In some embodiments, the service monitoring system 108 may be implemented using one or more computer systems, such as the example described below with reference to Figure 7 as described.
[0087] Figures 3 - 6 illustrates an example method for managing the network service level threshold 212 according to some embodiments. Although each of these figures is primarily described with respect to a single service metric and the corresponding service level threshold, the methods described with reference to each figure may be performed for multiple service metrics and the corresponding service level thresholds for those service metrics, including simultaneously, sequentially, or in other suitable ways. Each of these figures is described below.
[0088] Figure 3 illustrates an example method 300 for managing the service level threshold 212 according to some embodiments. In some embodiments, some or all of the operations associated with method 300 are performed by the service monitoring system 108 or an entity associated with the service monitoring system 108.
[0089] In the illustrated example, method 300 includes a service level threshold auto - adjustment phase 302 (shown as including steps 306 - 314) and a machine - learning analysis phase 304 (shown as including steps 316 - 318). During the auto - adjustment phase 302, method 300 may include automatically setting (e.g., adjusting) the value of service level threshold 212 based on an analysis performed using statistical model 214. During the machine - learning analysis phase 304, method 300 may include using one or more ML models (e.g., trained ML model 224) to analyze changes to the value of service level threshold 212 to identify patterns in the changes to the value of service level threshold 212. The auto - adjustment phase 302 and the machine - learning analysis phase 304 may run simultaneously or at different times, as appropriate.
[0090] For ease of description, method 300 is described with respect to a single service metric and a single associated service level threshold 212. However, it should be understood that method 300 may be performed for multiple service metrics, each having an associated service level threshold 212. Method 300 may be performed simultaneously, sequentially, or in another manner to manage multiple service level thresholds 212, such as by automatically adjusting the values of multiple service level thresholds 212 and using one or more ML models (e.g., trained ML model 224) to analyze the results of the auto - adjustment values of one or more of the multiple service level thresholds 212.
[0091] At step 306, the service level threshold analysis engine 114 (e.g., service level threshold adjustment logic 208) initializes / adjusts one or more system parameters to prepare / adjust the auto - adjustment phase 302. For example, the service level threshold adjustment logic 208 may access configuration information for configuring the statistical model 214. For example, the configuration information may be part of the statistical model data 216. In certain embodiments, the service level threshold adjustment logic 208 may be stored in a configuration file and read therefrom, such as another markup language (YAML) file. Within an appropriate scope, a user (e.g., SLE) may provide input at step 306 to initially configure or adjust the system.
[0092] In certain embodiments, the configuration information includes a predicted distribution of service metric values for the service metric; a hypothesis for the service metric, which includes a null hypothesis and an alternative hypothesis; a predetermined initial value of the service level threshold 212 (e.g., the threshold_edge value); and a predetermined adjustment amount for the service level threshold 212 (e.g., the threshold_step value).
[0093] The service level threshold adjustment logic 208 can configure the statistical model 214 based on configuration information. On a first pass, this can include initializing the statistical model 214, and on subsequent passes, this can include adjusting the statistical model 214.
[0094] At step 308, the service level threshold adjustment logic 208 can obtain service metric data 112 related to the service 110. In some embodiments, the service monitoring system 108 obtains the service metric data 112 from one or more agents deployed throughout the system 100 and / or in any other suitable manner, as described above.
[0095] The service metric data 112 can include one or more of the following: service metric values for one or more service metrics, timestamp information, and / or any other suitable information. A service metric can be a parameter that captures aspects of the performance of the service 110. For example, a service metric can be considered an SLI. In some embodiments, the service metric measures aspects of the interaction of a user (e.g., a user of the client system 102) with the service 110 to determine whether the user may be experiencing or about to experience problems with the service 110.
[0096] At step 310, the service level threshold adjustment logic 208 can use the statistical model 214 to evaluate the values of the service metrics to determine whether those values are outliers. The statistical model 214 can include a predicted distribution of the values of the service metrics and a range of normal values within the predicted distribution of the values of the service metrics. The predicted distribution of the values of the service metrics can be, for example, a so-called normal distribution, a Poisson distribution, or another suitable distribution. The appropriate distribution can be different for different service metrics and / or can be different for a given service metric at different times. An outlier can be a value of a service metric that falls outside the range of normal values.
[0097] At step 312, the service level threshold adjustment logic 208 can automatically set the value of the service level threshold 212 for the service metric based on whether one or more service metric values of the service metric are outliers. Setting the service level threshold 212 can vary in different situations. For example, if the service level threshold 212 has not been established for the service metric, setting the service level threshold 212 can include establishing the service level threshold 212 (e.g., defining a variable for the service level threshold and setting the service level threshold 212 to an initial value (e.g., the threshold_edge value), keeping the current value of the service level threshold 212 unchanged, adjusting the current value of the service level threshold 212 by a predefined adjustment amount (e.g., the threshold_step value), or taking other appropriate actions regarding the service level threshold 212.
[0098] At step 314, the service level threshold adjustment logic 208 may store information associated with the value of the service level threshold 212 and the associated value of the service metric in the service level threshold adjustment log 218. That is, the service level threshold adjustment logic 208 may store the result of setting the value of the service level threshold 212 for the service metric. For example, the service level threshold adjustment logic 208 may store historical values of the service level threshold 212, which may include changes to the value of the service level threshold 212, the timing of changes to the service level threshold 212, the circumstances that caused the change to the service level threshold 212, and / or any other suitable information.
[0099] In some embodiments, the service level threshold analysis engine 114 may store the result of setting the value of the service level threshold for the service metric as time series data. For example, the service level threshold adjustment log 218 may include time series data of the value of the service level threshold 212 over time.
[0100] At step 316, the time series data analysis logic 210 may use one or more machine learning models (e.g., the trained ML model 224) to analyze the time series data. One or more machine learning models (e.g., the trained ML model 224) may be trained to identify one or more patterns in the value of the service level threshold 212. Such patterns may include irregular fluctuations in the value of the service level threshold 212, cyclic patterns in the value of the service level threshold 212, trends in the value of the service level threshold 212, seasonal characteristics in the value of the service level threshold 212, and / or any other suitable patterns.
[0101] At step 318, the time series data analysis logic 210 may output the ML analysis result 226, which may be stored in the storage device 206 or another suitable location. The ML analysis result 226 may provide the ability to display visualizations of the machine learning analysis results, including potentially viewing spreadsheets, charts, graphs, or other suitable visualizations to facilitate identifying patterns in the value of the service level threshold 212.
[0102] Although a single iteration of method 300 has been described, in some embodiments, method 300 includes one or more iterative processes that may be repeated at appropriate regular or irregular intervals. For example, one or more of phases 302 and 304 may be repeated at regular or irregular intervals, as indicated by iteration symbols 320 and 322, respectively.
[0103] Figure 4FIG. illustrates an example method 400 for managing a service level threshold 212 in accordance with certain embodiments. For example, in accordance with certain embodiments, method 400 may involve automatically adjusting the value of a service level threshold. In certain embodiments, some or all of the operations associated with method 400 are performed by a service monitoring system 108 or an entity associated with the service monitoring system 108. In certain embodiments, method 400 at least partially corresponds to Figure 3 the automatic adjustment phase 302 of method 300.
[0104] For ease of description, method 400 is described with respect to a single service metric and a single associated service level threshold 212. However, it should be understood that method 400 may be performed with respect to multiple service metrics, each having an associated service level threshold 212. Method 400 may be performed simultaneously, sequentially, or in another manner to manage multiple service level thresholds 212, such as by automatically adjusting the values of multiple service level thresholds 212.
[0105] At step 402, a service level threshold analysis engine 114 (e.g., service level threshold adjustment logic 208) initializes / adjusts one or more system parameters to prepare / adjust the automatic adjustment phase 302. For example, the service level threshold adjustment logic 208 may access configuration information for configuring a statistical model 214. For example, the configuration information may be part of statistical model data 216. In certain embodiments, the service level threshold adjustment logic 208 may be stored in and read from a configuration file such as YAML. Within an appropriate scope, a user (e.g., SRE) may provide input at step 402 to initially configure or adjust the system.
[0106] In certain embodiments, the configuration information includes a predicted distribution of service metric values for a service metric; a hypothesis for the service metric, the hypothesis including a null hypothesis and an alternative hypothesis; a predetermined initial value for the service level threshold 212 (e.g., a threshold_edge value); and a predetermined adjustment amount for the service level threshold 212 (e.g., a threshold_step value).
[0107] The service level threshold adjustment logic 208 may configure the statistical model 214 based on the configuration information. On a first pass, this may include initializing the statistical model 214, and on subsequent passes, this may include adjusting the statistical model 214.
[0108] At step 404, the service level threshold adjustment logic 208 can monitor the values of service metrics associated with service 110 over time. For example, the service level threshold adjustment logic 208 can obtain service metric data 112 associated with service 110. In some embodiments, the service monitoring system 108 obtains the service metric data 112 from one or more agents deployed throughout the system 100 and / or in any other suitable manner, as described above.
[0109] The service metric data 112 can include one or more of the following: service metric values of one or more service metrics, timestamp information, and / or any other suitable information. A service metric can be a parameter that captures aspects of the performance of service 110. For example, a service metric can be considered an SLI. In some embodiments, the service metric measures aspects of the interaction of a user (e.g., a user of the client system 102) with service 110 to determine whether the user may be experiencing or about to experience a problem with service 110.
[0110] At step 406, the service level threshold adjustment logic 208 can use the statistical model 214 to evaluate the values of the service metrics to determine whether those values are outliers. The statistical model 214 can include a predicted distribution of the values of the service metrics and a normal value range within the predicted distribution of the values of the service metrics. The predicted distribution of the values of the service metrics can be, for example, a so-called normal distribution, a Poisson distribution, or another suitable distribution. The appropriate distribution can be different for different service metrics and / or can be different for a given service metric at different times. An outlier can be a value of a service metric that falls outside the normal value range.
[0111] Example implementation details of the statistical model 214 and the statistical model data 216 are described above in connection with Figure 2 and are incorporated by reference.
[0112] At step 408, the service level threshold adjustment logic 208 can determine whether a performance issue has been detected for service 110. In other words, the service level threshold adjustment logic 208 can determine whether service 110 has a performance issue. In some embodiments, a service having a performance issue can include one or more of the following: one or more service level thresholds 212 for service 110 are breached by one or more values of the corresponding service metric; a user of the client system 102 identifies an issue (e.g., a customer support request from a user of the client system 102); or a user associated with the service provider system 104 (e.g., an SRE) indicates an issue. Regarding breaching one or more service level thresholds 212, in some embodiments, the service level threshold adjustment logic 208 determines whether one or more service level thresholds 212 for service 110 are breached by one or more values of the corresponding service metric. Additionally or alternatively, another component of system 100 can determine whether one or more service level thresholds 212 for service 110 are breached by one or more values of the corresponding service metric and report the breach to the service level threshold adjustment logic 208.
[0113] Step 408 can be an explicit determination or not. For example, in some embodiments, such as in response to a notification from another component (e.g., a component that evaluates whether the value of a service metric breaches the corresponding service level threshold) or user input, the service level threshold adjustment logic 208 can simply detect that service 110 is having a performance issue.
[0114] If the service level threshold adjustment logic 208 detects at step 408 that service 110 does not have a performance issue (or does not detect service 110 having a performance issue at all), then method 400 can return to step 404 to continue monitoring the values of the service metrics. If appropriate, then method 400 can return to step 402 to adjust one or more system parameters or perform other suitable configuration / adjustment operations.
[0115] If the service level threshold adjustment logic 208 detects at step 408 that service 110 has a performance issue, then at step 410, the service level threshold adjustment logic 208 can determine whether there is a service level threshold 212 for the service metric. For example, the storage module 206 can store a record of which service metrics have corresponding service level thresholds, including a mapping of possible service level thresholds to service metrics, which can be implemented in any suitable manner. In an example where the values of multiple service metrics are monitored, the service level threshold adjustment logic 208 can perform similar operations (and subsequent operations) for each service metric.
[0116] If the service level threshold adjustment logic 208 determines at step 410 that there is no service level threshold 212 for the service metric, then the service level threshold adjustment logic 208 may determine at step 412 whether one or more service metric values for the service metric are abnormal. For example, based on the evaluation performed at step 406, the service level threshold adjustment logic 208 may determine whether one or more service metric values for the service metric are abnormal.
[0117] In some embodiments, one or more values for the service metric are related in time to the time associated with the performance issue. For example, the performance issue detected at step 408 may be associated with a time (e.g., a specific time or time range). In some embodiments, to determine whether one or more service metric values for the service metric are abnormal, the service level threshold adjustment logic 208 may consider one or more values of such a service metric that reflect measurements taken at or about the same time as the time when the performance issue occurred. The specific correlation amount may vary according to a particular implementation. For example, the performance issue and one or more values of the service metric may have the same time, overlapping times, times within a specified proximity (e.g., ±n seconds before or after the time when the performance issue occurred, where n is a predefined value), or other suitable relationships.
[0118] If the service level threshold adjustment logic 208 determines at step 412 that there are no (or insufficient) abnormal values for the service metric, then method 400 may return to step 404 to continue monitoring the value of the service metric. For example, the service level threshold adjustment logic 208 may determine that no service metric values for the service metric that are temporally related to the detected performance issue (of step 408) have been identified (or that there are insufficient). If appropriate, then method 400 may return to step 402 to adjust one or more system parameters or perform other suitable configuration / adjustment operations.
[0119] If the service level threshold adjustment logic 208 determines at step 412 that one or more service metric values for the service metric are abnormal, then method 400 can proceed to step 414. For example, the service level threshold adjustment logic 208 can determine that one or more service metric values for the service metric that are related in time to the detected performance issue (of step 408) are identified. At step 414, the service level threshold adjustment logic 208 can establish a service level threshold 212 for the service metric. For example, the service level threshold adjustment logic 208 can define appropriate variables and other information in the storage module 206 to define the service level threshold for the service metric. At step 416, the service level threshold adjustment logic 208 can assign an initial value to the service level threshold established at step 414. In some embodiments, the initial value is a predefined value, which in one example can be represented as the value threshold_edge.
[0120] Thus, in some embodiments of method 400, automatically setting the value of the service level threshold 212 for a service metric based on whether one or more values of the service metric are abnormal (e.g., steps 406 and 412) includes: in response to determining that one or more values of the service metric are abnormal (e.g., step 412) and in response to determining that a service level threshold for the service metric does not exist (e.g., step 410), automatically establishing a service level threshold 212 for the service metric (e.g., step 414) and setting the value of the service level threshold 212 to a predetermined initial value (e.g., step 416).
[0121] Method 400 can then return to step 404 to continue monitoring the value of the service metric. If appropriate, method 400 can return to step 402 to adjust one or more system parameters or perform other appropriate configuration / adjustment operations.
[0122] Returning to step 410, if the service level threshold adjustment logic 208 determines at step 410 that a service level threshold 212 exists for the service metric, then at step 418, the service level threshold adjustment logic 208 can determine whether the current value of the service level threshold 212 is manually set. In some embodiments, the statistical model data 216 (or another suitable element of the service monitoring system 108) includes a flag for indicating whether the current value of the service level threshold 212 is manually set. For example, the flag can be a boolean value indicating whether the current value of the service level threshold 212 is or is not manually set. The service level threshold adjustment logic 208 can access the current value of the flag in order to determine at step 418 whether the current value of the service level threshold 212 is manually set.
[0123] If the service level threshold adjustment logic 208 determines at step 418 that the current value of the service level threshold 212 is manually set, then at step 420 the service level threshold adjustment logic 208 may retain the value of the service level threshold 212 as the current value of the service level threshold 212. At step 422, the service level threshold adjustment logic 208 may transmit an alert and a recommendation for an updated value of the service level threshold (to a user of the service provider system 104, such as an SRE). If desired, the user or another appropriate entity may manually update the value of the service level threshold 212 to the recommended value or another value. The method 400 may return to step 404 to continue monitoring the value of the service metric. If appropriate, the method 400 may return to step 402 to adjust one or more system parameters or perform other appropriate configuration / adjustment operations.
[0124] Thus, in some embodiments, the method 400 includes determining whether a service level threshold 212 for a service metric exists (e.g., step 410) and determining whether the current value of the service level threshold 212 is manually set (e.g., step 418), and automatically setting the value of the service level threshold 212 for the service metric based on whether one or more values of the service metric are abnormal (e.g., steps 406 and 424) includes: in response to determining that a service level threshold 212 for the service metric exists (e.g., step 410) and the current value of the service level threshold 212 is manually set (e.g., step 418), automatically retaining the value of the service level threshold 212 as the current value of the service level threshold 212 (e.g., step 420) and transmitting an alert including a recommended new value of the service level threshold 212 (e.g., step 422).
[0125] Returning to step 418, if the service level threshold adjustment logic 208 determines at step 418 that the current value of the service level threshold 212 is not manually set, then the service level threshold adjustment logic 208 may determine at step 424 whether one or more service metric values for the service metric are abnormal. For example, based on the evaluation performed at step 406, the service level threshold adjustment logic 208 may determine whether one or more service metric values for the service metric are abnormal.
[0126] As described above with reference to step 412, in some embodiments, one or more values of the service metric are related in time to a time associated with a performance issue. The description associated with step 412 is incorporated by reference.
[0127] If the service level threshold adjustment logic 208 determines at step 424 that there are no (or insufficient) service metric values for the service metric that are anomalous, then method 400 may return to step 404 to continue monitoring the value of the service metric. For example, the service level threshold adjustment logic 208 may determine that no service metric values for the service metric that are temporally related to the detected performance issue (of step 408) have been identified (or that insufficient ones have been identified). If appropriate, method 400 may return to step 402 to adjust one or more system parameters or perform other suitable configuration / adjustment operations.
[0128] Thus, in some embodiments, method 400 includes determining whether a service level threshold 212 for the service metric exists (e.g., step 410), and automatically setting the value of the service level threshold 212 for the service metric based on whether one or more values of the service metric are anomalous (e.g., steps 406 and 424) includes: in response to determining that one or more values of the service metric are not anomalous (e.g., step 424) (and also potentially in response to determining that the current value of the service level threshold 212 is not manually set (e.g., step 418)), automatically retaining the value of the service level threshold 212 as the current value of the service level threshold 212 (e.g., step 424 returns to step 404 after the "no" decision).
[0129] If the service level threshold adjustment logic 208 determines at step 424 that one or more service metric values for the service metric are anomalous, then method 400 may proceed to step 426. For example, the service level threshold adjustment logic 208 may determine that one or more service metric values for the service metric that are temporally related to the detected performance issue (of step 408) have been identified.
[0130] At step 426, the service level threshold adjustment logic 208 may determine whether one or more service metric values for a service metric breach the current value of the service level threshold. The service level threshold adjustment logic 208 may determine whether a service metric value for a service metric breaches the current value of the service level threshold for that service metric in any suitable manner. For example, the service level threshold adjustment logic 208 may compare an individual service metric value, multiple service metric values (e.g., over a particular time period and / or a particular number of service metric values), or a value of multiple service metric values (e.g., an average value) to the current value of the service level threshold for that service metric to determine whether the service metric value for that service metric breaches the service level threshold for that service metric. As another example, one or more service metric values for a service metric that breach the current value of the service level threshold for that service metric may include: one or more service metric values that exceed (e.g., are greater than, or greater than or equal to, depending on the implementation) the current value of the service level threshold, one or more service metric values that are below (e.g., are less than, or less than or equal to, depending on the implementation) the current value of the service level threshold, or any other suitable implementation.
[0131] In some embodiments, the service level threshold adjustment logic 208 determines whether one or more service level thresholds 212 for service 110 are breached by one or more values of the corresponding service metric. Additionally or alternatively, another component of system 100 may determine whether one or more service level thresholds 212 for service 110 are breached by one or more values of the corresponding service metric and report the breach to the service level threshold adjustment logic 208.
[0132] If the service level threshold adjustment logic 208 determines at step 426 that one or more service metric values for a service metric breach the service level threshold 212, then method 400 may return to step 404 to continue monitoring the value of the service metric. If appropriate, method 400 may return to step 402 to adjust one or more system parameters or perform other suitable configuration / adjustment operations.
[0133] If the service level threshold adjustment logic 208 determines at step 426 that one or more service metric values for a service metric do not breach the current value of the service level threshold 212, then at step 428 the service level threshold adjustment logic 208 may adjust the current value of the service level threshold 212 by a predetermined adjustment amount (e.g., the threshold_step value) for the service level threshold 212.
[0134] Accordingly, in some embodiments, method 400 includes determining whether a service level threshold 212 for a service metric exists (e.g., step 410) and determining whether one or more values of the service metric breach the current value of the service level threshold 212 (e.g., step 426), and automatically setting the value of the service level threshold 212 for the service metric based on whether one or more values of the service metric are anomalous (e.g., steps 406 and 424) includes: in response to determining that one or more values of the service metric are anomalous (e.g., step 424), that a service level threshold 212 for the service metric exists (e.g., step 410), and that one or more values of the service metric do not breach the current value of the service level threshold 212 (e.g., step 426), automatically adjusting the current value of the service level threshold 212 by a predetermined adjustment amount (e.g., step 428).
[0135] Method 400 can then return to step 404 to continue monitoring the value of the service metric. If appropriate, method 400 can return to step 402 to adjust one or more system parameters or perform other suitable configuration / adjustment operations.
[0136] In some embodiments, some or all of steps 408 - 428 can be considered part of or otherwise associated with the operation of automatically setting the value of the service level threshold for a service metric based on whether one or more values of the service metric are anomalous (and other combinations of potential factors). As described above, setting the service level threshold can include establishing the service level threshold (e.g., step 414) and setting the service level threshold to an initial value (e.g., step 416), keeping the current value of the service level threshold unchanged (e.g., step 420 and after the "no" decision at step 424), adjusting the current value of the service level threshold by a predefined adjustment amount (e.g., step 428), or taking another appropriate action with respect to the service level threshold.
[0137] As described above, in some embodiments, method 400 can be performed for multiple service metrics, each having an associated service level threshold 212. In certain implementations of such an example, in response to determining at step 408 that service 110 is experiencing a performance issue, the service level threshold adjustment logic 208 can examine the anomalous values among the multiple and possibly all of the service metrics (e.g., the service metric values of those service metrics) to determine whether to adjust the corresponding service level threshold 212. In other words, in some embodiments, in response to determining at step 410 that service 110 is experiencing a performance issue, for multiple and possibly all of the service metrics, the service level threshold adjustment logic 208 can perform the appropriate steps among steps 412 to 428 to determine how to set the corresponding service level threshold 212 for those service metrics.
[0138] Although in some embodiments, method 400 can operate substantially autonomously to automatically adjust the value of one or more service level thresholds 212, in some embodiments, a user (such as an SRE or other suitable user) can provide manual input at any point during method 400. For example, the user can provide one or more instructions 228 to the service monitoring system (e.g., to the service threshold analysis engine 114 / service level threshold adjustment logic 208).
[0139] As discussed above with reference to Figure 3 method 300, in some embodiments, the service level threshold adjustment logic 208 can store information in the service level threshold adjustment log 218 that is associated with the value of the service level threshold 212 and the associated value of the service metric. That is, the service level threshold adjustment logic 208 can store the result of setting the value of the service level threshold 212 for the service metric. For example, the service level threshold adjustment logic 208 can store historical values of the service level threshold 212, which can include changes to the value of the service level threshold 212, the timing of changes to the service level threshold 212, the circumstances that caused the change to the service level threshold 212, and / or any other suitable information.
[0140] Figure 5 Illustrated is an example method 500 for freezing and thawing the value of the service level threshold 212 according to some embodiments. In some embodiments, some or all of the operations associated with method 500 are performed by the service monitoring system 108 or an entity associated with the service monitoring system 108. In some embodiments, method 500 can be performed during the automatic adjustment phase 302 of Figure 3 method 300.
[0141] For ease of description, method 500 is described for a single service metric and a single associated service level threshold 212. However, it should be understood that method 500 can be performed for multiple service metrics, each having an associated service level threshold 212. Method 500 can be performed simultaneously, sequentially, or in another manner to manage multiple service level thresholds 212, such as by freezing and / or thawing the values of multiple service level thresholds 212. Whether to freeze / thaw the value of one or more service level thresholds 212, in some embodiments, if appropriate, the values of one or more other service level thresholds 212 can continue to be automatically adjusted by the service level threshold adjustment logic 208.
[0142] At step 502, the service level threshold analysis engine 114 (e.g., the service level threshold adjustment logic 208) can receive an instruction at a first time to freeze the value of the service level threshold 212 at a specific value. The specific value can be the current value of the service level threshold 212 or another suitable value. In some embodiments, the instruction to freeze the value of the service level threshold 212 is received from a user (such as an SRE or other suitable user).
[0143] At step 504, in response to the instruction to freeze the value of the service level threshold 212 at a specific value, the service level threshold adjustment logic 208 can cause the value of the service level threshold 212 to be maintained at that specific value. In some embodiments, the statistical model data 216 includes a flag for indicating whether the current value of the service level threshold 212 is manually set, and the service level threshold adjustment logic 208 can set the value of the flag to indicate that the value of the service level threshold 212 has been manually set in response to the instruction to freeze the value of the service level threshold 212 at a specific value.
[0144] At step 506, the service level threshold adjustment logic 208 can receive an instruction at a second time to unfreeze the value of the service level threshold 212 from the specific value. In some embodiments, the instruction to unfreeze the value of the service level threshold 212 is received from a user (such as an SRE or other suitable user). Additionally or alternatively, a time period can be specified using the freeze instruction received at step 502, and the service level threshold adjustment logic 208 can set a timer according to the time period when freezing the value of the service level threshold 212. In one embodiment, the expiration of the timer can be considered as the instruction received at the second time to unfreeze the value of the service level threshold 212 from the specific value.
[0145] At step 508, in response to the instruction to unfreeze the value of the service level threshold 212 from the specific value, the service level threshold adjustment logic 208 can allow the value of the service level threshold 212 to be adjusted from the specific value. For example, the service level threshold adjustment logic 208 can resume the automatic adjustment of the value of the service level threshold 212. To the extent that the statistical model data 216 includes a flag for indicating whether the current value of the service level threshold 212 is manually set, in some embodiments, the service level threshold adjustment logic 208 can set the value of the flag to indicate that the value of the service level threshold 212 is not manually set in response to the instruction to unfreeze the value of the service level threshold 212 from the specific value.
[0146] Although a single iteration of method 500 has been described, in some embodiments, method 500 can be initiated at any suitable one or more times during the automatic adjustment phase 302.
[0147] Figure 6 Illustrated is an example method 600 for managing a service level threshold 212 according to certain embodiments. For example, according to certain embodiments, method 600 may involve using one or more machine learning models to analyze changes to the value of service level threshold 212 to identify patterns in the changes to the value of service level threshold 212. In certain embodiments, some or all of the operations associated with method 600 are performed by service monitoring system 108 or an entity associated with service monitoring system 108. In certain embodiments, method 600 at least partially corresponds to Figure 3 the machine learning analysis phase 304 of method 300 of
[0148] For ease of description, method 600 is described for a single service metric and a single associated service level threshold 212. However, it should be understood that method 600 may be performed for multiple service metrics, each having an associated service level threshold 212. Method 600 may be performed simultaneously, sequentially, or in another manner to manage multiple service level thresholds 212, such as by automatically analyzing the results of autoregulated values of one or more of the multiple service level thresholds 212 using one or more ML models (e.g., trained ML model 224).
[0149] At step 602, the time series data analysis logic 210 may train one or more machine learning models. For example, one or more untrained ML models 222 and / or one or more trained ML models 224 may be retrained. For example, whether training an untrained ML model 22 or retraining a trained ML model 224, one or more machine learning models may be trained using training data 220.
[0150] Although the training of one or more machine learning models is described as being performed by the time series data analysis logic 210, the training of some or all of the one or more machine learning models may be performed by another module of the service monitoring system 108, or may be performed by another processing system different from the service monitoring system 108 and loaded onto the service monitoring system 108 (e.g., loaded into the storage device 206) or otherwise made accessible to the time series data analysis logic 210 as a trained ML model 224.
[0151] At step 604, the time series data analysis logic 210 may access time series data. As described above, the service level threshold adjustment log 218 may include time series data for the value of the service level threshold 212 over time. In certain embodiments, the time series data analysis logic 210 may access the time series data that is part of the service level threshold adjustment log 218.
[0152] At step 606, the time series data analysis logic 210 can analyze the time series data using one or more machine learning models (e.g., the trained ML model 224). One or more machine learning models (e.g., the trained ML model 224) can be trained to identify one or more patterns for values of the service level threshold 212. Such patterns can include irregular fluctuations in the values of the service level threshold 212, cyclic patterns in the values of the service level threshold 212, trends in the values of the service level threshold 212, seasonal characteristics in the values of the service level threshold 212, and / or any other suitable patterns.
[0153] At step 608, the time series data analysis logic 210 can output the ML analysis result 226, which can be stored in the storage device 206 or another suitable location. The ML analysis result 226 can provide the ability to display visualizations of the machine learning analysis results, including potentially viewing spreadsheets, charts, graphs, or other suitable visualizations, to facilitate identifying patterns in the values of the service level threshold 212.
[0154] Although in some embodiments, the method 600 is capable of operating substantially autonomously using one or more machine learning models to analyze changes to the values of the service level threshold 212 to identify patterns in the changes to the values of the service level threshold 212, in some embodiments, a user (such as an SRE or other suitable user) can provide manual input at any point during the method 600. For example, the user can provide one or more instructions 228 to the service monitoring system (e.g., to the service threshold analysis engine 114 / time series data analysis logic 210).
[0155] Although a single iteration of the method 600 has been described, in some embodiments, the method 600 includes one or more iterative processes that can be repeated at appropriate regular or irregular intervals, as indicated by the iteration symbol 610.
[0156] The methods 300, 400, 500, and 600 can be combined and executed using the systems and apparatuses described herein. Although shown in a logical order, the arrangement and numbering of the steps of the methods 300, 400, 500, and 600 are not intended to be limiting. The steps of the methods 300, 400, 500, and 600 can be executed in any suitable order or simultaneously with each other, as will be apparent to those skilled in the art.
[0157] Figure 7 A block diagram of an example computing device 700 in accordance with certain embodiments is illustrated. As discussed above, embodiments of the present disclosure can be implemented using computing devices. For example, one or more computing devices can be used at least in part to implementFigures 1 - 2 all or any part of the components shown in (e.g., client system 102, service provider system 104, and service monitoring system (if applicable, including service level threshold adjustment logic 208 and / or time series data analysis logic 210 and their associated storage devices 206). As another example, one or more computing devices (such as computing device 700) can be used to implement at least in part Figures 3 - 6 all or any part of the methods shown in
[0158] Computing device 700 can include one or more computer processors 702, non-persistent storage device 704 (e.g., volatile memory such as random access memory (RAM), cache memory, etc.), persistent storage device 706 (e.g., hard disk, optical drive such as compact disc (CD) drive or digital versatile disc (DVD) drive, flash memory, etc.), communication interface 712 (e.g., Bluetooth interface, infrared interface, network interface, optical interface, etc.), input device 710, output device 708, and many other elements and functionalities. Each component is described below.
[0159] In certain embodiments, the computer processor(s) 702 can be an integrated circuit for processing instructions. For example, the computer processor(s) can be one or more cores or microcores of a processor. Processor 702 can be a general-purpose processor configured to execute program code included in software executed on computing device 700. Processor 702 can be a dedicated processor, where certain instructions are incorporated into the processor design. Although Figure 7 only one processor 702 is shown in
[0160] The computing device 700 may also include one or more input devices 710, such as a touch screen, keyboard, mouse, microphone, touchpad, stylus, motion sensor, or any other type of input device. The input device 710 may allow a user to interact with the computing device 700. In some embodiments, the computing device 700 may include one or more output devices 708, such as a screen (e.g., a liquid crystal display (LCD), plasma display, touch screen, cathode ray tube (CRT) display, projector, or other display device), printer, external storage device, or any other output device. One or more of the output devices may be the same as or different from the (one or more) input devices. The (one or more) input and output devices may be locally or remotely connected to the (one or more) computer processors 702, non-persistent storage device 704, and persistent storage device 706. There are many different types of computing devices, and the foregoing (one or more) input and output devices may take other forms. In some instances, a multimodal system may allow a user to provide multiple types of input / output to communicate with the computing device 700.
[0161] In addition, the communication interface 712 may facilitate connecting the computing device 700 to a network (e.g., a LAN, WAN), such as the Internet, a mobile network, or any other type of network) and / or connecting to another device, such as another computing device. The communication interface 712 may use wired and / or wireless transceivers to perform or facilitate receiving and / or transmitting wired or wireless communications, including using an audio jack / plug, microphone jack / plug, universal serial bus (USB) port / plug, port / plug, Ethernet port / plug, fiber optic port / plug, proprietary wired port / plug, wireless signal transmission, low power (BLE) wireless signal transmission, Wireless signal transmission, RFID wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, WLAN signal transmission, visible light communication (VLC), worldwide interoperability for microwave access (WiMAX), IR communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof. The communication interface 712 may also include one or more global navigation satellite system (GNSS) receivers or transceivers, which are used to determine the location of the computing device 700 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the US-based GPS, the Russian-based global navigation satellite system (GLONASS), the Chinese-based BeiDou navigation satellite system (BDS), and the European-based Galileo global navigation satellite system. There is no limitation on the operation on any particular hardware arrangement, and thus when improved hardware or firmware arrangements are developed, the basic features here can be easily replaced for them.
[0162] The term computer-readable medium includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium may include non-transitory media, where data can be stored and which do not include carrier waves and / or transient electronic signals propagated wirelessly or via a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as CDs or DVDs, flash memory, memories, or memory devices. The computer-readable medium may have code and / or machine-executable instructions stored thereon, which may represent a process, function, subroutine, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. Code segments may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means, suitable means including memory sharing, message passing, token passing, network transmission, and the like.
[0163] All or any part of the components of computing device 700 may be implemented in circuitry. For example, the components may include and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. In some aspects, computer-readable storage devices, media, and memories may include cables or wireless signals that contain bitstreams and the like. However, when mentioned, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and signals themselves.
[0164] Certain embodiments may not provide the following technical advantages, provide some or all of the following technical advantages. These and other potential technical advantages may be described elsewhere in this disclosure or may otherwise be apparent to those skilled in the art based on this disclosure.
[0165] Certain embodiments may provide improved values of service level thresholds, which may improve overall system performance. Certain embodiments are capable of automatically adjusting service level thresholds based on statistical models (which provide more reliable and up-to-date information than may be possible with manual user intervention), rather than relying solely on sporadic manual intervention by SREs or other users to modify service level thresholds.
[0166] In certain embodiments, by automatically adjusting the values of one or more service level thresholds, the burden on SREs or other users of manually setting service level thresholds (e.g., SLOs) for service metrics (e.g., SLI) may be reduced or eliminated. This may be particularly useful for service metrics for which the impact on the end-customer experience is currently uncertain.
[0167] Certain embodiments may provide traceability of changes to service level thresholds, such as by maintaining a log of changes to the values of service level thresholds.
[0168] Certain embodiments may allow the values of one or more service level thresholds to "float" during specific tests designed to induce poor service behavior and allow the values of those one or more service level thresholds to be set thereby via system testing (e.g., via an automatic adjustment mechanism).
[0169] Certain embodiments may allow SRE engineers or other users to focus on a specific subset of important service level thresholds for manual intervention and have the system set and adjust (via an automated automatic adjustment mechanism) other service level thresholds.
[0170] If desired, some embodiments may allow for the manual freezing of service level thresholds at any time.
[0171] Some embodiments may provide a time series analysis of the values of service level thresholds adjusted using an autoregulation mechanism, which can provide insights and / or predictions.
[0172] It should be understood that the systems and methods described in this disclosure may be combined in any suitable manner.
[0173] Although this disclosure describes or illustrates particular operations occurring in a particular order, this disclosure contemplates operations occurring in any suitable order. Additionally, this disclosure contemplates repeating any suitable operation one or more times in any suitable order. Although this disclosure describes or illustrates particular operations occurring in sequence, this disclosure contemplates any suitable operations occurring substantially simultaneously, where appropriate. Any suitable operation or sequence of operations described or illustrated herein may be interrupted, suspended, or otherwise controlled by another process, such as an operating system or kernel, where appropriate. These actions may operate in an operating system environment or as a stand-alone routine that occupies all or most of the system's processing.
[0174] While this disclosure has been described with reference to illustrative embodiments, this description is not intended to be construed in a limiting sense. After reviewing the specification, various modifications and combinations of the illustrative embodiments of this disclosure, as well as other embodiments, will be apparent to those skilled in the art. Accordingly, the appended claims are intended to cover any such modifications or embodiments.
Claims
1. A computer system comprising: one or more processors; as well as One or more non-transitory computer-readable storage media storing programming executed by the one or more processors, the programming comprising instructions to: monitoring, over time, values for service metrics associated with providing a computerized service over a communications network; evaluating a value for the service metric according to a statistical model to determine whether the value is an outlier, the statistical model comprising a predicted distribution for the values of the service metric and a range of normal values within the predicted distribution for the values of the service metric, an outlier being a value for the service metric outside the range of normal values; detecting performance issues with the computerized service; responsive to detecting the performance problem with respect to the computerized service, determining whether one or more of the values for the service metric are abnormal; automatically setting a value of a service level threshold for the service metric according to whether the one or more values for the service metric are abnormal; determining whether a service level threshold exists for the service metric; as well as Automatically setting the value of the service level threshold for the service metric according to whether the one or more values for the service metric are abnormal includes: in response to determining that the one or more values for the service metric are abnormal and in response to determining that the service level threshold for the service metric does not exist, automatically performing the following operations: establishing the service level threshold for the service metric; and The value of the service level threshold is set to a predetermined initial value.
2. The computer system of claim 1, wherein: The programming also includes instructions to determine whether the value for the service metric breaches a current value of the service level threshold; as well as The instructions to detect the performance problem with the computerized service include instructions to determine the current value at which a particular one of the values for the service metric breaches the service level threshold. 3 . The computer system of claim 1 , wherein the one or more values for the service metric are temporally correlated to a time associated with the performance issue.
4. The computer system of claim 1 , wherein the instructions to automatically set the value of the service level threshold for the service metric based on whether the one or more values for the service metric are abnormal include at least one of the following: An instruction for setting the initial value of the service level threshold to a predetermined initial value; or Instructions for adjusting a current value of the service level threshold by a predetermined adjustment amount.
5. The computer system of claim 1, wherein: The programming also includes instructions to: determining whether a service level threshold exists for the service metric; as well as determining whether one or more values for the service metric breach a current value of the service level threshold; as well as The instructions for automatically setting the value of the service level threshold for the service metric based on whether the one or more values for the service metric are abnormal include: instructions for automatically adjusting the current value of the service level threshold by a predetermined adjustment amount in response to determining that the one or more values for the service metric are abnormal, a service level threshold for the service metric exists, and the one or more values for the service metric does not exceed the current value of the service level threshold.
6. The computer system of claim 1, wherein: The programming further includes instructions for determining whether a service level threshold exists for the service metric; as well as The instructions for automatically setting the value of the service level threshold for the service metric based on whether the one or more values for the service metric are abnormal include: instructions for automatically retaining the value of the service level threshold as the current value of the service level threshold in response to determining that the one or more values for the service metric are not abnormal and the service level threshold for the service metric exists.
7. The computer system of claim 1, wherein: The programming also includes instructions to: determining whether a service level threshold exists for the service metric; as well as determining whether the current value of the service level threshold is manually set; The instructions to automatically set the value of the service level threshold for the service metric based on whether the one or more values for the service metric are abnormal include: instructions to automatically retain the value of the service level threshold as the current value of the service level threshold in response to determining that a service level threshold for the service metric exists and the current value of the service level threshold is manually set; as well as The programming also includes instructions to, in response to determining that the current value of the service level threshold is manually set, transmit an alert including a suggested new value for the service level threshold.
8. The computer system of claim 1, wherein the programming further comprises instructions to store information associated with the value of the service level threshold and the associated value of the service metric in a service level threshold adjustment log.
9. The computer system of claim 8, wherein: wherein the service level threshold adjustment log includes time series data of the value of the service level threshold over time; and The programming also includes instructions to analyze the time series data using one or more machine learning models, wherein the one or more machine learning models are trained to identify one or more patterns of the values for the service level thresholds.
10. The computer system of claim 1, wherein the programming further comprises instructions to: receiving, at a first time, an instruction to freeze the value of the service level threshold at a particular value; In response to the instruction to freeze the value of the service level threshold at a specific value, causing the value of the service level threshold to be maintained at the specific value; receiving, at a second time, an instruction to unfreeze the value of the service level threshold from the particular value; as well as In response to the instruction to unfreeze the value of the service level threshold from the particular value, the value of the service level threshold is allowed to be adjusted from the particular value.
11. The computer system of claim 1 , wherein the programming further comprises instructions to: Accessing configuration information for configuring the statistical model, the configuration information comprising: the predicted distribution of the values for the service metric; Assumptions for the service metrics, including null hypotheses and alternative hypotheses; a predetermined initial value of the service level threshold; as well as a predetermined adjustment amount for the service level threshold; as well as The statistical model is configured according to the configuration information.
12. A computer-implemented method comprising: monitoring, over time, values for service metrics associated with providing a computerized service over a communications network; evaluating the value for the service metric according to a statistical model to determine whether the value is an outlier, the statistical model comprising a predicted distribution for the values of the service metric and a range of normal values within the predicted distribution for the values of the service metric, an outlier being a value for the service metric outside the range of normal values; detecting performance issues with the computerized service; responsive to detecting the performance problem with respect to the computerized service, determining whether one or more of the values for the service metric are abnormal; automatically setting a value of a service level threshold for the service metric according to whether the one or more values for the service metric are abnormal; determining whether a service level threshold exists for the service metric; Automatically setting the value of the service level threshold for the service metric according to whether the one or more values for the service metric are abnormal includes: in response to determining that the one or more values for the service metric are abnormal and in response to determining that the service level threshold for the service metric does not exist, automatically: establishing the service level threshold for the service metric; as well as The value of the service level threshold is set to a predetermined initial value.
13. The computer-implemented method of claim 12, wherein: The method further comprises: determining whether a service level threshold exists for the service metric; and determining whether one or more values for the service metric breach a current value of the service level threshold; and Automatically setting the value of the service level threshold for the service metric based on whether the one or more values for the service metric are abnormal includes: in response to determining that the one or more values for the service metric are abnormal, the service level threshold for the service metric exists, and the one or more values for the service metric does not exceed the current value of the service level threshold, automatically adjusting the current value of the service level threshold by a predetermined adjustment amount.
14. The computer-implemented method of claim 12, wherein: The method further includes determining whether a service level threshold exists for the service metric; as well as Automatically setting the value of the service level threshold for the service metric based on whether the one or more values for the service metric are abnormal includes: in response to determining that the one or more values for the service metric are not abnormal and the service level threshold for the service metric exists, automatically retaining the value of the service level threshold as the current value of the service level threshold.
15. The computer-implemented method of claim 12, wherein: The method further comprises: determining whether a service level threshold exists for the service metric; and determining whether the current value of the service level threshold is manually set; Automatically setting the value of the service level threshold for the service metric based on whether the one or more values for the service metric are abnormal includes: in response to not determining that the one or more values for the service metric are abnormal, a service level threshold for the service metric exists, and the current value of the service level threshold is manually set, automatically retaining the value of the service level threshold as the current value of the service level threshold; and The method also includes generating an alert including a suggested new value for the service level threshold in response to determining that the current value of the service level threshold is manually set.
16. The computer-implemented method of claim 12, further comprising: Information associated with the value of the service level threshold and the associated value of the service metric is stored in a service level threshold adjustment log.
17. The computer-implemented method of claim 16, wherein: wherein the service level threshold adjustment log includes time series data of the value of the service level threshold over time; and The method also includes analyzing the time series data using one or more machine learning models, wherein the one or more machine learning models are trained to identify one or more patterns of the values for the service level threshold.
18. One or more non-transitory computer-readable storage media storing programming executed by one or more processors, the programming comprising instructions to: monitoring, over time, values for service metrics associated with providing a computerized service over a communications network; evaluating the value for the service metric according to a statistical model to determine whether the value is an outlier, the statistical model comprising a predicted distribution for the values of the service metric and a range of normal values within the predicted distribution for the values of the service metric, an outlier being a value for the service metric outside the range of normal values; detecting performance issues with the computerized service; responsive to detecting the performance problem with respect to the computerized service, determining whether one or more of the values for the service metric are abnormal; automatically setting a value of a service level threshold for the service metric according to whether the one or more values for the service metric are abnormal; determining whether a service level threshold exists for the service metric; Automatically setting the value of the service level threshold for the service metric according to whether the one or more values for the service metric are abnormal includes: in response to determining that the one or more values for the service metric are abnormal and in response to determining that the service level threshold for the service metric does not exist, automatically: establishing the service level threshold for the service metric; as well as The value of the service level threshold is set to a predetermined initial value.
Citation Information
Patent Citations
Visual early warning system and method based on dynamic threshold method
CN113342623A
Service anomaly detection method and device, equipment and storage medium
CN114356734A