Anomaly detection cloud resource management system that receives external information and includes short-term resource planning

The integration of internal and external metrics in cloud resource management systems improves anomaly detection and resource allocation by predicting external events, ensuring proactive and efficient resource management in cloud environments.

JP7766189B2Active Publication Date: 2025-11-07TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024518244
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-10-08
Filing Date
2022-10-06
Publication Date
2025-11-07
Estimated Expiration
2042-10-06

AI Technical Summary

Technical Problem

Existing anomaly detection and resource management systems in cloud computing are ineffective in detecting external events that cause performance degradation, leading to delayed corrective actions and suboptimal resource allocation, especially in telecommunications applications.

Method used

Anomaly detection resource management systems that integrate internal and external metrics to predict anomalies, combining them to generate composite metrics for proactive resource allocation, utilizing short-term and long-term optimization policies.

Benefits of technology

Enhances anomaly detection accuracy and resource allocation efficiency by anticipating external impacts, preventing performance degradation and optimizing resource use both in the short and long term.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007766189000020
    Figure 0007766189000020
  • Figure 0007766189000021
    Figure 0007766189000021
  • Figure 0007766189000022
    Figure 0007766189000022
Patent Text Reader

Abstract

An anomaly detection resource management system (14) in a cloud computing system (10) monitors telecommunications applications running in the cloud (10) and detects or predicts anomalies based on internal metrics regarding the application's performance and / or resource usage and external metrics extracted from information obtained from systems external to the cloud. The internal and external metrics are merged (210) to generate composite metrics that are stored. Anomalies are detected or predicted (212) based on the composite metrics and historical data. Telecommunications traffic is predicted based in part on the detected or predicted anomalies. A short-term resource computation of application resource allocation is performed based in part on the predicted traffic and a short-term optimization policy. A long-term optimization of application resource allocation is performed based on the short-term computation and the long-term optimization policy.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 253898, filed October 8, 2021, the entire contents of which are incorporated herein by reference.

[0002] (Technical field) The present invention relates generally to computer system management, and more particularly to anomaly-detecting cloud resource management systems and methods that receive external information and include short-term and long-term resource planning. [Background technology]

[0003] "Cloud" is a collective term for computing systems in numerous privately and publicly hosted data centers connected by various networks, including the Internet. Each data center provides a shared pool of computing resources, including, for example, servers and other computing hardware, data storage, network interfaces, operating systems, applications, and services. Subscribers run applications remotely on cloud servers and store data in cloud data storage facilities. Subscribers typically access their data and interface with applications over a network such as the Internet. Data center operators allocate computing resources, such as computing hardware and data storage, to each application.

[0004] The cloud offers subscribers many benefits, including the ability to access their data and run their applications from any device with an internet connection. Data center operators perform routine technical tasks such as replacing failed hardware, backing up data, upgrading software, and providing rapid protection from evolving malware threats. Data centers have multiple, redundant power sources to prevent local power outages. Data centers can be geographically distributed, making the cloud resistant to the effects of local weather or other natural disasters. The cloud relieves subscribers of the expense and need for technical expertise to own and operate their own information technology (IT) resources.

[0005] Subscriber applications can range from very small, such as an individual accessing an email server, to large, such as implementing the functions of some or all of the core network nodes in a regional or national telecommunications network. Data center operators allocate computing resources to applications according to their size and need. Such allocations can be dynamic, with resources from a shared pool being allocated to applications according to the application's ongoing needs. Data center operators and subscribers negotiate predetermined ranges of values ​​for the application's expected performance parameters (e.g., key performance indicators, or KPIs) and agree to predetermined ranges of expected resource usage by the application to achieve the required performance. KPIs and other metadata may be recorded, and the ranges of expected performance / resource parameters are periodically adjusted to accommodate actual usage. The predetermined ranges of expected application performance and resource usage may be quantified in a service level agreement (SLA).

[0006] Anomalies in application performance and / or resource usage are known and can arise from many different causes. For example, an increase in users accessing an application (load spikes), component or network failures, malicious attacks, etc. can all adversely affect application performance. As used herein, an "anomaly" in a computing system refers to an application's performance falling outside of a predetermined range of its expected performance and / or an application's computing resource needs falling outside of a predetermined range of the application's expected resource usage. In the face of such anomalies, data center operators can increase the computing resources allocated to an application manually or through an automated anomaly detection and resolution system (ADRS) in an attempt to maintain performance within SLA limits. For example, Kardani-Moghaddam et al. describe such a system in their paper "ADRL: A Hybrid Anomaly-Aware Deep Reinforcement Learning-Based Resource Scaling in Clouds," published in IEEE Transactions on Parallel and Distributed Systems, 32, no. 3, pp. 514-526, March 1, 2021. No. 6,229,699, the disclosure of which is incorporated herein by reference in its entirety.

[0007] Such anomaly detection cloud resource management tools detect anomalous patterns and take corrective action to mitigate or even prevent performance degradation of cloud applications. treatmentThey can monitor some metrics of an application and calculate the probability or score of having an anomaly. However, internal metrics of performance and resource usage (meaning those captured from events or conditions within a computing system, such as CPU usage, memory usage, data or message throughput, latency, and Quality of Service (QoS)) do not always show a strong correlation with anomalies, especially when the anomaly is triggered by an external event or condition. For example, in telecommunication applications, external events in the cloud, such as a traffic incident or earthquake, can result in a significant increase in traffic as users make more calls. However, traditional anomaly detection and cloud resource management tools only detect anomalies when the impact reaches the application, i.e., when the traffic load overwhelms some core network nodes. Therefore, they are unable to take any corrective action, such as allocating additional resources to handle the increased call volume. treatment However, it will inevitably be too slow. Application performance will already be degraded, some calls may be dropped, users may not be able to access the network, and other degradations to QoS may occur.

[0008] Another well-known area of ​​cloud management is resource optimization. Resource optimization algorithms are widely used to host and run applications as cost-efficiently as possible. However, these techniques typically only optimize in the short term. Returning to telecommunications applications as an example, user traffic can be predicted with little advance accuracy. guess It may be assumed that this is not possible. If the cost of reallocating resources to applications is not negligible, short-term optimization may lead to suboptimal resource management in the long term.

[0009] The Background section of this specification is provided to place embodiments of the present invention in a technical and operational context and to assist those skilled in the art in understanding their scope and usefulness. The approaches described in the Background section could be pursued, but are not necessarily approaches that have been previously conceived or pursued. Unless expressly identified as such, nothing in this specification is admitted to be prior art merely by inclusion in this Background section. Summary of the Invention

[0010] The following presents a simplified summary of the present disclosure in order to provide a basic understanding to those skilled in the art. This summary is not an extensive overview of the disclosure and is not intended to identify key / critical elements of embodiments of the present invention or to delineate the scope of the present invention. Its sole purpose is to present some concepts disclosed herein in a simplified form as a prelude to the more detailed description that is presented later.

[0011] According to one or more embodiments described and claimed herein, an anomaly detection resource management system in a cloud computing system monitors applications running in the cloud, such as telecommunications applications, and detects or predicts anomalies based on internal metrics related to the application's performance and / or resource usage and external metrics extracted from information obtained from systems external to the cloud. An external information extraction and analysis function generates external metrics from the external information. A merging function combines the internal and external metrics to generate composite metrics, which are stored. Anomalies are detected or predicted based on the composite metrics and historical data. Anomalies occur when resource usage by an application falls outside a predetermined range of expected resource usage and / or application performance falls outside a predetermined range of expected performance. Telecommunications traffic is predicted based in part on the detected or predicted anomalies. Short-term resource calculations for application resource allocation are performed based in part on the predicted traffic and a short-term optimization policy. Long-term optimization of application resource allocation is performed based in part on the short-term optimization policy. calculation and is executed based on a long-term optimization policy.

[0012] One embodiment relates to a method for managing computing resources in a computing system. An application is executed in the computing system with a predetermined range of expected resource usage by the application and a predetermined range of expected performance of the application. Execution of the application is monitored and internal metrics related to resource usage by the application and performance of the application are generated. Information related to external events of the computing system is received. External metrics are extracted from the received information. The external and internal metrics are merged to generate a composite metric. Based on the composite metric, an anomaly is detected or predicted, where resource usage by the application falls outside the predetermined range of expected resource usage and / or performance of the application falls outside the predetermined range of expected performance. Computing resources required by the application are determined based on the detected or predicted anomaly.

[0013] Another embodiment relates to an anomaly detection resource management system executed on a computing system. The computing system runs a telecommunications application and receives information from an external system. The anomaly detection resource management system includes a data store and computing resources. The computing resources are configured to implement: a system monitoring function configured to monitor the application and generate internal metrics related to the application's performance and / or resource usage; an information extraction and analysis function configured to receive information from the external system and receive historical data from the data store and further configured to generate external metrics; and a feature merging function configured to receive the internal and external metrics and generate composite metrics. The data store is configured to store the composite metrics. The anomaly detection resource management system further includes: an anomaly detection function configured to receive the composite metrics and the historical data from the data store and further configured to detect or predict anomalies where resource usage by the application is outside a predetermined range of expected resource usage and / or where performance of the application is outside a predetermined range of expected performance; and a traffic estimation function configured to receive the composite metrics, the historical data from the data store, and the detected or predicted anomalies and further configured to estimate communication traffic. The anomaly detection resource management system is configured to determine computing resources required by an application based on traffic inferences and detected or predicted anomalies.

[0014] Yet another embodiment relates to a non-transitory computer-readable medium including instructions operable to cause a computing resource in a computing system to implement an anomaly detection resource management system configured to cause the computing resource to perform the steps of: executing an application on the computing system with predetermined ranges of expected resource usage by the application and predetermined ranges of expected performance of the application, monitoring the application execution to generate internal metrics related to resource usage and performance of the application, receiving information about events external to the computing system, extracting external metrics from the received information, merging the external and internal metrics to generate composite metrics, detecting or predicting an anomaly based on the composite metrics in which resource usage by the application is outside the predetermined ranges of expected resource usage and / or performance by the application is outside the predetermined ranges of expected performance, and determining computing resources required by the application based on the detected or predicted anomaly. [Brief explanation of the drawings]

[0015] The present invention will now be described in more detail with reference to the accompanying drawings, which show embodiments of the invention. However, the invention should not be construed as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. Like reference characters refer to like elements throughout. [Figure 1] FIG. 1 is a block diagram of a computing system that executes applications and an anomaly detection resource management system. [Figure 2] FIG. 2 is a flow diagram of a method for resource management in a computing system. [Figure 3]FIG. 3 is a flow diagram of a method for managing computing resources in a computing system. DETAILED DESCRIPTION OF THE INVENTION

[0016] The present invention will be described, for simplicity and exemplary purposes, primarily by reference to exemplary embodiments thereof. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be readily apparent to those skilled in the art that the present invention may be practiced without being limited to these specific details. Well-known methods and structures are not described in detail herein so as not to unnecessarily obscure the present invention. To illustrate aspects of embodiments of the present invention, a specific example of a telecommunications application executed in a large-scale computing system (also referred to as a cloud) is presented. Those skilled in the art will readily recognize that this exemplary application is not a limitation of the embodiments claimed herein, and that the inventive concepts described herein may be readily and advantageously applied to many different applications in computing systems.

[0017] FIG. 1 illustrates a computing system 10, also known as a cloud, a representative telecommunications application 12 running within the computing system 10, and an anomaly detection resource management system 14 that monitors the application 12 and receives information from an external system 16.

[0018] FIG. 2 illustrates steps in a method 100 for managing resources within a computing system 10 that executes an application 12 .

[0019] FIG. 3 illustrates steps in a method 200 for managing computing resources in a computing system.

[0020] Figures 1, 2, and 3 are referenced simultaneously in the following discussion.

[0021] As mentioned above, several anomaly detection resource management tools are known that operate to detect and correct anomalies in the performance and / or resource utilization of application 12. Some functions of such tools may include a system monitoring function 18 that generates internal metrics, an anomaly detection function 20 that detects anomalies from internal metrics and historical data stored in and retrieved from a data store 22, and, at least in the case of telecommunications applications 12, a traffic estimation function 24. These functions 18, 20, 22, and 24, in the context of traditional resource management systems, may detect anomalies in the performance and / or resource utilization of the monitored application 12 based on the internal metrics and historical data. Internal metrics are metrics detected within computing system 10 and may include metrics such as CPU load, memory usage, the number or timing of memory accesses, cache hit rates, the number or rate of context switches, the number or rate of interrupts processed, input / output volume and timing, power consumption, or other computing events or resource utilization within computing system 10 that may be detected by system monitoring function 18. In step 102, the ongoing collection of internal metrics is shown in FIG.

[0022] In many applications, such as telecommunications applications 12, more accurate prediction of traffic loads is useful for predictive resource planning. Resources can be speculatively added when increased traffic loads are predicted, rather than reacting solely to detected performance degradation due to increased traffic loads. Performance degradation According to an embodiment of the present invention, information is obtained from external systems 16 (step 104) and processed (step 108) by information extraction and analysis function 26, along with historical data from data store 22, to generate external metrics useful for anomaly detection (step 106).

[0023] External system 16 can include many types of information sources that generate various types of information. For example, a crowdsourced navigation application generates near-real-time road traffic congestion data, which can be supplemented by monitoring police, fire, and emergency medical services (EMS) communications, video from traffic cameras, and the like. Congestion data can be useful for predicting traffic because drivers and passengers stuck in traffic can call others to arrange meetings, access congestion routing apps to find alternative routes, or otherwise access telecommunications networks. Similarly, weather forecasts can be monitored because increased telecommunications traffic can correlate with bad or severe weather. Other examples include emergency broadcasts that may warn of severe weather or other natural disasters (e.g., fires, earthquakes, tsunamis, etc.), schedules of sports arenas, concert venues, etc., financial market data, news headlines, etc.

[0024] In fact, wireless telecommunications has become so embedded in modern life that network traffic is surprisingly correlated with a variety of seemingly unrelated factors. A recent paper by Rostami-Tabar et al., “Forecasting COVID-19 daily cases using phone call data,” published in Elsevier Applied Soft Computing in 2021, demonstrates a correlation between daily phone calls to healthcare facilities and daily COVID-19 caseloads, suggesting that monitoring phone call traffic can accurately predict COVID-19 caseloads. The inverse of this correlation can be used to optimize telecommunications applications12; i.e., as the number of COVID-19 caseloads increases, increased call traffic to healthcare facilities can be predicted and appropriate resources allocated to the application12. This correlation is an example of the diversity of information sources and types of information external to the cloud that can be leveraged to detect or predict anomalies in applications such as telecommunications applications.

[0025] Information from external systems 16 comes in many forms and requires processing to extract useful external metrics from it. The information extraction and analysis function 26 processes the received external information (Step 104) The external information is then processed (step 106) along with historical data from data store 22 to generate useful external metrics (step 108). The information extraction method applied depends on the type of external data. For example, for text, advanced text processing techniques such as summary generation, sentiment analysis, and word2vec may be applied. For images and videos, deep learning techniques such as convolutional neural networks (CNNs), transfer learning, large pre-trained networks (e.g., InceptionV3), or other image processing methods may be appropriate. Numerical data may be processed using statistical models, machine learning algorithms, or other numerical methods. In some cases, the external information may be partially or fully pre-processed by external system 16, simplifying the interpretation task of information extraction and analysis function 26.

[0026] The feature merging function 28 merges the internal metrics generated by the system monitoring function 18 with the external metrics generated by the information extraction and analysis function 26 to generate a combined metric (step 110). Metrics are stored in data store 22 for later use during traffic estimation and further external information extraction (step 112). The composite metrics, along with historical data, are input to anomaly detection function 20, which may operate similarly to known anomaly detection systems that utilize only internal metrics, but with the additional ability to utilize components of external metrics or aspects of the composite metrics. Based on the composite metrics, anomaly detection function 20 detects whether the performance of application 12 or its resource utilization is outside of a predetermined expected range, and whether it is likely to do so in the short term (step 114).

[0027] Traffic estimation function 30 receives the anomaly indications and associated composite metrics and historical data from anomaly detection function 20 and provides long-term predictions regarding future traffic load dynamics (step 116). Traffic estimation function 30 may use machine learning methods, such as recurrent neural networks (e.g., long short-term memory networks, or LSTM), or statistical methods (e.g., autoregressive integrated moving average, or ARIMA), to generate traffic estimates.

[0028] Short-term resource calculation function 32 receives traffic estimates from traffic estimation function 30, anomaly detection output, and short-term optimization policies from operator policy function 34. Short-term resource calculation function 32 uses the predicted future traffic to perform resource calculations for the predicted traffic segments (step 118). For example, short-term resource calculation function 32 may receive traffic estimates for the next two hours and calculate resources every 10 minutes.

[0029] The results of the short-term resource calculation, along with the anomaly detection output and long-term optimization policy from operator policy function 34, are used by long-term resource optimization function 36 to calculate an optimal resource allocation for the entire predicted traffic (step 120). Long-term resource optimization function 36 may modify the calculated resources to satisfy the long-term optimization policy. If necessary (step 122), the resources allocated to application 12 are updated (step 124) according to the results of long-term optimization function 36.

[0030] In one embodiment, short-term and long-term resource calculations and optimizations are performed as follows. Traffic predictions are available for the next k time intervals (e.g., a one-day prediction at 5-minute intervals), horizontal resource scaling is applied, there is an approximately linear relationship between traffic values and resource usage (i.e., traffic 2 or 3 times larger results in approximately 2 or 3 times the resource usage), and it is assumed that there is a function f that can calculate the resource usage expected from the traffic value.

[0031] The short-term resource calculation is based on a threshold. For each predicted traffic volume v i , the resource usage u i can be calculated using the assumed function f as u i = f(v i ). The operator can define a threshold th (0 < th ≤ 1) to be used as an overprovisioning / cost optimization target. th = x means setting the resources to use only (x × 100)% of the resources. The other part needs to be reserved. For example, th = 0.5 is set for 50% overprovisioning. Depending on the threshold, the short-term resource usage can be calculated for each u resource usage according to the following formula: i where s

[0032] The long-term resource optimization function 36 receives s i from the short-term resource calculation function 32. These define the short-term optimized resource usage, considered as a proposal or starting point. The long-term resource optimization function 36 outputs a decision d i which is the final optimized resource allocation for each i interval. The following cost function is used for long-term resource optimization: where c ris a cost ratio describing the weighting of two cost components (explained below), and d i is the final resource decision in the ith time interval. During optimization, these are the variables whose optimal values ​​are determined. s i is the recommended resource usage for the i-th interval (from the short-term resource calculation), and k is the number of intervals.

[0033] The cost function consists of two components. The first part, called the idle cost, defines the cost of running additional resources beyond the short-term optimal value. This also means that the following constraints hold for all i: s i ≦d i The second part of the cost function is called the adaptation cost, which reflects the cost of changing the resources allocated to application 12 during interval i. The adaptation cost is defined here as the square of the change in resources during the subsequent time interval. For example, d i-1 5 CPU cores are allocated in d i If application 12 has 3 CPU cores allocated, the resource change will be -2 CPU cores. The adaptation cost is its square, so the sum for this particular pair of time intervals is 4. Note that because of the square function, the adaptation cost is the same whether additional resources are allocated to application 12 or whether excess resources are removed from application 12.

[0034] The long-term optimization process consists of two steps: first, the optimal solution is determined without constraints; second, intervals are checked to see if corrective action is required.

[0035] To obtain the optimal decision, the gradient of the cost function ▽C is calculated, where the vector operator ▽ contains partial derivatives with respect to the decision.

[0036] TIFF0007766189000003.tifThe gradient of the cost function 6530 has the following form.

[0037] TIFF0007766189000004.tifThe goal is to minimize the cost function. When the gradient of the cost function is 0, there may be extreme values for certain d1, d2,..., d k where the equation ▽C = 0 can be written as the determinant Ad 0 = b, and all variables d i can be separated and collected into the vector d 0 The structure of the matrix is as follows.

[0038] TIFF0007766189000005.tifThe matrix A is invertible, and thus there exists a solution for the vector d 0 which is the unconstrained optimal solution.

[0039] d 0 Having obtained d 0 i constraints were introduced. If the first decision d i in the i-th interval is greater than the proposed value s 0 i there are no conditions for the interval. However, if the decision d

[0040] is less than the proposed value, the condition that the decision must be equal to the proposal is imposed. M The vector g ∈ R

[0041] is the constraint vector, where M ≤ k. M The indices of the intervals are collected, where the state is defined for the vector m = [m1, m2,..., m M For example, if d2 < s2 and d3 < s3 are true, then m = [2, 3] and In this approach, inequalities are avoided and the conditions are reduced to simple equations. The Lagrange multiplier method is applied to solve this constrained extreme value problem. The state space is then expanded by a Lagrange multiplier λ∈R with the same dimension as the constraint vector g. M The new cost function is: TIFF0007766189000007.tif1268The same approach as above applies. 1 The gradient of is calculated, which should be equal to 0 for the optimal solution, and the determinant is constructed based on the gradient. Note that the introduction of Lagrange multipliers expands the unknown state. The difference operator is changed. TIFF0007766189000008.tif11137The gradient of the new cost function is TIFF0007766189000009.tif2952▽ 1 C 1 After correcting the terms from =0, the matrix equation becomes: The matrices A and b were defined previously. The vector b contains the unknown decisions for a given interval, and λ contains the unknown Lagrange multipliers. To take constraints into account, bg∈R M and σ∈R kxM will be introduced.

[0042] TIFF0007766189000011.tif54130where m j are the elements of the m vector indicating the interval over which the states are defined. After solving the final matrix equation for the decision vector d, the optimal solution allowed by the constraints is obtained.

[0043] If horizontal scaling is assumed, the last element of the decision vector d should be rounded up to get the correct resource value.

[0044] The paper "LSSO: Long Short-term Scaling Optimizer," by Balazs Fodor, Laszlo Toka, and Balazs Sonkoly of the MTA-BME Network Softwarization Research Group, Budapest University of Technology and Economics, presents a long-term resource scaling optimization applicable to telecommunications applications in cloud environments. short The Long-Term Scaling Optimization Problem (LSSOP) considers long-term forecasts for the number of instances and a predefined cost model for global optimization. The optimization method for solving the LSSOP is based on a transformation into a shortest path problem, providing an optimal solution in polynomial time. We consider the impact of inaccurate forecasts using two forecasting methods: today we use yesterday's traffic, and the other uses a noisy one, weakening the forecast accuracy. While the scaling cost is low, the maximum cost gain cannot be reached. Also, when the scaling cost is small, short-term resource proposals are preferred, and optimization using inaccurate values ​​may perform unnecessary scaling actions that increase costs, so the cost gain is short term The cost gain can be negative compared to the optimal allocation of . As scaling costs increase, the impact of inaccurate predictions becomes less noticeable because the optimized allocation uses more instances than recommended to prevent scaling behavior. Also, slight inaccuracies in instance numbers are not effective. As a result, the cost gain is fairly close to that of perfect prediction. Therefore, as scaling costs increase, the prediction accuracy becomes less noticeable and near-maximum cost gain can be achieved.

[0045] 3 illustrates steps in a method 200 for managing computational (computing) resources within a computing system. An application is executed in the computing system and has a predetermined range of expected resource usage by the application and a predetermined range of expected performance for the application (block 202). Execution of the application is monitored, and internal metrics related to resource usage by the application and performance of the application are generated (block 204). In parallel, information related to events external to the computing system is received (block 206), and external metrics are extracted from the received information (block 208). The external and internal metrics are merged to generate a composite metric (block 210). Based on the composite metric, an anomaly is detected or predicted, whereby resource usage by the application falls outside the predetermined range of expected resource usage and / or performance of the application falls outside the predetermined range of expected performance (block 212). Computing resources required by the application are determined based on the detected or predicted anomaly (block 214).

[0046] Embodiments of the present invention offer numerous advantages over the prior art: Anomaly detection / prediction is more accurate when external events impact application performance. By generating external metrics, events that impact application performance are detected sooner and application performance degradation can be proactively avoided. By performing both short-term calculations and long-term optimization of application resource allocation, resource allocation is more robust and cost-effective.

[0047] According to various embodiments of the present invention, the methods described herein are intended to operate as software programs executed on a computer processor or other suitable computing resource. Specialized hardware implementations, including but not limited to application specific integrated circuits, programmable logic arrays, and other hardware devices, may likewise be configured to implement the methods described herein. Furthermore, alternative software implementations, including but not limited to distributed processing or component / object distributed processing, parallel processing, or virtual machine processing, may also be configured to implement the methods described herein.

[0048] It should also be noted that the software implementations of the invention described herein are optionally stored on non-transitory computer-readable tangible storage media, such as magnetic media such as disks or tapes, magneto-optical or opto-magnetic media such as disks, or solid-state media such as memory cards or other packages containing one or more read-only (non-volatile) memories, random access memories, or other re-writable (volatile) memories. A digital file attachment to an email or other self-contained information archive or set of archives is considered a distribution medium equivalent to a non-transitory computer-readable tangible storage medium. Accordingly, the embodiments of the invention described herein are considered to include non-transitory computer-readable tangible storage media or distribution media, including art-recognized equivalents and successor media, enumerated herein and on which the software implementations of the invention are stored.

[0049] In general, all terms used herein should be interpreted according to their ordinary meaning in the relevant technical field unless a different meaning is clearly given and / or is implied from the context in which it is used. All references to a / an / element, apparatus, component, means, step, etc. should be interpreted openly as referring to at least one instance of the element, apparatus, component, means, step, etc., unless explicitly stated otherwise. The steps of any method disclosed herein need not be performed in the exact order disclosed, unless a step is explicitly described as following or preceding another step and / or it is implicit that a step must follow or precede another step. Any feature of any of the embodiments disclosed herein may be applied to any other embodiment, whenever appropriate. Similarly, any advantage of any embodiment may be applied to any other embodiment, and vice versa. Other objects, features, and advantages of the accompanying embodiments will become apparent from the following description. As used herein, the term "configured" means configured, organized, adapted, or arranged to operate in a particular manner; it is synonymous with the term "designed." As used herein, the term "substantially" means approximately or to the extent, but not necessarily entirely. The term encompasses and accounts for mechanical or component tolerances, measurement errors, random variations, and similar sources of inaccuracy.

[0050] Of course, the present invention may be practiced otherwise than as specifically described herein without departing from its essential characteristics. The present embodiments are to be considered in all respects as illustrative and not restrictive, and all changes which come within the meaning and range of equivalency of the appended claims are intended to be embraced therein.

Claims

1. A method (200) for managing computational resources in a computing system, comprising: Executing the application (202) on the computing system with a predetermined range of expected resource usage by the application and a predetermined range of expected performance of the application; monitoring (204) the execution of the application to generate internal metrics regarding resource usage and performance by the application; receiving (206) information regarding an event external to the computing system; Extracting extrinsic metrics from the received information (208); Merging the external and internal metrics to generate a composite metric (210); Detecting or predicting anomalies (212) based on the composite metrics, wherein resource usage by the application is outside the predetermined range of expected resource usage and / or performance of the application is outside the predetermined range of expected performance; determining (214) computing resources required by the application based on the detected or predicted anomaly; A computer-implemented method (200) comprising:

2. 2. The method of claim 1, further comprising storing the composite metric, wherein detecting or predicting anomalies is further based on historical values ​​of the composite metric.

3. 3. The method of claim 2, wherein the application is a telecommunications application; further comprising inferring traffic based on the detected or predicted anomalies and current and / or historical values ​​of the composite metric; The method (200), wherein determining the computational resources required by the application is further based on traffic estimation.

4. 4. The method of claim 3, wherein determining (214) the computational resources required by the application based on the detected or predicted anomaly comprises: calculating short-term resource needs of the application; optimizing the long-term resource needs of the application based on the short-term resource calculation; A method (200) comprising:

5. 5. The method of claim 4, wherein calculating short-term resources required by the application comprises calculating the short-term resources required based on inferred traffic, a short-term optimization policy provided by an operator of the computing system, and the detected or predicted anomaly.

6. 5. The method of claim 4, wherein optimizing the long-term resources required by the application is further based on a long-term optimization policy provided by an operator of the computing system and the detected or predicted anomaly (200).

7. 7. The method of claim 6, further comprising allocating the optimized long-term resources to the application.

8. 7. The method of claim 6, wherein calculating the short-term resources required by the application comprises, for each of a plurality of time intervals i: Using the function f, we calculate the estimated traffic value v for the interval. i resource usage u for the interval i based on i , u i = f (v i ), and defining a threshold th as an overprovisioning target in the range 0<th≦1; The number of short-term instances of the resource for the interval i, s i of and calculating by A method (200) comprising:

9. 9. The method of claim 8, wherein optimizing the long-term resources required by the application based on the calculation of the short-term resources comprises: allocating short-term resources s i For each time interval i for which σ is calculated, a decision d represents a final optimized allocation of resources for the interval i based on a cost function including: an idle cost, which represents the cost of allocating additional resources beyond the short-term optimal value; and an adaptation cost, which reflects the cost of changing the resource allocation between intervals i. i The method (200) includes determining:

10. 10. The method of claim 9, wherein the cost function is: and c r is a cost ratio describing the weight of idle and adaptive costs, d i is the final resource decision in the i-th interval, s i is the short-term resource allocation calculated for the i-th interval, The method (200) wherein k is the number of time intervals.

11. 11. The method of claim 10, wherein for each interval i, s i ≦d i The method (200) is as follows.

12. 12. The method of claim 11, wherein computing an optimal solution to the cost function without constraints comprises: Calculating the gradient ▽C of the cost function, , and Vector d of unconstrained optimized resource allocation decisions 0 Solving ▽C=0 to obtain That is, solving and The method (200) further comprises calculating by:

13. 13. The method of claim 12, wherein the unconstrained optimal allocation d i 0 <s i If , then for each interval i, i 0 =s i The method (200) further comprises applying the constraint that

14. An anomaly detection resource management system (14) running on a computing system (10), the computing system (10) running a telecommunications application and receiving information from an external system (16), the anomaly detection resource management system (14) comprising: a data store (22); A computing resource, a system monitoring function (18) configured to monitor the application and generate internal metrics regarding the performance and / or resource usage of the application; an information extraction and analysis function (26) configured to receive information from the external systems and historical data from the data store, and further configured to generate external metrics; a feature merging function (28) configured to receive the internal and external metrics and further configured to generate a composite metric; an anomaly detector (20) configured to receive the composite metrics and historical data from the data store (22) and further configured to detect or predict anomalies, where resource usage by the application is outside a predetermined range of expected resource usage and / or performance of the application is outside a predetermined range of expected performance; a traffic estimation function (30) configured to receive the composite metrics and historical data and the detected or predicted anomalies from the data store (22), and further configured to estimate telecommunications traffic; a computing resource configured to implement the Equipped with the data store (22) is configured to store the composite metrics; the anomaly detection resource management system (14) is configured to determine computing resources required by the application based on the traffic inference and detected or predicted anomalies. An anomaly detection resource management system (14).

15. 15. The system (14) of claim 14, wherein the computing resource comprises: a short-term resource calculation function (32) configured to receive traffic estimates and historical data from the data store (22) and a short-term optimization policy from an operator of the computing system, the short-term resource calculation function (32) further configured to calculate short-term resource allocations for the application; a long-term resource optimization function (36) configured to receive the calculated short-term resource allocation, historical data from the data store (22), and a long-term optimization policy from an operator of the computing system, the long-term resource optimization function (36) further configured to optimize long-term resource allocation for the application; The system (14) is further configured to implement:

16. 16. The system (14) of claim 15, wherein the short-term resource calculation function (32) calculates a short-term resource allocation for the application for each of a plurality of time intervals i as follows: Using the function f, we calculate the estimated traffic value v for the interval. i resource usage u for the interval i based on i , u i = f (v i ), and defining a threshold th as an overprovisioning target in the range 0<th≦1; The number of short-term instances of the resource for the interval i, s i of and calculating by A system (14) configured to calculate by:

17. 16. The system of claim 15, wherein the long-term resource optimization function (36) optimizes short-term resource allocations s i For each time interval i for which σ is calculated, a decision d i based on a cost function including an idle cost, which represents the cost of allocating additional resources beyond the short-term optimization value, and an adaptation cost, which reflects the cost of modifying the resource allocation during the interval i.

18. 18. The system of claim 17, wherein the cost function is and c r is a cost ratio describing the weight of idle and adaptive costs, d i is the final resource decision in the i-th interval, s i is the short-term resource allocation calculated for the i-th interval, k is the number of time intervals in the system (14).

19. 19. The system of claim 18, wherein for each interval i, s i ≦d i This is the system (14).

20. 20. The system of claim 19, wherein the long-term resource optimization function (36) finds an optimal solution to the cost function without constraints by: Calculating the gradient ▽C of the cost function, , and Vector d of unconstrained optimized resource allocation decisions 0 Solving ▽C=0 to obtain That is, solving and The system (14) is further configured to calculate by:

21. 21. The system of claim 20, wherein the long-term resource optimization function (36) determines that the unconstrained optimal allocation is d i 0 <s i If , then for each interval i, i 0 =s i The system (14) is further configured to enforce the constraint that

22. 1. A non-transitory computer-readable medium comprising instructions operable to cause a computing resource in a computing system to implement an anomaly detection resource management system, the anomaly detection resource management system causing the computing resource to: Executing an application within the computing system (202) with a predetermined range of expected resource usage by the application and a predetermined range of expected performance of the application; monitoring (204) the execution of the application to generate internal metrics regarding resource usage and performance of the application; receiving (206) information regarding an event external to the computing system; Extracting extrinsic metrics from the received information (208); Merging the external and internal metrics to generate a composite metric (210); Detecting or predicting anomalies (212) based on the composite metrics, wherein resource usage by the application is outside the predetermined range of expected resource usage and / or the performance of the application is outside the predetermined range of expected performance; determining (214) computing resources required by the application based on the detected or predicted anomaly; A non-transitory computer-readable medium for causing the execution of

Citation Information

Patent Citations

  • Configuration method, apparatus, system and computer readable medium for determining new configuration of computing resources

    JP2017530482A

  • Cloud application scaling framework

    US20130254384A1