Automatic Resource Scaling Based on LSTM-RNN and Attention Mechanism
By using prediction models of LSTM-RNN and attention mechanism in cloud computing systems, the resource utilization rate is monitored and resource allocation is actively adjusted, and the application resilience and availability problems caused by passive response in the existing technology are solved, and more efficient resource management is achieved.
Patent Information
- Application Number
- CN202010275401.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-28
- Filing Date
- 2020-04-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-07-22
AI Technical Summary
The resource scaling technologies of existing cloud computing systems are mostly passive responses, resulting in application resilience and availability issues, making it difficult to actively adjust before resource utilization reaches the critical threshold.
A model based on long and short-term memory recursive neural network (LSTM-RNN) and attention mechanism is adopted to monitor resource consumer utilization, generate prediction models, and actively regulate resource allocation to adapt to future needs.
Active resource scaling is achieved before resource utilization reaches the critical threshold, improving the application resilience and availability of cloud platforms, and avoiding service level agreement violations and market share loss.
Smart Images

Figure CN112015543B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a non - transitory machine - readable storage medium, method, and system for automatic resource scaling. More specifically, it relates to a non - transitory machine - readable storage medium, method, and system for automatic resource scaling based on a long short - term memory recurrent neural network (LSTM - RNN) and an attention mechanism. Background Art
[0002] Elasticity can be an important feature of cloud computing systems. With this feature, a provisioning platform can dynamically adapt to changing workloads. In some cases, the platform scales out to add more resources. In other cases, the platform scales in to remove unused resources. The scalable feature allows users to run their applications in an elastic manner, using only the computing resources they need and paying only for what they use.
[0003] Various existing scaling techniques for cloud platforms include manual approaches or rules based on static definitions. Many automatic scaling techniques act in a passive manner, where resource availability is modified when a threshold is exceeded. After the worst - case scenario occurs, service providers respond to the situation by adjusting resource capacity. This sometimes affects the resilience and availability of applications on the cloud platform. The result may be a violation of the service - level agreement or loss of market share. Some existing techniques can also adjust resource capacity in a pre - determined manner. Whenever a threshold is reached, resource availability is modified by a predefined constant value. Summary of the Invention
[0004] In some embodiments, a non - transitory machine - readable medium stores a program executable by at least one processing unit of a computing device. The program monitors the utilization of a resource set by a resource consumer running on the computing device. Based on the monitored utilization of the resource set, the program also generates a model including a plurality of long short - term memory recurrent neural network (LSTM - RNN) layers and a set of attention mechanism layers. The model is configured to predict the future utilization of the resource set. Based on the monitored utilization of the resource set and the model, the program also determines a set of predicted values representing the utilization of the resource set by the resource consumer running on the computing device.
[0005] In some embodiments, monitoring the utilization of a resource set by a resource consumer running on a computing device can include measuring, for each of a plurality of time intervals, the utilization of the resource set by the resource consumer running on the computing device, and storing, according to a set of values of a set of metrics representing the utilization of the resource set by the resource consumer running on the computing device, the utilization measured for each of the plurality of time intervals. Generating the model can include training the model using the set of values of the set of metrics measured for each of a subset of the plurality of time intervals.
[0006] In some embodiments, the program can calculate a set of error metrics based on a plurality of sets of predicted values representing the utilization of a resource set by a resource consumer running on the computing device and a plurality of corresponding sets of values of a set of metrics, determine whether the value of an error metric in the set of error metrics is greater than a defined threshold, and update the model when it is determined that the value of an error metric in the set of error metrics is greater than the defined threshold. Updating the model can include training the model using the set of values of the set of metrics measured for each of a set of most recent time intervals, and storing the updated model in a storage device.
[0007] In some embodiments, based on the set of predicted values, the program can also adjust the allocation of resources in the resource set. The program can also include sending a notification to a client device to warn of high utilization of resources in the resource set.
[0008] In some embodiments, a method executable by a computing device monitors the utilization of a resource set by a resource consumer running on the computing device. Based on the monitored utilization of the resource set, the method also generates a model including a plurality of long short-term memory recurrent neural network (LSTM-RNN) layers and a set of attention mechanism layers. The model is configured to predict future utilization of the resource set. Based on the monitored utilization of the resource set and the model, the method also determines a set of predicted values representing the utilization of the resource set by the resource consumer running on the computing device.
[0009] In some embodiments, monitoring the utilization of a resource set by a resource consumer running on a computing device can include measuring, for each of a plurality of time intervals, the utilization of the resource set by the resource consumer running on the computing device, and storing, according to a set of values of a set of metrics representing the utilization of the resource set by the resource consumer running on the computing device, the utilization measured for each of the plurality of time intervals. Generating the model can include training the model using the set of values of the set of metrics measured for each of a subset of the plurality of time intervals.
[0010] In some embodiments, the method may further calculate an error metric set based on multiple sets of predicted values and multiple corresponding sets of values of a metric set representing the utilization of a resource set by a resource consumer running on a computing device, determine whether the value of an error metric in the error metric set is greater than a defined threshold, and update the model when it is determined that the value of an error metric in the error metric set is greater than the defined threshold. Updating the model may include training the model using the set of values of the metric set measured in each time interval in a most recent set of time intervals, and storing the updated model in a storage device.
[0011] In some embodiments, the program may further adjust the allocation of resources in the resource set based on the set of predicted values. The program may also send a notification to a client device to alert of high utilization of resources in the resource set.
[0012] In some embodiments, the system includes a set of processing units and a non-transitory machine-readable medium storing instructions. The instructions cause at least one processing unit to monitor the utilization of a resource set by a resource consumer running on the system. Based on the monitored utilization of the resource set, the instructions further cause at least one processing unit to generate a model including a set of multiple long short-term memory recurrent neural network (LSTM-RNN) layers and an attention mechanism layer. The model is configured to predict future utilization of the resource set. Based on the monitored utilization of the resource set and the model, the instructions further cause at least one processing unit to determine a set of predicted values representing the utilization of the resource set by a resource consumer running on the system.
[0013] In some embodiments, monitoring the utilization of a resource set by a resource consumer running on a computing device may include measuring the utilization of the resource set by the resource consumer running on the system in each of a plurality of time intervals, and storing the utilization measured in each of the plurality of time intervals according to a set of values of a metric set representing the utilization of the resource set by the resource consumer running on the system. Generating the model may include training the model using the set of values of the metric set measured in each time interval in a subset of the plurality of time intervals.
[0014] In some embodiments, the instructions may also cause at least one processing unit to: calculate a set of error metrics based on a set of multiple predicted value sets and a set of corresponding value sets of a set of metrics representing the utilization of a set of resources by a resource consumer running on a computing device, determine whether the value of an error metric in the set of error metrics is greater than a defined threshold, and update the model when it is determined that the value of an error metric in the set of error metrics is greater than the defined threshold. Updating the model may include training the model using the set of values of the set of metrics measured in each time interval in a set of recent time intervals, and storing the updated model in a storage device. The instructions may also cause at least one processing unit to adjust the allocation of resources in the set of resources based on the set of predicted values.
[0015] The following detailed description and the accompanying drawings provide a better understanding of the nature and advantages of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Illustrates a system for automatically scaling resource utilization based on a long short-term memory recurrent neural network and an attention mechanism according to some embodiments.
[0017] Figure 2 Illustrates a model including a long short-term memory recurrent neural network layer and an attention mechanism layer according to some embodiments.
[0018] Figure 3 Illustrates a long short-term memory unit according to some embodiments.
[0019] Figure 4 Illustrates an attention mechanism layer according to some embodiments.
[0020] Figure 5 Illustrates a process for predicting resource usage according to some embodiments.
[0021] Figure 6 Illustrates an exemplary computer system in which various embodiments may be implemented.
[0022] Figure 7 Illustrates an exemplary computing device in which various embodiments may be implemented.
[0023] Figure 8 Illustrates an exemplary system in which various embodiments may be implemented. DETAILED DESCRIPTION
[0024] In the following description, for purposes of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention as defined by the claims may include some or all of the features of these examples, either alone or in combination with other features described below, and may also include modifications and equivalents of the features and concepts described herein.
[0025] Resource auto-scaling techniques based on Long Short-Term Memory Recurrent Neural Networks (LSTM-RNNs) and attention mechanisms are described herein. In some embodiments, a computing system includes one or more resource consumers, each configured to consume resources provided by the computing system (e.g., processors, memory, network bandwidth, etc.). The computing system can monitor the resource utilization of the resource consumers and collect data associated with the resource utilization. Based on the collected data, the computing system can generate a model that includes several LSTM-RNN layers and one or more attention mechanism layers. The model can be configured to predict the resource utilization at one or more future time points based on the collected data. Based on the predicted resource utilization, the computing system can auto-scale (e.g., scale out and / or scale in) resources.
[0026] The techniques described in this application provide numerous benefits and advantages over traditional methods of auto-scaling resources. First, by using a model that predicts resource utilization, the computing system can auto-scale resources in a proactive manner, which is different from traditional techniques that scale resources in a reactive manner. In this way, the computing system can scale resources before the resource utilization reaches a critical threshold. Additionally, the computing system can be able to scale resources by a sufficient amount to meet the predicted demand for resources (as opposed to scaling resources by a defined amount that may not meet the demand for resources). Second, by using a model that includes LSTM-RNN layers and attention mechanism layers, the computing system can generate more accurate predictions of resource utilization.
[0027] Figure 1 A system 100 for auto-scaling resources based on Long Short-Term Memory Recurrent Neural Networks and attention mechanisms is shown according to some embodiments. As shown, system 100 includes client devices 105a-n and a computing system 110. Each of the client devices 105a-n can be configured to communicate with and interact with the computing system 110. For example, a user of the client devices 105a-n can access and interact with one or more of the resource consumers 115a-k running on the computing system 110.
[0028] As Figure 1As shown, the computing system includes resource consumers 115a-k, a resource monitor 120, a model manager 125, a prediction engine 130, and storage devices 135-145. The resource utilization storage device 135 is configured to store metrics representing the utilization of resources provided to the computing system 110. Examples of such resources include processors, memory (e.g., random access memory (RAM)), storage space (e.g., hard disk storage space), network bandwidth, etc. Examples of metrics representing the utilization of such resources include processor utilization metrics, memory utilization metrics, storage device utilization metrics, response time metrics (e.g., the time taken to respond to a request, such as a request for memory, a request to read from or write to a hard disk, a request to query data from a database, or a request for any other type of computing-related task), bandwidth utilization metrics (e.g., throughput), etc. For example, in some embodiments, the resource utilization storage device 135 may store separate metrics for each resource consumer 115a-k. For example, in some such embodiments, the resource utilization storage device 135 may store a first set of metrics representing the utilization of resources by the resource consumer 115a, a second set of metrics representing the utilization of resources by the resource consumer 115b, a third set of metrics representing the utilization of resources by the resource consumer 115c, etc.
[0029] The model storage device 140 may store models generated by the model manager 125. In some embodiments, each model stored in the model storage device 140 is configured to predict the utilization of a set of resources by a resource consumer 115 based on the utilization history of the set of resources by the resource consumer 115. That is, separate models are used to predict the resource utilization of different resource consumers 115. For example, a first model may be configured to predict the utilization of a set of resources by the resource consumer 115a, a second model may be configured to predict the utilization of a set of resources by the resource consumer 115b, a third model may be configured to predict the utilization of a set of resources by the resource consumer 115c, etc. The predicted value storage device 145 is configured to store metrics representing the predicted values of resource utilization. In some embodiments, the storage devices 135-145 are implemented in a single physical storage device, while in other embodiments, the storage devices 135-145 may be implemented across several physical storage devices. Although Figure 1 the storage devices 135-145 are shown as part of the computing system 110, those of ordinary skill in the art will understand that in some embodiments, the resource utilization storage device 135, the model storage device 140, and / or the predicted value storage device 145 may be external to the computing system 110.
[0030] Resource consumers 115a-k are each configured to consume resources provided by computing system 110. As described above, examples of resources provided by computing system 110 include processors, memory, storage space, network bandwidth, etc. Each resource consumer 115 can be an application running on computing system 110, a service running on computing system 110, a virtual machine running on computing system 110, a container instantiated by computing machine 110 and running on computing machine 110, a dedicated kernel running on computing system 110, or any other type of computing element that can consume resources provided by computing system 110.
[0031] Resource monitor 120 is responsible for monitoring the utilization of resources by each of resource consumers 115a-k. For example, resource monitor 120 can monitor the utilization of resources by resource consumer 115a, monitor the utilization of resources by consumer 115b, monitor the utilization of resources by resource consumer 115c, etc. In some embodiments, resource monitor 120 monitors the utilization of resources by resource consumer 115 by measuring the utilization of resources by resource consumer 115 and storing the measured resource utilization in resource utilization storage device 135 according to a metric representing resource utilization. As described above, examples of such metrics include processor utilization metrics, memory utilization metrics, storage device utilization metrics, response time metrics (e.g., the time taken to respond to a request, such as a request for memory, a request to read from or write to a hard disk, a request to query data from a database, or a request for any other type of computing-related task), bandwidth utilization metrics (e.g., throughput), etc. In some embodiments, resource monitor 120 measures the utilization of resources by resource consumer 115 and stores the measured resource utilization according to the metric at defined intervals (e.g., once per second, once every ten seconds, once every thirty seconds, once per minute, etc.).
[0032] Model manager 125 processes the generation of models. In some embodiments, the models generated by model manager 125 are configured to predict the utilization of resources by resource consumer 115 at one or more future time points based on the historical utilization of resources by resource consumer 115. In some such embodiments, model manager 125 generates a model including several LSTM-RNN layers and one or more attention mechanism layers.
[0033] Figure 2Illustrated is a model 200 including a long short-term memory recurrent neural network layer and an attention mechanism layer according to some embodiments. In some embodiments, the model manager 125 generates the model 200. As shown, the model 200 includes an input layer 205, an LSTM-RNN layer 210, an attention mechanism layer 215, an LSTM-RNN layer 220, an attention mechanism layer 225, an LSTM-RNN layer 230, and an output layer 235. The input layer 205 is configured to receive input data. In some embodiments, the input data received by the input layer 205 includes a set of time series data associated with the utilization of a resource set by the resource consumer 115. For example, such a set of time series data may include a set of metrics representing the utilization of the resource set by the resource consumer 115 at a series of past time instances. In some embodiments, the input layer 205 performs a set of activation function calculations based on the input data and then forwards the calculation results to the LSTM-RNN layer 210. Each of the LSTM-RNN layers 210, 220, and 230 may be implemented by an LSTM-RNN cell described below with reference to Figure 3 Each of the attention mechanism layers 215 and 225 may be implemented by an attention mechanism described below with reference to Figure 4 The output layer 235 is configured to output predicted values representing the resource utilization in a set of future time intervals (e.g., a predicted value representing the resource utilization in the next second, a predicted value representing the resource utilization in the next two seconds, a predicted value representing the resource utilization in the next three seconds, etc.).
[0034] Figure 2An example of a model configured to predict the utilization of a resource by a resource consumer 115 at one or more future time points based on the historical utilization of the resource by the resource consumer 115 is shown. Any model can be used that includes an input layer, followed by an LSTM-RNN layer, followed by an attention mechanism layer, followed by zero or more pairs of an LSTM-RNN layer and an attention mechanism layer, followed by an LSTM-RNN layer, and finally an output layer. For example, one such model can include an input layer, followed by an LSTM-RNN layer, followed by an attention mechanism layer, followed by an LSTM-RNN layer, and finally an output layer. This particular model has zero pairs of an LSTM-RNN layer and an attention mechanism layer. Another example of such a model can include an input layer, followed by an LSTM-RNN layer, followed by an attention mechanism layer, followed by an LSTM-RNN layer, followed by an attention mechanism layer (the first pair of an LSTM-RNN layer and an attention mechanism layer), followed by an LSTM-RNN layer, followed by an attention mechanism layer (the second pair of an LSTM-RNN layer and an attention mechanism layer), followed by an LSTM-RNN layer, and finally an output layer. Those of ordinary skill in the art will understand that there is no additional limitation on the number of pairs of an LSTM-RNN layer and an attention mechanism layer.
[0035] Figure 3 A long short-term memory recurrent neural network (LSTM-RNN) cell 300 according to some embodiments is shown. In some embodiments, the LSTM-RNN cell 300 can be used to implement each of the LSTM-RNN layers 210, 220, and 230. As shown, the LSTM-RNN cell 300 includes a forget gate layer (f t ), an input gate layer (i t ), an input gate modulation layer (z t ), and an output gate layer (o t ). In addition, the LSTM-RNN cell 300 receives the cell state (C t-1 ) from the previous time interval, the hidden state (h t-1 ) from the previous time interval, and the input variable (x t ) for the current time interval. In some embodiments, the input variable x t can be a vector of input variables. As shown, the LSTM-RNN cell 300 outputs the cell state (C t ) for the current time interval and the hidden state (h t ) for the current time interval. The parameters of the LSTM-RNN cell 300 can be determined using the following equations:
[0036] f t = σ(W f · |h t-1 , x t| + b f )
[0037] i t = σ(W i ·|h t-1 , x t | + b i )
[0038] z t = tanh(W C ·|h t-1 , x t | + b C )
[0039] C t = (f t × C t-1 ) + (i t × z t )
[0040] o t = σ(W o ·|h t-1 , x t | + b o )
[0041] h t = o t × tanh(C t )
[0042] Among them, f t is the forget gate layer at time interval t, σ is the sigmoid activation function, W f is the weight matrix of the forget gate layer, h t-1 is the hidden state at time interval t - 1, x t is the input variable at time interval t, b f is the bias of the forget gate layer, i t is the input gate layer at time interval t, W i is the weight matrix of the input gate layer, b i is the bias of the input gate layer, z t is the input gate modulation layer at time interval t, tanh is the hyperbolic tangent activation function, W C is the weight matrix of the input gate modulation layer, b c is the bias of the input gate modulation layer, C t is the cell state at time interval t, C t-1 is the cell state at time interval t - 1, o t is the output gate layer at time interval t, W o is the weight matrix of the output gate layer, b o is the bias of the output gate layer, h tis the hidden state at time interval t.
[0043] The forget gate layer is configured to determine what information to discard from the cell state. To make this determination, the forget gate layer applies the sigmoid activation function to the hidden state at time interval t-1 and applies the sigmoid activation function to the input variables at time interval t. The forget gate layer can output a value between 0 and 1 for each digit in the cell state at time interval t-1. A value of 1 indicates that the corresponding digit is to be completely retained, while a value of 0 indicates that the corresponding value is to be completely removed. The next step is to decide what new information to store in the cell state at time interval t. First, the input gate layer applies the sigmoid activation function to determine which values are to be updated. Next, the input gate modulation layer applies the hyperbolic tangent activation function to produce a vector of new candidate values that can be added to the cell state. Then, the results of the input gate layer and the input gate modulation layer are combined to update the cell state. Finally, the output gate layer applies the sigmoid activation function to determine which parts of the cell state will be included in the output. Then, the hyperbolic tangent activation function is applied to the cell state at time interval t in order to produce a value between -1 and 1, which is then multiplied by the output in order to output only the determined parts.
[0044] Figure 4 Illustrates an attention mechanism 400 according to some embodiments. In some embodiments, the attention mechanism 400 can be used to implement each of the attention mechanism layers 215 and 225. As shown, the attention mechanism 400 receives n inputs y1, y2, y3, ……, y n and context C. In some embodiments, the output of the LSTM-RNN cell 300h i is the input of y i and the cell state C t of the LSTM-RNN cell 300 is the input of context C. The attention mechanism 400 outputs a vector z. In some embodiments, the vector z is the weighted arithmetic mean of y i . In some embodiments, the weights are determined according to the relevance of each y i to the given context C.
[0045] Back to Figure 1, after the model manager 125 generates a model, the model manager 125 can retrieve data associated with the utilization rate of resources by the resource consumer 115 (e.g., a metric representing the utilization rate of resources by the resource consumer 115) from the resource utilization storage device 135 for a period of time (e.g., a period starting from the time point when the resource consumer 115 starts using the resources and ending at the current time point). Next, the model manager 125 can use the retrieved data to train the model. Once the model manager 125 finishes training the model, the model manager 125 stores it in the model storage device 140.
[0046] The model manager 125 can also perform updates on the model. In some cases, the model manager 125 performs updates on the model stored in the model storage device 140 at defined intervals (e.g., once an hour, once a day, once a week, etc.). To update the model, the model manager 125 retrieves it from the model storage device 140. Next, the model manager 125 calculates a set of error metrics. In some embodiments, the model manager 125 calculates a normalized root mean square error (NRMSE) metric and a normalized mean absolute percentage error (NMAPE) metric. The model manager 125 can calculate the NRMSE metric using the following equation:
[0047]
[0048]
[0049] where RMSE is the root mean square error value, y i is the actual value of the resource utilization retrieved by the model manager 125 from the resource utilization storage device 135, is the predicted value of the resource utilization retrieved by the model manager 125 from the predicted value storage device 145, n is the number of observed values in the dataset for which predictions have been determined, NRMSE is the normalized root mean square error value, y max is the maximum value in the dataset for which predictions have been determined, and y min is the minimum value in the dataset for which predictions have been determined.
[0050] The model manager 125 can calculate the NMAPE metric using the following equation:
[0051]
[0052]
[0053] where MAPE is the mean absolute percentage error value, y i is the actual value of the resource utilization retrieved by the model manager 125 from the resource utilization storage device 135, is the predicted value of the resource utilization retrieved by the model manager 125 from the predicted value storage device 145, n is the number of observations in the dataset for which predictions have been determined, NMAPE is the normalized mean absolute percentage error value, y max is the maximum value in the dataset for which predictions have been determined, and y min is the minimum value in the dataset for which predictions have been determined.
[0054] After calculating the set of error metrics, the model manager 125 then determines whether one of the error metric values is greater than a defined threshold (e.g., five percent, ten percent, fifteen percent, twenty-five percent, etc.). If so, the model manager 125 updates the model by retrieving data associated with the utilization of resources by the resource consumer 115 that has not been used for training the model from the resource utilization storage device 135 and training the model with the retrieved data. When the training of the model is complete, the model manager 125 replaces the old version of the model in the model storage device 140 with the updated model.
[0055] The prediction engine 130 is responsible for predicting the utilization of resources by the resource consumers 115a-k. For example, to predict the utilization of resources by the resource consumer 115, the prediction engine 130 can retrieve from the model storage device 140 a model configured to predict the utilization of resources by the resource consumer 115. Next, the prediction engine 130 retrieves from the resource utilization storage device 135 data associated with the past utilization of resources by the resource consumer 115. In some embodiments, the retrieved data is a defined number of most recent data (e.g., the most recent five data, the most recent ten data, etc.) associated with the past utilization of resources by the resource consumer 115. In other embodiments, the retrieved data is the most recent data within a defined time period associated with the past utilization of resources by the resource consumer 115 (e.g., the most recent thirty seconds of data, the most recent five minutes of data, the most recent fifteen minutes of data, etc.). The prediction engine 130 then uses the model to generate a set of predicted values representing the utilization of resources by the resource consumer 115 by using the retrieved data as input to the model. Finally, the prediction engine 130 stores the generated set of predicted values in the predicted value storage device 145. In some embodiments, the prediction engine 130 determines the predicted values representing the utilization of resources by the resource consumer 115 at defined intervals (e.g., once per second, once per ten seconds, once per thirty seconds, once per minute, etc.).
[0056] The prediction engine 130 may perform various operations after generating a set of predicted values representing the utilization of resources by the resource consumer 115. In some cases, the prediction engine 130 may scale the resources utilized by the resource consumer 115 to meet the predicted demand for resources. For example, if the resource consumer 115 is an application running on the computing system 110 that is currently using 1 gigabyte (GB) of memory, and the predicted utilization of the memory five seconds from now is 2 GB of memory, then the prediction engine 130 may allocate an additional 1 GB of memory (i.e., expand) for the application in order to meet the expected utilization of 2 GB of memory. If the predicted utilization of the memory is less than 1 GB of memory, then the prediction engine 130 may also scale down (e.g., deallocate) the memory used for the application. As another example, if the resource consumer 115 is a virtual machine running on the computing system 110, and the predicted utilization of the processor of the virtual machine five seconds from now is 90%, then the prediction engine 130 may instantiate an additional similarly configured virtual machine to handle the expected increased processing load. Alternatively, or in combination with scaling resources in response to the predicted utilization of resources, the prediction engine 130 may send a notification to the client device 105 that may be using the resource consumer 115 to alert the user of the client device 105 that the resource utilization is high.
[0057] Figure 5 A process 500 for predicting resource usage in accordance with some embodiments is shown. In some embodiments, the computing system 110 executes the process 500. The process 500 begins at 510 by monitoring the utilization of a set of resources by a resource consumer running on the computing system. Refer Figure 1 As an example, the resource monitor 120 may monitor the utilization of a set of resources by the resource consumer 115 running on the computing system 110. In some embodiments, the resource monitor 120 measures the utilization of the set of resources by the resource consumer 115 and stores the measured resource utilization according to a metric representing the utilization of the set of resources.
[0058] Next, based on the monitored utilization of the set of resources, the process 500 generates, at 520, a model that includes a plurality of long short-term memory recurrent neural network (LSTM-RNN) layers and a set of attention mechanism layers, the model being configured to predict future utilization of the set of resources. Refer Figure 1 and Figure 2 As an example, the model manager 125 may generate a model that includes a plurality of LSTM-RNN layers and a set of attention mechanism layers, such as Figure 2The model 200 shown. After generating the model, the model manager 125 can retrieve data associated with the utilization of resources by the resource consumer 115 over a period of time from the resource utilization storage device 135, train the model with the retrieved data, and then store the model in the model storage device 140 when the model manager 125 finishes training the model.
[0059] Finally, based on the monitored utilization of the resource set and the model, process 500 determines, at 530, a set of predicted values representing the utilization of the resource set by the resource consumers running on the computing system. Refer to Figure 1 As an example, the prediction engine 130 can determine the set of predicted values by the following steps: First, retrieve from the model storage device 140 a model configured to predict the utilization of resources by the resource consumer 115, retrieve data associated with the past utilization of resources by the resource consumer 115 from the resource utilization storage device 135, and then use the retrieved data as input to the model to generate the set of predicted values using the model. The prediction engine 130 then stores the determined set of predicted values in the predicted value storage device 145.
[0060] Figure 6 An exemplary computer system 600 for implementing the various embodiments described above is shown. For example, the computer system 600 can be used to implement the client devices 105a-n and the computing system 110. The computer system 600 can be a desktop computer, a laptop computer, a server computer, or any other type of computer system or a combination thereof. Some or all of the elements of the resource consumers 115a-n, the resource monitor 120, the model manager 125, the prediction engine 130, or a combination thereof can be included or implemented in the computer system 600. In addition, the computer system 600 can implement many of the operations, methods, and / or processes described above (e.g., process 500). As Figure 6 shown, the computer system 600 includes a processing subsystem 602 that communicates with an input / output (I / O) subsystem 608, a storage subsystem 610, and a communication subsystem 624 via a bus subsystem 626.
[0061] The bus subsystem 626 is configured to facilitate communication between the various components and subsystems of the computer system 600. Although the bus subsystem 626 is shown in Figure 6is shown as a single bus in the figure, but those of ordinary skill in the art will understand that the bus subsystem 626 can be implemented as multiple buses. The bus subsystem 626 can be any one of several types of bus structures using any one of various bus architectures (e.g., memory bus or memory controller, peripheral bus, local bus, etc.). Examples of bus architectures can include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, Peripheral Component Interconnect (PCI) bus, Universal Serial Bus (USB), etc.
[0062] The processing subsystem 602, which can be implemented as one or more integrated circuits (e.g., traditional microprocessors or microcontrollers), controls the operation of the computer system 600. The processing subsystem 602 can include one or more processors 604. Each processor 604 can include one processing unit 606 (e.g., a single-core processor such as processor 604-1) or several processing units 606 (e.g., a multi-core processor such as processor 604-2). In some embodiments, the processors 604 of the processing subsystem 602 can be implemented as independent processors, while in other embodiments, the processors 604 of the processing subsystem 602 can be implemented as multiple processors integrated into a single chip or multiple chips. Nevertheless, in some embodiments, the processors 604 of the processing subsystem 602 can also be implemented as a combination of independent processors and multiple processors integrated into a single chip or multiple chips.
[0063] In some embodiments, the processing subsystem 602 can execute various programs or processes in response to program code and can maintain multiple simultaneously-executing programs or processes. At any given time, some or all of the program code to be executed can reside in the processing subsystem 602 and / or the storage subsystem 610. Through appropriate programming, the processing subsystem 602 can provide various functions, such as the functions described above with reference to process 500, etc.
[0064] The I / O subsystem 608 may include any number of user interface input devices and / or user interface output devices. User interface input devices may include keyboards, pointing devices (e.g., mice, trackballs, etc.), touchpads, touchscreens incorporated into displays, rollers, click wheels, dials, buttons, switches, keypads, audio input devices with voice recognition systems, microphones, image / video capture devices (e.g., webcams, image scanners, barcode readers, etc.), motion sensing devices, gesture recognition devices, eye (e.g., blink) recognition devices, biometric input devices, and / or any other type of input device.
[0065] User interface output devices may include visual output devices (e.g., display subsystems, indicator lights, etc.), audio output devices (e.g., speakers, headphones, etc.), and the like. Examples of display subsystems may include cathode ray tubes (CRTs), flat panel devices (e.g., liquid crystal displays (LCDs), plasma displays, etc.), projection devices, touchscreens, and / or any other type of device and mechanism for outputting information from the computer system 600 to a user or another device (e.g., a printer).
[0066] As Figure 6 shown, the storage subsystem 610 includes system memory 612, computer-readable storage media 620, and a computer-readable storage media reader 622. The system memory 612 may be configured to store software in the form of program instructions that can be loaded and executed by the processing subsystem 602, as well as data generated during the execution of the program instructions. In some embodiments, the system memory 612 may include volatile memory (e.g., random access storage device (RAM)) and / or non-volatile storage devices (e.g., read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc.). The system memory 612 may include different types of memory, such as static random access memory (SRAM) and / or dynamic random access memory (DRAM). In some embodiments, the system memory 612 may include a basic input / output system (BIOS) configured to store basic routines to facilitate the transfer of information between components within the computer system 600 (e.g., during startup). Such BIOS may be stored in ROM (e.g., a ROM chip), flash memory, or any other type of storage device configured to store the BIOS.
[0067] As Figure 6As shown, system memory 612 includes application programs 614, program data 616, and an operating system (OS) 618. The OS 618 can be one of the following: various versions of Microsoft Windows, Apple Mac OS, Apple OS X, Apple macOS, and / or Linux operating systems, various commercially available UNIX or UNIX-like operating systems (including but not limited to various GNU / Linux operating systems, Google OS, etc.) and / or mobile operating systems such as Apple iOS, Windows Phone, Windows Mobile, Android, BlackBerry OS, BlackBerry 10, and Palm OS, WebOS operating systems.
[0068] The computer-readable storage medium 620 can be a non-transitory computer-readable medium configured to store software (e.g., programs, code modules, data structures, instructions, etc.). Many of the above components (e.g., resource consumers 115a-n, resource monitor 120, model manager 125, and prediction engine 130) and / or processes (e.g., process 500) can be implemented as software that, when executed by a processor or processing unit (e.g., the processor or processing unit of processing subsystem 602), performs the operation of such components and / or processes. The storage subsystem 610 can also store data for software execution or data generated during software execution.
[0069] The storage subsystem 610 can also include a computer-readable storage medium reader 622 configured to communicate with the computer-readable storage medium 620. The computer-readable storage medium 620, together with and optionally in combination with the system memory 612, can comprehensively represent remote, local, fixed, and / or removable storage devices plus storage media for temporarily and / or more permanently containing, storing, transmitting, and retrieving computer-readable information.
[0070] The computer-readable storage medium 620 can be any suitable medium known or used in the art, including storage media such as volatile, non-volatile, removable, and non-removable media implemented in any method or technology for storing and / or transmitting information. Examples of such storage media include RAM, ROM, EEPROM, flash memory or other storage technologies, compact disc read-only memory (CD-ROM), digital versatile disk (DVD), Blu-ray Disc (BD), cassette tapes, magnetic tapes, disk storage devices (e.g., hard disk drives), zip drives, solid-state drives (SSD), flash memory cards (e.g., secure digital (SD) cards, compact flash cards, etc.), USB flash drives, or any other type of computer-readable storage medium or device.
[0071] The communication subsystem 624 serves as an interface for receiving data from and transmitting data to other devices, computer systems, and networks. For example, the communication subsystem 624 may allow the computer system 600 to connect to one or more devices via a network (e.g., a personal area network (PAN), a local area network (LAN), a storage area network (SAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a global area network (GAN), an intranet, the Internet, a network of any number of different types of networks, etc.). The communication subsystem 624 may include any number of different communication components. Examples of such components may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using cellular technologies such as 2G, 3G, 4G, 5G, etc., wireless data technologies such as Wi-Fi, Bluetooth, ZigBee, etc., or a combination thereof), global positioning system (GPS) receiver components, and / or other components. In some embodiments, in addition to or instead of components configured for wireless communication, the communication subsystem 624 may also provide components configured for wired communication (e.g., Ethernet).
[0072] Those of ordinary skill in the art will recognize thatFigure 6 The architecture shown is merely an example architecture of computer system 600, and computer system 600 may have more or fewer components than shown, or components with different configurations. Figure 6 The various components shown can be implemented in hardware, software, firmware, or any combination thereof, including one or more signal processing and / or application specific integrated circuits.
[0073] Figure 7 An exemplary computing device 700 for implementing the various embodiments described above is shown. For example, computing device 700 can be used to implement client devices 105a - n. Computing device 700 can be a mobile phone, a smart phone, a wearable device, an activity tracker or manager, a tablet computer, a personal digital assistant (PDA), a media player, or any other type of mobile computing device or a combination thereof. As Figure 7 shown, computing device 700 includes a processing system 702, an input / output (I / O) system 708, a communication system 718, and a storage system 720. These components can be coupled via one or more communication buses or signal lines.
[0074] The processing system 702, which can be implemented as one or more integrated circuits (e.g., a conventional microprocessor or microcontroller), controls the operation of computing device 700. As shown, the processing system 702 includes one or more processors 704 and a memory 706. The processors 704 are configured to run or execute various software and / or instruction sets stored in the memory 706 to perform the various functions of computing device 700 and process data.
[0075] Each of the processors 704 in the processor 704 can include one processing unit (e.g., a single - core processor) or several processing units (e.g., a multi - core processor). In some embodiments, the processors 704 of the processing system 702 can be implemented as separate processors, while in other embodiments, the processors 704 of the processing system 702 can be implemented as multiple processors integrated into a single chip. Nevertheless, in some embodiments, the processors 704 of the processing system 702 can also be implemented as a combination of separate processors and multiple processors integrated into a single chip.
[0076] The memory 706 can be configured to receive and store software (e.g., the operating system 722, applications 724, I / O module 726, communication module 728, etc. from the storage system 720) in the form of program instructions that can be loaded and executed by the processor 704 and in the form of data generated during the execution of the program instructions. In some embodiments, the memory 706 can include volatile memory (e.g., random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc.), or a combination thereof.
[0077] The I / O system 708 is responsible for receiving inputs through various components and providing outputs through various components. As shown in this example, the I / O system 708 includes a display 710, one or more sensors 712, a speaker 714, and a microphone 716. The display 710 is configured to output visual information (e.g., a graphical user interface (GUI) generated and / or rendered by the processor 704). In some embodiments, the display 710 is a touch screen configured to also receive touch-based inputs. The display 710 can be implemented using liquid crystal display (LCD) technology, light-emitting diode (LED) technology, organic light-emitting diode (OLED) technology, organic electroluminescence (OEL) technology, or any other type of display technology. The sensors 712 can include any number of different types of sensors for measuring physical quantities (e.g., temperature, force, pressure, acceleration, direction, light, radiation, etc.). The speaker 714 is configured to output audio information, and the microphone 716 is configured to receive audio inputs. Those of ordinary skill in the art will understand that the I / O system 708 can include any number of additional, fewer, and / or different components. For example, the I / O system 708 can include a keypad or keyboard for receiving inputs, ports for transmitting data, receiving data, and / or power and / or communicating with another device or component, an image capture component for capturing photos and / or videos, etc.
[0078] The communication system 718 serves as an interface for receiving data from and transmitting data to other devices, computer systems, and networks. For example, the communication system 718 may allow the computing device 700 to connect to one or more devices via a network (e.g., a personal area network (PAN), a local area network (LAN), a storage area network (SAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a global area network (GAN), an intranet, the Internet, a network of any number of different types of networks, etc.). The communication system 718 may include any number of different communication components. Examples of such components may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using cellular technologies such as 2G, 3G, 4G, 5G, etc., wireless data technologies such as Wi-Fi, Bluetooth, ZigBee, etc., or combinations thereof), a global positioning system (GPS) receiver component, and / or other components. In some embodiments, in addition to or instead of components configured for wireless communication, the communication system 718 may also provide components configured for wired communication (e.g., Ethernet).
[0079] The storage system 720 handles the storage and management of data for the computing device 700. The storage system 720 may be implemented by one or more non-transitory machine-readable media configured to store software (e.g., programs, code modules, data structures, instructions, etc.) and store data used for the execution of the software or data generated during the execution of the software.
[0080] In this example, the storage system 720 includes an operating system 722, one or more applications 724, an I / O module 726, and a communication module 728. The operating system 722 includes various programs, instruction sets, software components, and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitating communication between various hardware and software components. The operating system 722 may be one of the following: various versions of Microsoft Windows, Apple Mac OS, Apple OS X, Apple macOS, and / or Linux operating systems, various commercially available UNIX or UNIX-like operating systems (including but not limited to various GNU / Linux operating systems, Google OS, etc.), and / or mobile operating systems such as Apple iOS, Windows Phone, Windows Mobile, Android, BlackBerry OS, BlackBerry 10, and Palm OS, WebOS operating systems.
[0081] The application 724 may include any number of different applications installed on the computing device 700. Examples of such applications may include a browser application, an address book application, a contact list application, an email application, an instant messaging application, a word processing application, a JAVA-enabled application, an encryption application, a digital rights management application, a voice recognition application, a location determination application, a mapping application, a music player application, and the like.
[0082] The I / O module 726 manages information received via input components (e.g., the display 710, the sensor 712, and the microphone 716) and information to be output via output components (e.g., the display 710 and the speaker 714). The communication module 728 facilitates communication with other devices via the communication system 718 and includes various software components for processing data received from the communication system 718.
[0083] Those of ordinary skill in the art will recognize that Figure 7 the illustrated architecture is merely an example architecture of the computing device 700, and the computing device 700 may have more or fewer components than shown, or a different configuration of components. Figure 7 The various components shown may be implemented in hardware, software, firmware, or any combination thereof, including one or more signal processing and / or application specific integrated circuits.
[0084] Figure 8 An exemplary system 800 for implementing the various embodiments described above is shown. For example, the client devices 802 - 808 may be used to implement the client devices 105a - n, and the cloud computing system 812 may be used to implement the computing system 110. As shown, the system 800 includes client devices 802 - 808, one or more networks 810, and a cloud computing system 812. The cloud computing system 812 is configured to provide resources and data to the client devices 802 - 808 via the network 810. In some embodiments, the cloud computing system 800 provides resources to any number of different users (e.g., customers, tenants, organizations, etc.). The cloud computing system 812 may be implemented by one or more computer systems (e.g., servers), virtual machines running on the computer systems, or a combination thereof.
[0085] As shown, the cloud computing system 812 includes one or more applications 814, one or more services 816, and one or more databases 818. The cloud computing system 800 may provide the applications 814, services 816, and databases 818 to any number of different customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner.
[0086] In some embodiments, the cloud computing system 800 may be adapted to automatically provision, manage, and track customer subscriptions to services provided by the cloud computing system 800. The cloud computing system 800 may provide cloud services via different deployment models. For example, cloud services may be provided under a public cloud model, in which the cloud computing system 800 is owned by an organization that sells cloud services, and the cloud services are available to the public or different industry enterprises. As another example, cloud services may be provided under a private cloud model, in which the cloud computing system 800 runs only for a single organization and may provide cloud services to one or more entities within that organization. Cloud services may also be provided under a community cloud model, in which the cloud computing system 800 and the cloud services provided by the cloud computing system 800 are shared by several organizations in a related community. Cloud services may also be provided under a hybrid cloud model, which is a combination of two or more of the above different models.
[0087] In some cases, any of the applications 814, services 816, and databases 818 that are available to client devices 802 - 808 from the cloud computing system 800 via the network 810 are referred to as "cloud services". Generally, the servers and systems that make up the cloud computing system 800 are different from the customer's local servers and systems. For example, the cloud computing system 800 may host applications, and a user of one of the client devices 802 - 808 may order and use the applications via the network 810.
[0088] The application 814 may include software applications configured to execute on a cloud computing system 812 (e.g., a computer system or a virtual machine running on a computer system) and be accessed, controlled, managed, etc. via the client devices 802 - 808. In some embodiments, the application 814 may include server applications and / or middleware applications (e.g., HTTP (Hypertext Transfer Protocol) server applications, FTP (File Transfer Protocol) server applications, CGI (Common Gateway Interface) server applications, JAVA server applications, etc.). The service 816 is a software component, module, application, etc. configured to execute on the cloud computing system 812 and provide functionality to the client devices 802 - 808 via the network 810. The service 816 may be a network - based service or an on - demand cloud service.
[0089] The database 818 is configured to store and / or manage data accessed by the application 814, the service 816, and / or the client devices 802 - 808. For example, the storage devices 135 - 145 may be stored in the database 818. The database 818 may reside on a non - transitory storage medium local to (and / or within) the cloud computing system 812, in a storage area network (SAN), or on a non - transitory storage medium local to a location remote from the cloud computing system 812. In some embodiments, the database 818 may include a relational database managed by a relational database management system (RDBMS). The database 818 may be a column - oriented database, a row - oriented database, or a combination thereof. In some embodiments, some or all of the database 818 is an in - memory database. That is, in some such embodiments, the data of the database 818 is stored and managed in a memory (e.g., random access memory (RAM)).
[0090] The client devices 802 - 808 are configured to execute and run client applications (e.g., web browsers, proprietary client applications, etc.) that communicate via the network 810 with the application 814, the service 816, and / or the database 818. Thus, when the application 814, the service 816, and the database 818 are running (e.g., hosted) on the cloud computing system 800, the client devices 802 - 808 can access the various functions provided by the application 814, the service 816, and the database 818. The client devices 802 - 808 may be the computer system 600 or the computing device 700 as described above with reference to Figure 6 and Figure 7 respectively. Although the system 800 is shown as having four client devices, any number of client devices may be supported.
[0091] The network 810 can be any type of network configured to facilitate data communication between the client devices 802 - 808 and the cloud computing system 812 using any one of a variety of network protocols. The network 810 can be a personal area network (PAN), a local area network (LAN), a storage area network (SAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a global area network (GAN), an intranet, the Internet, a network of any number of different types of networks, etc.
[0092] The foregoing description illustrates various embodiments of the present invention and examples of how aspects of the present invention may be implemented. The above examples and embodiments should not be considered the only embodiments and are presented to illustrate the flexibility and advantages of the present invention as defined by the appended claims. Based on the foregoing disclosure and the following claims, other arrangements, embodiments, implementations, and equivalents will be apparent to those skilled in the art and may be employed without departing from the spirit and scope of the present invention as defined by the claims.
Claims
1. A non-transitory machine-readable medium storing a program executable by at least one processing unit of a computing device, the program including an instruction set for the following operations: Monitoring the utilization rate of a resource set by a resource consumer running on a computing device; Generate a model including a plurality of long short-term memory recurrent neural network (LSTM-RNN) layers and a set of attention mechanism layers based on the utilization rate of the resource set, where the model is configured to predict the future utilization rate of the resource set; And Based on the monitored utilization rate of the resource set and the model, determining a set of predicted values representing the utilization rate of the resource set by the resource consumer running on the computing device, Wherein, Monitoring the utilization rate of the resource set by a resource consumer running on a computing device includes: In each of a plurality of time intervals, measuring the utilization rate of the resource set by a resource consumer running on the computing device; and Storing the utilization rate measured in each of the plurality of time intervals according to a set of values of a set of metrics representing the utilization rate of the resource set by a resource consumer running on the computing device, Wherein, generating the model includes training the model using a set of values of a set of metrics measured in each time interval in a subset of the plurality of time intervals, Wherein, the attention mechanism layer receives n inputs y1, y2, y3,..., yn and a context C, Among them, the output of the LSTM-RNN layer h i is the input of the attention mechanism layer y i and the cell state C of the LSTM-RNN layer t is the input of the context C, where h i is the hidden state at time interval i Wherein, the attention mechanism layer outputs a vector z, and the vector z is a weighted arithmetic mean of the n inputs y1, y2, y3,..., yn, and Wherein, the weights are determined according to the correlation between the given context C and each yi of the n inputs y1, y2, y3,..., yn, Wherein, each input yi in the n inputs y1, y2, y3,..., yn is weighted by a corresponding weight si, Wherein, si = softmax(mi)i, mi = tanh(yi, C), and Among them, softmax(m)i is the softmax function used to evaluate the i-th component mi of the vector m, where m = (m1, m2, m3, …, m n ).
2. The non-transitory machine-readable medium according to claim 1, wherein, The program further includes an instruction set for the following operations: Calculating a set of error metrics based on a plurality of sets of predicted values representing the utilization rate of the resource set by a resource consumer running on the computing device and a plurality of corresponding sets of values of a set of metrics; Determining whether the value of an error metric in the set of error metrics is greater than a defined threshold; And When it is determined that the value of an error metric in the set of error metrics is greater than a defined threshold, updating the model.
3. The non-transitory machine-readable medium according to claim 2, wherein, Updating the model includes: Training the model using a set of values of a set of metrics measured in each time interval in a set of most recent time intervals; and Storing the updated model in a storage device.
4. The non-transitory machine-readable medium according to claim 1, wherein, The program further includes an instruction set for adjusting the allocation of resources in the resource set based on the set of predicted values.
5. The non-transitory machine-readable medium according to claim 1, wherein, The program further includes an instruction set for sending a notification to a client device to warn of high utilization of resources in the resource set.
6. A method executable by a computing device, including: Monitoring the utilization rate of a resource set by a resource consumer running on a computing device; Based on the monitored utilization rate of the resource set, generating a model including a plurality of long short-term memory recurrent neural network LSTM-RNN layers and a set of attention mechanism layers, the model being configured to predict the future utilization rate of the resource set; And Based on the monitored utilization rate of the resource set and the model, determine a set of predicted values representing the utilization rate of the resource set by resource consumers running on the computing device. Wherein, monitoring the utilization rate of the resource set by resource consumers running on the computing device includes: In each of a plurality of time intervals, measuring the utilization rate of the resource set by resource consumers running on the computing device; and Storing the utilization rate measured in each of the plurality of time intervals according to a set of values of a set of metrics representing the utilization rate of the resource set by resource consumers running on the computing device. Wherein, generating the model includes training the model using a set of values of a set of metrics measured in each time interval in a subset of the plurality of time intervals. Wherein, the attention mechanism layer receives n inputs y1, y2, y3, ……, yn and a context C. Among them, the output of the LSTM-RNN layer h i is the input of the attention mechanism layer y i and the cell state C t of the LSTM-RNN layer is the input of the context C, where h i is the hidden state at time interval i Wherein, the attention mechanism layer outputs a vector z, and the vector z is the weighted arithmetic mean of the n inputs y1, y2, y3, ……, yn, and Wherein, the weights are determined according to the correlation between the given context C and each yi of the n inputs y1, y2, y3, ……, yn. Wherein, each input yi in the n inputs y1, y2, y3, ……, yn is weighted by a corresponding weight si. Wherein, si = softmax(mi), mi = tanh(yi, C), and Among them, softmax(m)i is the softmax function used to evaluate the i-th component mi of the vector m, where m = (m1, m2, m3, …, m n ).
7. The method according to claim 6, further comprising: Calculating a set of error metrics based on a set of multiple predicted values representing the utilization rate of the resource set by resource consumers running on the computing device and a set of corresponding values of a set of metrics. Determining whether the value of an error metric in the set of error metrics is greater than a defined threshold; And When it is determined that the value of an error metric in the set of error metrics is greater than the defined threshold, updating the model.
8. The method according to claim 7, wherein Updating the model includes: Training the model using a set of values of a set of metrics measured in each time interval in a set of most recent time intervals; and Storing the updated model in a storage device.
9. The method according to claim 6, wherein The method further includes adjusting the allocation of resources in the resource set based on the set of predicted values.
10. The method according to claim 6, wherein, The method further includes sending a notification to a client device to warn of high utilization of resources in the resource set.
11. A system, comprising: A set of processing units; And A non-transitory machine-readable medium that stores instructions that, when executed by at least one processing unit in the set of processing units, cause the at least one processing unit to: Monitor the utilization rate of a resource set by resource consumers running on the system; Based on the monitored utilization rate of the resource set, generate a model including a set of multiple long short-term memory recurrent neural network (LSTM-RNN) layers and an attention mechanism layer, the model being configured to predict future utilization of the resource set; And Based on the monitored utilization rate of the resource set and the model, determine a set of predicted values representing the utilization rate of the resource set by resource consumers running on the system. Among them, monitoring the utilization rate of the resource consumer running on the system for the resource set includes: In each of a plurality of time intervals, measuring the utilization rate of the resource consumer running on the system for the resource set; and Storing the utilization rate measured in each of the plurality of time intervals according to a set of values of a set of metrics representing the utilization rate of the resource consumer running on the system for the resource set, Among them, generating the model includes training the model using a set of values of a set of metrics measured in each of a subset of the plurality of time intervals, Among them, the attention mechanism layer receives n inputs y1, y2, y3, ……, yn and context C, where the output of the LSTM-RNN layer h i is the input of the attention mechanism layer y i and the cell state C t of the LSTM-RNN layer is the input of context C, where h i is the hidden state at time interval i Among them, the attention mechanism layer outputs a vector z, and the vector z is the weighted arithmetic mean of n inputs y1, y2, y3, ……, yn, and Among them, the weights are determined according to the relevance between the given context C and each yi of the n inputs y1, y2, y3, ……, yn, Among them, each input yi in the n inputs y1, y2, y3, ……, yn is weighted by a corresponding weight si, Among them, si = softmax(mi), mi = tanh(yi, C), and Among them, softmax(m)i is the softmax function used to evaluate the i-th component mi of the vector m, where m = (m1, m2, m3, …, m n ).
12. The system according to claim 11, wherein The instruction also causes the at least one processing unit: Calculating a set of error metrics based on a set of multiple predicted values representing the utilization rate of the resource consumer running on the system for the resource set and a set of corresponding values of a set of metrics; Determining whether the value of an error metric in the set of error metrics is greater than a defined threshold; And When it is determined that the value of an error metric in the set of error metrics is greater than the defined threshold, updating the model.
13. The system according to claim 11, wherein, Updating the model includes: Training the model using a set of values of a set of metrics measured in each of the most recent set of time intervals; and Storing the updated model in a storage device.
14. The system according to claim 11, wherein, The instruction also causes the at least one processing unit to adjust the allocation of resources in the resource set based on the set of predicted values.
Citation Information
Patent Citations
Kubernetes dispatching optimization method based on neural network
CN108874542A
Flood prediction method based on an attention model long-short-term memory network
CN109583565A