Host load prediction method and device

By combining the SARIMA model and the Transformer model to predict the host load, the problem of insufficient model accuracy in the cloud server system is solved, and efficient resource management and energy optimization are achieved.

CN120336114APending Publication Date: 2025-07-18SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510338184.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the cloud server system with high complexity, long-term dependence and large data scale, it is difficult to learn shallow single linear features of time series through the SARIMA model, and effectively capture the non-single linear timing features and long-distance dependence in the residual, resulting in insufficient host load prediction accuracy.

Method used

The SARIMA model is used to learn shallow single linear features of the time series, and the model residuals are recursively modeled through Transformer until the residuals pass the white noise test. Finally, the output of the SARIMA model and the output of multiple Transformer models are summed as the prediction result of the host load at the future moment.

Benefits of technology

It improves the accuracy of host load prediction, reduces business decision-making risks caused by insufficient prediction accuracy, helps cloud computing systems optimize resource allocation, reduces energy consumption and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336114A_ABST
    Figure CN120336114A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cloud computing, and particularly provides a host load prediction method and device, and the method comprises the following steps: S1, collecting historical workload data; s2, according to channel-index setting, the multivariate time sequence is disassembled into C unit time sequences, and each unit time sequence independently executes the following steps; s3, performing normalization processing on the historical workload data; s4, constructing training data by adopting a sliding window method; s5, an SARIMA model is constructed; s6, calculating a residual error of the SARIMA model; s7, carrying out Transform modeling on the residual error recursion; and S8, the sum of the output of the SARIMA model and the outputs of the plurality of Transform models is used as a prediction result of the host load at the future moment. Compared with the prior art, the resource management system can be helped to better and automatically allocate resources to adapt to the change of the workload of each server, and the high-level service quality is maintained while the energy consumption is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing, and particularly provides a host load prediction method and device. Background Art

[0002] One of the main characteristics of a cloud computing system is elasticity, that is, the resource management system can automatically allocate resources to adapt to the changes in the workloads of each server. Through reasonable prediction of the host load, on the one hand, cloud service providers can prevent potential undersupply during service resource scheduling, avoid problems such as performance degradation, increased latency, or system crashes caused by excessive host load, ensure the continuity and stability of the service, reduce the risk of violating the Service Level Agreement (SLA), and improve the user experience; on the other hand, they can promptly handle over-supply, reduce unnecessary hardware investment and maintenance costs, improve the resource utilization rate of the data center, and reduce energy consumption. Through reasonable prediction of the host load, on the one hand, cloud service providers can prevent potential undersupply during service resource scheduling, avoid problems such as performance degradation, increased latency, or system crashes caused by excessive host load, ensure the continuity and stability of the service, reduce the risk of violating the Service Level Agreement (SLA), and improve the user experience; on the other hand, they can promptly handle over-supply, reduce unnecessary hardware investment and maintenance costs, improve the resource utilization rate of the data center, and reduce energy consumption.

[0003] As a key technology for cluster resource management, the load prediction method collects data on resource information such as CPU and memory for each server at regular intervals on the premise that the server is running normally. By analyzing the historical data collected from the cloud data center, it grasps the trend and variation law of the load data, so as to predict the load value of the next cycle. The load prediction of cloud computing resources is a typical time series prediction problem, and establishing an accurate model is the focus of the research work.

[0004] In the research of related fields, the solutions for host load prediction have undergone a transformation from traditional statistical methods to machine learning techniques and then to deep learning-based approaches. For example, the patent "A Host Load Prediction Method Based on Long Short-Term Memory Network" (CN106502799A) uses an LSTM network; the patent "A Host Load Prediction Method in a Cloud Environment" (CN108196957A) uses an ARMA (Auto-Regressive Moving Average) model; the patent "A Cloud Computing Host Load Prediction Method Combining Attention Mechanism and Gated Recurrent Unit" (CN113076196A) uses an attention mechanism and a GRU (Gate Recurrent Unit).

[0005] Meanwhile, as the sequence modeling architecture Transformer has shone in various natural language processing tasks, the number of Transformer-based time series solutions has been surging. For example, the paper "Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting, NeurIPS 2019" proposed the LogTrans model; the paper "Informer: Beyond efficient transformer for long sequence time-series forecasting, AAAI 2021" proposed the Informer model; the paper "Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting, NeurIPS 2021" proposed the Autoformer model; the paper "Pyraformer: Low-complexity pyramidal attention for long-range time series modeling and forecasting, ICLR 2022 Oral" proposed the Pyraformer model; the paper "Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting, ICML 2022" proposed the FEDformer model.

[0006] For a cloud server system with characteristics of high complexity, long-term dependence, and large data scale, how to learn the shallow single linear features of time series through the SARIMA model, and then use Transformer to recursively fit the model residuals, effectively improving the prediction performance of the model is an urgent problem for those skilled in the art. Summary of the Invention

[0007] The present invention aims at the deficiencies of the above-mentioned prior art and provides a host load prediction method with strong practicability.

[0008] A further technical task of the present invention is to provide a host load prediction device with reasonable design, safety and applicability.

[0009] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0010] A host load prediction method has the following steps:

[0011] S1. Collect historical workload data;

[0012] S2. Decompose the multivariate time series into C unit time series according to the channel-independence setting, and each unit time series independently executes the following steps;

[0013] S3. Normalize the historical workload data;

[0014] S4. Construct training data by using the sliding window method;

[0015] S5. Construct a SARIMA model;

[0016] S6. Calculate the residuals of the SARIMA model and perform white noise test to evaluate whether the residual data is completely random or has any identifiable patterns or trends;

[0017] S7. Recursively perform Transformer modeling on the residuals until the model residuals pass the white noise test;

[0018] S8. Add the output of the SARIMA model and the outputs of multiple Transformer models as the prediction result of the host load at future moments.

[0019] Further, in step S1, obtain the historical workload data of M hosts with a look-back window L through the resource monitoring system of the cloud computing service cluster center, including C dimensions; for host m, m ∈ {1, 2,..., M}, the historical data constitutes a multivariate time series sample set of length L:

[0020]

[0021] Among them, represents the load of host m in the i-th dimension at time t.

[0022] Furthermore, in step S3, the Min-Max normalization method is used to process the original sequence, and the processed data conforms to a minimum value of 0 and a maximum value of 1. The calculation formula of Min-Max normalization is as follows:

[0023]

[0024] where t = 1, 2,..., L, m = 1, 2,..., M, i = 1, 2,..., C, is the maximum value of the historical data of the host load in the i-th dimension, is the minimum value of the historical data of the host load in the i-th dimension.

[0025] Furthermore, in step S4, the sliding window method is adopted to construct L-S groups of training data with a learning length of s and a prediction length of 1. In the n-th group of training data is the feature, is the label.

[0026] Furthermore, in step S5, the SARIMA model eliminates the local level or trend of the sequence through the differencing method, and the time series is Its mathematical expression is:

[0027]

[0028] where φ1, φ2,..., φ p are autoregressive coefficients, Φ1, Φ2,..., Φ P are seasonal autoregressive coefficients, θ1, θ2,..., θ q are moving average coefficients, Θ1, Θ2,..., Θ Q are seasonal moving average coefficients, B is the differencing operator, B s = B s is the seasonal delay operator with a period equal to s, d is the number of differencing times, D is the number of seasonal differencing times, ∈ t is the error term;

[0029] The hyperparameter s in the SARIMA model is determined through periodicity tests, d and D are determined through stationarity tests, and p, q, P, Q are determined through the BIC criterion.

[0030] Further, in step S6, if the residual sequence passes the pure randomness test, it means that there is no autocorrelation between the values in the sequence, that is, the occurrence of each value does not depend on the previous value sequence; otherwise, it is considered that the residual sequence has a trend or periodic structure that has not been extracted by the SARIMA model, and other models need to be used to capture it.

[0031] Further, in step S7, the parameter space of the Transformer is large, which can capture non-single linear time series features and better capture long-distance dependence relationships through the self-attention mechanism.

[0032] Further, in step S8, the output of each Transformer model is a compensation for the residuals of the previous model. Finally, the sum of the output of the SARIMA model and the outputs of multiple Transformer models is used as the prediction result of the host load at future moments.

[0033] A host load prediction device includes: at least one memory and at least one processor;

[0034] The at least one memory is used to store machine-readable programs;

[0035] The at least one processor is used to call the machine-readable program and execute a host load prediction method.

[0036] Compared with the prior art, a host load prediction method and device of the present invention have the following outstanding beneficial effects:

[0037] The present invention accelerates the convergence in the training process of the deep learning algorithm through the normalization preprocessing module; learns the shallow single linear features of the time series through the SARIMA model; captures the remaining non-single linear time series features and long-distance dependence relationships in the residuals through the Transformer, which can effectively improve the prediction effect and reduce the business decision-making risk caused by insufficient model prediction accuracy.

[0038] The method of the present invention can predict the host load in real time and accurately, and help the resource management of the cloud computing system to automatically allocate resources to adapt to the changes in the workload of each server. On the one hand, it prevents potential supply shortages and avoids problems such as performance degradation, increased latency, or system crashes caused by high host loads, ensures the continuity and stability of services, reduces the risk of violating service level agreements, and improves the user experience; on the other hand, it can timely handle over-supply, reduce unnecessary hardware investment and maintenance costs, and reduce energy consumption. Description of the Drawings

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Attached Figure 1 is a flowchart of a method for predicting host load;

[0041] Attached Figure 2 is a flowchart of the SARIMA model in a method for predicting host load. Detailed implementation manners

[0042] To enable those skilled in the art to better understand the solution of the present invention, the following further details the present invention in combination with specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0043] The following gives a best embodiment:

[0044] As Figure 1 shown, the framework diagram of the host load prediction method based on the SARIMA model and Transformer residual compensation includes processes such as historical workload data collection, multivariate time series decomposition, normalization preprocessing, sliding window construction of training data, SARIMA modeling, Transformer residual modeling, and calculation of prediction results at future moments.

[0045] A method for predicting host load in this embodiment has the following steps:

[0046] S1. Collect historical workload data;

[0047] In this implementation, historical workload data of M hosts with a look-back window b is obtained through the resource monitoring system of the cloud computing service cluster center, including a total of C dimensions such as CPU load sequence, memory load sequence, and disk I / O load sequence. For host m, m ∈ {1, 2,..., M}, this historical data constitutes a multivariate time series sample set of length L:

[0048]

[0049] Among them, represents the load of host m at the i-th dimension at time t.

[0050] S2. Decompose the multivariate time series into C unit time series according to the channel - independence setting, and each unit time series independently performs the following steps.

[0051] S3. Normalize the historical workload data;

[0052] The values of the historical load data vary greatly in different time intervals, and it is necessary to pre - process the original data by normalization. The pre - processed data can accelerate the convergence in the subsequent training process of deep learning algorithms. In this implementation, the Min - Max normalization method is used to process the original sequence, and the processed data meets the minimum value of 0 and the maximum value of 1.

[0053] The calculation formula of Min - Max normalization is as follows:

[0054]

[0055] where \(t = 1,2,\cdots,L\), \(m = 1,2,\cdots,M\), \(i = 1,2,\cdots,C\), is the maximum value of the historical data of the host load in the \(i\) - th dimension, is the minimum value of the historical data of the host load in the \(i\) - th dimension.

[0056] S4. Construct training data using the sliding window method;

[0057] Use the sliding window method to construct \(L - S\) groups of training data with a learning length of \(s\) and a prediction length of 1. In the \(n\) - th group of training data, is the feature, is the label.

[0058] S5. Construct a SARIMA model;

[0059] As shown in the appendix Figure 2 Construct a SARIMA model, including processes such as periodicity test, stationarity test, differencing, order determination by BIC criterion, and parameter estimation for modeling. The SARIMA model is a time - series prediction method developed from the ARMA model and is used to solve the prediction problem of non - stationary data in the actual production environment. The SARIMA model considers the influence of periodic factors and eliminates the local level or trend of the sequence through differencing. Taking the time series as an example, its mathematical expression is:

[0060]

[0061] where \(\varphi_1,\varphi_2,\cdots,\varphi\) p are autoregressive coefficients (AR parameters), \(\Phi_1,\Phi_2,\cdots,\Phi\) Pare seasonal autoregressive coefficients, θ1, θ2, ..., θ q are moving average coefficients (MA parameters), Θ1, Θ2, ..., Θ Q are seasonal moving average coefficients, B is the backshift operator, B s = B s is the seasonal backshift operator with a period equal to s, d is the order of differencing, D is the order of seasonal differencing, ∈ t is the error term.

[0062] The hyperparameter s in the SARIMA model is determined through periodicity tests, d and D are determined through stationarity tests, and p, q, P, Q are determined through the BIC criterion (Bayesian Information Criterion). There are multiple methods for periodicity tests and stationarity tests. The former includes Fourier transform and autocorrelation coefficient, etc., and the latter includes unit root test and KPSS test, etc.

[0063] S6. Calculate the residuals of the SARIMA model and conduct a white noise test to evaluate whether the residual data is completely random or has any identifiable patterns or trends;

[0064] If the residual sequence passes the pure randomness test, it means there is no autocorrelation between the values in the sequence, that is, the occurrence of each value does not depend on the previous values in the sequence. Otherwise, it can be considered that there are trends, periodicities, or other structures in the residual sequence that have not been extracted by the SARIMA model and more complex models are needed to capture them. There are multiple methods for white noise tests, including autocorrelation function test, turning point test, and mixed test (Q statistic test), etc.

[0065] S7. Recursively model the residuals using Transformer until the model residuals pass the white noise test; compared with the SARIMA model, Transformer has a larger parameter space and can capture non-single linear time series features. In addition, Transformer can better capture long-distance dependencies through the self-attention mechanism, effectively avoiding the problems of gradient disappearance or gradient explosion existing in neural network methods, and is suitable for the long sequence modeling requirement of host load prediction.

[0066] S8. The output of each Transformer model is the compensation for the residual of the previous model. Finally, the sum of the output of the SARIMA model and the outputs of multiple Transformer models is used as the prediction result of the host load at future moments.

[0067] Based on the above method, a host load prediction device in this embodiment includes: at least one memory and at least one processor;

[0068] The at least one memory is used to store machine-readable programs;

[0069] The at least one processor is used to call the machine-readable program and execute a host load prediction method.

[0070] The above specific embodiments are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific embodiments. Any technical solution that conforms to the technical solutions described in the above specific embodiments of the present invention and any appropriate changes or substitutions made by those of ordinary skill in the art shall fall within the patent protection scope of the present invention.

[0071] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A host load prediction method, characterized in that, It has the following steps: S1. Collect historical workload data; S2. Decompose the multivariate time series into C unit time series according to the channel-independence setting, and each unit time series independently executes the following steps; S3. Normalize the historical workload data; S4. Use the sliding window method to construct training data; S5. Construct a SARIMA model; S6. Calculate the residuals of the SARIMA model and perform a white noise test to evaluate whether the residual data is completely random or has any identifiable patterns or trends; S7. Recursively model the residuals with a Transformer until the model residuals pass the white noise test; S8. Sum the outputs of the SARIMA model and the outputs of multiple Transformer models as the prediction result of the host load at future moments.

2. The host load prediction method according to claim 1, wherein In step S1, obtain the historical workload data of M hosts with a lookback window length L through the resource monitoring system of the cloud computing service cluster center, including C dimensions; for host m, m ∈ {1, 2,..., M}, this historical data constitutes a multivariate time series sample set with a lookback window length of L: Among them, represents the load of the host m in the i-th dimension at time t.

3. A host load prediction method according to claim 2, characterized in that, In step S3, use the Min-Max normalization method to process the original sequence. The processed data conforms to a minimum value of 0 and a maximum value of 1. The calculation formula for Min-Max normalization is as follows: where t = 1, 2, ..., L, m = 1, 2, ..., M, i = 1, 2, …, C, is the maximum value of the historical data of the host load in the i-th dimension, is the minimum value of the historical data of the host load in the i-th dimension.

4. The host load prediction method according to claim 3, wherein In step S4, the L-S group training data with a learning length of s and a prediction length of 1 is constructed by using the sliding window method. In the nth group of training data, is a feature, is a label.

5. A method for predicting the host load according to claim 3, characterized in that, In step S5, the SARIMA model eliminates the local level or trend of the sequence through the differencing method, and the time series is Its mathematical expression is: Among them, φ1, φ2, …, φ p are autoregressive coefficients, Φ1, φ2, …, φ P are seasonal autoregressive coefficients, θ1, θ2, …, θ q are moving average coefficients, Θ1, Θ2, …, Θ Q are seasonal moving average coefficients, B is the difference operator, B s = B s is the seasonal delay operator with a period equal to s, d is the number of differences, D is the number of seasonal differences, ∈ t is the error term; The hyperparameter s in the SARIMA model is determined through periodicity tests, d and D are determined through stationarity tests, and p, q, P, Q are determined through the BIC criterion.

6. A host load prediction method according to claim 5, characterized in that In step S6, if the residual sequence passes the pure randomness test, it means that there is no autocorrelation between the values in the sequence, that is, the occurrence of each value does not depend on the sequence of previous values; otherwise, it is considered that the residual sequence has a trend or periodic structure that has not been extracted by the SARIMA model and needs to be captured by other models.

7. A method for predicting host load according to claim 6, characterized in that, In step S7, the Transformer has a large parameter space, captures non-single linear time series features, and better captures long-distance dependencies through the self-attention mechanism.

8. A host load prediction method according to claim 7, characterized in that, In step S8, the output of each Transformer model is a compensation for the residuals of the previous model. Finally, sum the outputs of the SARIMA model and the outputs of multiple Transformer models as the prediction result of the host load at future moments.

9. A host load prediction device, characterized in that, It includes: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program and execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Host load prediction method based on long and short term memory network

    CN106502799A

  • Host load prediction method under cloud environment

    CN108196957A

  • Cloud computing host load prediction method combining attention mechanism and gating circulation unit

    CN113076196A