Edge Prediction Method Based on Deep Autoregressive Recurrent Neural Network

Through the edge prediction method based on deep autoregressive recurrent neural network, the problem of unbalanced load distribution in edge computing systems is solved, the probability distribution prediction of load is realized, the efficiency and accuracy of resource scheduling are improved, and it is suitable for resource configuration of edge computing systems.

CN115509752BActive Publication Date: 2025-06-27FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211200915.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-06-27
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

Existing load prediction technologies cannot effectively support the probability distribution prediction of loads in edge computing systems, resulting in unbalanced resource configuration and waste, and the load distribution of different edge servers varies greatly, affecting the accuracy of the prediction model.

Method used

The edge prediction method based on deep autoregressive recurrent neural network is adopted. By obtaining historical edge load data, building data sets, preprocessing, and using the edge server users as covariates, the deep autoregressive recurrent neural network model is used to predict loads, output the probability distribution of future loads, and assisting edge computing services in formulating resource scheduling solutions.

Benefits of technology

The load balancing of edge computing systems is realized, the efficiency and accuracy of resource allocation are improved, and the prediction effect of high accuracy and high confidence can be maintained at different prediction lengths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115509752B_ABST
    Figure CN115509752B_ABST
Patent Text Reader

Abstract

The present invention relates to an edge prediction method based on a deep regression recurrent neural network, comprising the following steps: Step S1: Obtain historical edge load data and construct a data set; Step S2: Preprocess the data set; Step S3: Take the users of the edge server as covariates; Step S4: Input the preprocessed data set and covariates into a prediction model to predict future edge loads; Step S5: Based on the edge load prediction results, assist the edge computing service in formulating a resource scheduling plan. The present invention realizes scalable and effective workload prediction, and effectively improves the resource allocation efficiency in cloud computing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer resource scheduling, and particularly to an edge prediction method based on a deep autoregressive recurrent neural network. Background Art

[0002] As an emerging computing paradigm in the Internet of Things era, edge computing can effectively reduce the response time of IoT applications and thus improve the user service experience. Load prediction is an important supporting technology in edge computing. By pre-configuring and allocating resources, reasonable and efficient resource supply can be achieved. For example, when a large number of service requests arrive at the edge server simultaneously, insufficient resource supply is likely to lead to an increase in service request time. When the edge server processes a small number of service requests for a long time, over-allocation of resources results in waste of resources. If it is predicted that the edge load will be at a low level in the future, the corresponding resource allocation is reduced; otherwise, the corresponding resource allocation is increased. Therefore, edge load prediction can better ensure the service level agreement (SLA) and effectively improve the reliability of the edge computing system.

[0003] Compared with the cloud data center with centralized management, the deployment of edge servers is more dispersed, and the load distribution of different edge servers is highly uneven. The maximum load gap between cross-site edge servers reaches 19.8 times, and the maximum load gap between edge servers at the same site is also 14.3 times. This uneven distribution of edge load data samples leads to different degrees of influence of different time series on the prediction model, and also brings challenges to the prediction of edge load. Therefore, this sample-level difference problem needs to be solved. In addition, most of the existing load prediction works only focus on single-point real-value prediction and do not support the prediction of load probability distribution. However, in many actual edge computing scenarios, obtaining the probability distribution of future load changes is more valuable in application than directly predicting the real value of future load. This is because the prediction of the probability distribution of edge load is more meaningful for grasping the future changes of edge load and is more helpful for the edge computing system to flexibly allocate resource supply. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide an edge prediction method based on a deep autoregressive recurrent neural network, which can efficiently schedule resources such as computing and storage according to the current and future load conditions, so as to achieve load balancing of the system.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] An edge prediction method based on a deep autoregressive recurrent neural network, comprising the following steps:

[0007] Step S1: Obtain historical edge load data and construct a data set;

[0008] Step S2: Preprocess the dataset;

[0009] Step S3: Take the users of the edge servers as covariates;

[0010] Step S4: Input the preprocessed dataset and covariates into the prediction model to predict the future edge load;

[0011] Step S5: Based on the edge load prediction results, assist the edge computing service in formulating a resource scheduling plan.

[0012] Furthermore, the historical edge load data includes the CPU usage rate, specifically:

[0013] For the load of an edge server i, z i,t represents the change in the CPU usage rate of edge server i over a period of time t, then the historical load sequence of this edge server is expressed as:

[0014]

[0015] The future predicted load sequence is expressed as

[0016] where t0 represents the starting point of load prediction, T is the total length of the load sequence, [1, t0 - 1] represents the time range of the historical load sequence, and [t0, T] represents the time range of the future predicted load sequence.

[0017] Furthermore, the data preprocessing includes data cleaning and resampling.

[0018] Furthermore, there are differences in the CPU usage rates between servers, so the mean value of the load data of each edge server is used as the scaling factor v i , and the load is divided by v when inputting into the prediction model i , and the corresponding load is multiplied by v when outputting from the prediction model i

[0019]

[0020] Furthermore, the specific content of Step S3 is as follows:

[0021] The scaled historical load and covariates are used as the input of the prediction model to predict the future edge load. The prediction model can model the future load situation based on the known historical load and obtain the probability distribution of the future load, which is defined as

[0022]

[0023] Furthermore, the prediction model is an edge load prediction model based on a deep autoregressive recurrent neural network, adopting a sequence-to-sequence architecture, including an encoder and a decoder. The encoder and the decoder adopt the same network structure and share weights. The initial inputs of the encoder ( and z i,0 ) are both initialized to 0. Within the time interval [1, t0 - 1], the encoder calculates to obtain and uses it as the initial input of the decoder.

[0024] Furthermore, the prediction model uses LSTM to extract temporal features. The input is the edge load data and covariates within a past period of time, and the goal is to predict the probability distribution of the edge load z i,t at each time step. The probability distribution is defined as

[0025]

[0026] where h i,t represents the output of an LSTM; taking the observed value z i,t-1 at the previous time step and the LSTM output h i,t-1 as inputs, h i,t can be calculated as

[0027] h i,t = h(h i,t-1 , z i,t-1 , x i,t , Θ h ) (Equation 8)

[0028] where h represents a neural network with a multi-layer LSTM structure, and Θ h is the set of parameters in the LSTM; the edge load data z i,1:t-1 at the previous time step and the hidden layer output h i,t-1 will be used to calculate the network output h i,t at the current time step; the likelihood function is a probability distribution, and its set of parameters, including the mean μ and variance σ, is calculated through , and represents the mapping from h i,t to the set of parameters of the likelihood function.

[0029] Furthermore, the training of the prediction model is as follows:

[0030] The input at each time step includes the covariates the load value at the previous time step and the network output at the previous time step Through training, the network output Furthermore, calculate the parameters of the likelihood function l(z|θ).

[0031] Next, optimize the loss function through the Stochastic Gradient Descent with Adaptive Momentum method.

[0032]

[0033] Where N is the total number of the edge load data sequences, is the parameter set of the likelihood function, and z i,t is the true load value.

[0034] The present invention has the following beneficial effects compared with the prior art:

[0035] 1. The present invention can efficiently schedule resources such as computing and storage according to the current and future load conditions, thereby achieving the load balance of the system;

[0036] 2. The present invention can achieve better prediction effects for different prediction lengths by selecting different probability distribution functions, and shows more excellent prediction accuracy under different prediction lengths, and obtains higher confidence levels within different prediction intervals. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a schematic flow chart of the method of the present invention;

[0038] Figure 2 is the training and prediction process of ELP-DAR in an embodiment of the present invention, where (a) is the training process and (b) is the prediction process;

[0039] Figure 3 is the LSTM cell structure in an embodiment of the present invention;

[0040] Figure 4 is the prediction effect of different methods when the prediction length is 3 days in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0041] The present invention will be further described below with reference to the drawings and embodiments.

[0042] Please refer to Figure 1 , the present invention provides an edge prediction method based on a deep autoregressive recurrent neural network, including the following steps:

[0043] Step S1: Obtain historical edge load data and construct a data set;

[0044] Step S2: Preprocess the data set;

[0045] Step S3: Use the users of the edge server as covariates;

[0046] Step S4: Input the preprocessed dataset and covariates into the prediction model to predict the future edge load;

[0047] Step S5: Based on the edge load prediction results, assist the edge computing service in formulating a resource scheduling plan.

[0048] In this embodiment, the CPU usage rate is regarded as the main prediction index, specifically:

[0049] For the load of an edge server i, z i,t represents the change in the CPU usage rate of edge server i over a period of time t. Then the historical load sequence of this edge server is expressed as:

[0050]

[0051] The future predicted load sequence is expressed as

[0052] where t0 represents the starting point of load prediction, T is the total length of the load sequence, [1, t0 - 1] represents the time range of the historical load sequence, and [t0, T] represents the time range of the future predicted load sequence.

[0053] In this embodiment, data preprocessing includes data cleaning and resampling.

[0054] In this embodiment, the users using the edge server are used as covariates to improve the learning effect of the edge load prediction model, because the introduction of covariates can help the prediction model better capture the correlation between sequences. Specifically:

[0055] Encode the user ID as a covariate, denoted as [x i,1 , x i,2 , …, x i,T : = x i,1 : T . In the edge computing system, there is a 10 - to 20 - fold difference in CPU usage rate between servers. To address the above data span difference and thus establish an accurate edge load prediction model, we use the mean of each edge server load data as the scaling factor v i (divide the load by v when inputting into the prediction model i , and multiply the corresponding load by v when outputting from the prediction model i ). This scaling factor v i is defined as

[0056]

[0057] In this embodiment, step S3 is specifically as follows:

[0058] The scaled historical load and the covariates will be used as the input of the prediction model to predict the future edge load. The prediction model can model the future load situation based on the known historical load and obtain the probability distribution of the future load. This probability distribution is defined as

[0059]

[0060] Finally, the edge load prediction result will assist the edge computing service provider to formulate a suitable resource scheduling scheme to achieve the load balance of the edge computing system.

[0061] In this embodiment, a novel ELP-DAR method is proposed to achieve more accurate probability distribution prediction for time series problems. To evaluate the accuracy of edge load prediction, we introduce performance metrics such as RMSE, ND, and mean wQuantileLoss. Their specific definitions are

[0062]

[0063]

[0064]

[0065] where, represents the predicted value of edge server i within a period of time t, and N is the total number of time series.

[0066]

[0067] where, is the τ-quantile of the prediction result distribution. wQuantileLoss selects a suitable quantile according to the expected value of positive or negative errors. For example, when τ = 0.25, wQuantileLoss will give more penalties to overestimated predicted values. And mean wQuantileLoss is the average value of wQuantileLoss when τ takes values from 0.1 to 0.9.

[0068] In this embodiment, the edge load prediction model (ELP-DAR) based on the deep autoregressive recurrent neural network has steps as shown in Algorithm 1.

[0069]

[0070]

[0071] In this embodiment, ELP-DAR uses LSTM to extract temporal features, and its input is the edge load data and covariates over a period of time in the past. The goal of the ELP-DAR method is to predict the edge load z at each time step i,t The probability distribution of. Based on (Equation 2), the above probability distribution is defined as

[0072]

[0073] where h i,t represents the output of an LSTM. The observed value z at the previous moment i,t-1 and the LSTM output h i,t-1 are used as inputs, and through calculation, h i,t is obtained as

[0074] h i,t = h(h i,t-1 , z i,t-1 , x i,t , Θ h ) (Equation 8)

[0075] where h represents a neural network with a multi-layer LSTM structure, and Θ h is the set of parameters in the LSTM. The edge load data z at the previous moment i,1:t-1 and the hidden layer output h i,t-1 will be used to calculate the network output h at the current moment i,t . The likelihood function is a probability distribution, and its set of parameters, including the mean μ and variance σ, etc., is calculated through , while realizes the mapping from h i,t to the set of parameters of the likelihood function.

[0076] Overall, ELP-DAR is a Sequence-to-Sequence (S2S) architecture, including an encoder and a decoder. Specifically, the encoder and decoder in ELP-DAR adopt the same network structure and share weights between them. The initial inputs of the encoder ( and z i,0 ) are both initialized to 0. In the time interval [1, t0 - 1], the encoder calculates and uses it as the initial input of the decoder.

[0077] Figure 2Shows the training and prediction processes of the ELP-DAR method. The encoder in the training and prediction processes is the same. In the model training stage, all load data are known. The encoder sequentially inputs the load data in the time interval [1, t0-1] into the LSTM and obtains the output of the last hidden layer. And uses it as the input to the decoder.

[0078] The training process of the ELP-DAR method is as Figure 2 shown in (a). The input at each time step includes covariates the load value at the previous moment and the network output at the previous moment Through training, the network output at the current moment is obtained and then the parameters of the likelihood function l(z|θ) are calculated Next, the loss function is optimized by the Adaptive Moment (Adam) stochastic optimization method

[0079]

[0080] where N is the total number of the marginal load data sequences, is the set of parameters of the likelihood function, z i,t is the true load value.

[0081] The prediction process of the ELP-DAR method is as Figure 2 shown in (b). Since the load data in the time interval [t0, T] are unknown, the load value at the previous time step cannot be directly input at the current time step We use the random sampling method to obtain the sampled load value and use it as the input for the next time step. Through continuous iteration, the load prediction results in the time interval [t0, T] can be obtained, and corresponding indicators such as quantiles and expectations are further calculated through the sampled values.

[0082] Figure 3 Shows the LSTM cell structure in ELP-DAR. The LSTM cell uses three gates to control the information flow into the cell, including the input gate i t , the forget gate f t and the output gate o t . At the same time, the LSTM cell also saves the current cell state and the cell state c at the previous moment t , which is defined as

[0083]

[0084] where the forget gate f tDetermines the information discarded from the cell state at the previous moment, c t-1 Is the cell state at time t-1. The input gate i t Determines the new information to be stored in the current cell state, Represents the candidate cell state, which is defined as

[0085]

[0086] Where, W c Represents the connection weights between the previous hidden layer and the inputs z t 、x t And b c Represents the corresponding bias.

[0087] The forget gate is defined as

[0088] f t =sig(W f [z t ,x t ,h t-1 +b f ) (Equation 12)

[0089] Where, W f Represents the connection weights between the previous hidden layer and the inputs z t 、x t And b f Represents the corresponding bias.

[0090] The update gate is defined as

[0091] i t =sig(W i [z t ,x t ,h t-1 +b i ) (Equation 13)

[0092] Where, W i Represents the connection weights between the previous hidden layer and the inputs z t 、x t And b i Represents the corresponding bias.

[0093] The output gate is defined as

[0094] o t =sig(W o [z t ,x t ,h t-1 +b o ) (Equation 14)

[0095] Where, Wo represents the connection weights between the previous hidden layer and the input z t , x t , and b o represents the corresponding bias. The output gate is used to calculate the hidden layer output h at time t t , which is defined as

[0096] h t = o t tanh(c t ) (Equation 15)

[0097] Different from existing load prediction related work, the proposed ELP-DAR does not directly output the real value of the load at the next moment, but outputs the probability distribution of the load prediction at the next moment. Considering the data feature differences in the edge server load, the ELP-DAR method considers likelihood functions based on different probability distributions, including Gaussian likelihood function and T-distribution likelihood function. Taking the Gaussian likelihood function as an example, only two parameters, the mean and variance, need to be considered, and its definition is

[0098]

[0099] where the mean is obtained by mapping the output of the last layer of the neural network to a linear layer, and its definition is

[0100]

[0101] while the variance is further obtained after activation by the softplus function, and its definition is

[0102]

[0103] Example 1:

[0104] This embodiment uses a real edge load dataset. This dataset contains the load data of 6,870 edge servers and records the load changes within 1 month at a sampling frequency of 1 minute. We use the CPU usage rate as the key performance indicator for sampling, and information such as the user ID to which the edge server belongs, the start recording time, the end recording time, and the sampling frequency are used as components of the edge load dataset. We input the data into the prediction model in batches, randomly dividing the dataset into three parts, namely the training set (50%), the validation set (25%), and the test set (25%). The training set is used for model training (calculating the weights of the neural network), the validation set is used for model selection (selecting hyperparameters and preventing overfitting), and the test set is used to evaluate the performance of the selected optimal model. In each dataset, multiple load instances are generated according to the load input length and the load prediction length. Among them, we set prediction lengths of 20 minutes, 40 minutes, 60 minutes, 1 day, 2 days, 3 days, etc. For the minute-level prediction scenarios, the input length is set to 60 minutes; for the day-level prediction scenarios, the input length is set to 3 days. At the same time, we encoded the user ID and used it as a covariate, which, together with the edge load data, is used as the input of the edge load prediction model. To evaluate the accuracy of edge load prediction, we introduced performance indicators such as RMSE, ND, and mean wQuantileLoss. In addition, the total number of training cycles is 100, the initial learning rate is 0.001, the number of LSTM layers is 3, the number of neurons in each layer is 120, and the batch size is 32.

[0105]

[0106]

[0107] Table 1 Performance Comparison of Using T-Distribution and Gaussian Distribution under Different Prediction Lengths

[0108] First, for different prediction length scenarios, we evaluated the influence of different probability distribution functions on the prediction performance of the proposed ELP-DAR method. As shown in Table 1. When the prediction length is less than or equal to 1 day, the ELP-DAR method using the Gaussian distribution can achieve better performance than using the T-distribution; when the prediction length is greater than 1 day, as the prediction length increases, the prediction accuracy of the ELP-DAR method using the Gaussian distribution gradually decreases, while the ELP-DAR method using the T-distribution still maintains a high prediction accuracy.

[0109]

[0110]

[0111] Table 2 Performance of Various Methods under Different Prediction Lengths

[0112] Secondly, we compared with the MQ-RNN, MQ-CNN and SFF benchmark methods. As shown in Table 2, in the minute-level prediction scenario, the ELP-DAR method demonstrated better performance in several different evaluation metrics. When the prediction length is short, since the load change pattern is relatively simple, the ELP-DAR method does not show a significant superiority in prediction accuracy compared to the other three benchmark methods. In particular, in the Mean wQuantileLoss evaluation metric (which reflects the calculation errors at different quantiles), the ELP-DAR method achieved higher prediction accuracy compared to the other three benchmark methods. This indicates that the ELP-DAR method not only has good performance in single-point prediction, but also in terms of probability distribution prediction, the quantiles obtained by the ELP-DAR method can better cover the true values.

[0113] Figure 4 Shows the prediction effects of different methods when the prediction length is 3 days. Among them, the 90% prediction interval means that there is a 90% probability that the predicted load value falls within this interval; the 50% prediction interval means that there is a 50% probability that the predicted load value falls within this interval. It can be found from the figure that the ELP-DAR method can accurately predict the probability distribution of the load within the next 3 days, and most of the true load values also fall within the 90% prediction interval. It can be concluded from the figure that the ELP-DAR method maintains a high confidence level, and its prediction interval can accurately reflect the change trend of the true load value.

[0114] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope of the present invention.

Claims

1. A method for edge prediction based on a deep autoregressive recurrent neural network, characterized in that, It includes the following steps: Step S1: Obtain historical edge load data and construct a data set; Step S2: Preprocess the data set; Step S3: Take the users of the edge server as covariates; Step S4: Input the preprocessed data set and covariates into a prediction model to predict future edge loads; Step S5: Based on the edge load prediction results, assist the edge computing service in formulating a resource scheduling plan; The historical edge load data includes CPU usage rate, specifically: For the load of an edge server i, z i,t represents the change in the CPU usage rate of edge server i over a period of time t. Then the historical load sequence of this edge server is expressed as: The future predicted load sequence is represented as Among them, t0 represents the starting point of load prediction, T is the total length of the load sequence, 1:t0 - 1 represents the time range [1, t0 - 1] of the historical load sequence, and t0:T represents the time range [t0, T] of the future predicted load sequence; The specific content of Step S3 is: The scaled historical load and the historical covariates will be used as inputs to a prediction model to predict future edge loads. The prediction model can model future load conditions based on known historical loads and obtain the probability distribution of future loads. The prediction model uses LSTM to extract temporal features, with the input being edge load data and covariates over a past period of time, and the goal being to predict the probability distribution of the edge load z i,t at each time step, and the probability distribution is defined as where h i,t represents the output of an LSTM; taking the observation value z i,t-1 at the previous time step and the LSTM output h i,t-1 as inputs, the following can be obtained through calculation: h i,t = h(h i,t-1 , z i,t-1 , x i,t , Θ h ) Among them, h represents a neural network with a multi-layer LSTM structure, and Θ h is the set of parameters in the LSTM; the edge load data z i,t-1 at the previous moment and the hidden layer output h i,t-1 will be used to calculate the network output h i,t at the current moment; the likelihood function l(z i,t |θ(h i,t , Θ l )) is a probability distribution, and its set of parameters, including the mean μ and variance σ, is calculated through θ(h i,t , Θ l ); Θ l represents the mapping from h i,t to the set of likelihood function parameters.

2. The edge prediction method based on a deep autoregressive recurrent neural network according to claim 1, characterized in that The preprocessing includes data cleaning and resampling.

3. The edge prediction method based on a deep autoregressive recurrent neural network according to claim 1, characterized in that There are differences in CPU usage among different servers, so the mean value of the load data of each edge server is used as the scaling factor v i , and when inputting into the prediction model, the load is divided by v i , and when outputting from the prediction model, the corresponding load is multiplied by v i , v i The calculation formula is as follows: z i,t is the true load value.

4. The edge prediction method based on a deep autoregressive recurrent neural network according to claim 3, wherein The prediction model is an edge load prediction model based on a deep autoregressive recurrent neural network, adopting a sequence-to-sequence architecture, including an encoder and a decoder. The encoder and the decoder adopt the same network structure and share weights. The encoder encodes and outputs and takes as the initial input of the decoder, where L c is the load input length.

5. The edge prediction method based on a deep autoregressive recurrent neural network according to claim 4, characterized in that The training of the prediction model is specifically as follows: The input at each time step includes the covariates at the current moment The load value at the previous moment And the network output at the previous moment The network output at the current moment is obtained through training Furthermore, calculate the likelihood function l(z i,t |θ(h i,t ,Θ l )) parameters Next, optimize the loss function through the stochastic optimization method with adaptive momentum where N is the total number of the edge load data sequences, and θ(h i,t , Θ l ) is the parameter set of the likelihood function.

Citation Information

Patent Citations

  • Load prediction method and device and resource scheduling method and device for multiple cloud data centers

    CN113220450A

  • Street-crossing pedestrian trajectory prediction method based on SFM-LSTM neural network model

    CN114462667A