Dynamic Workload Prediction Method Based on Data Flow and Batch Processing Integrated Architecture

By adopting a combination method of complementary attention mechanism, multi-scale time convolution network and lightweight decoder in the integrated data flow batch processing architecture, multiple shortcomings in workload prediction in the prior art are solved, and more efficient and accurate dynamic workload prediction is achieved.

CN119961010BActive Publication Date: 2025-06-13NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510443422.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-13
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

When handling workload prediction in data flow and batch tasks, the prior art faces problems such as insufficient global and local dependency modeling, limited multi-scale time feature modeling capabilities, balance of computing efficiency and prediction performance, and insufficient adaptability of dynamic environments.

Method used

Using a dynamic workload prediction method based on an integrated data flow batch processing architecture, multi-scale time-dependent features are extracted, long-term and short-term dependency characteristics are captured, and long-term and short-term dependency characteristics are predicted through lightweight decoders.

Benefits of technology

It significantly improves the accuracy and computing efficiency of dynamic workload prediction, and is especially suitable for complex time series data processing in dynamic environments, and can effectively balance system performance and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961010B_ABST
    Figure CN119961010B_ABST
Patent Text Reader

Abstract

This application belongs to the technical field of edge cloud devices, and discloses a dynamic workload prediction method based on a data stream batch processing integrated architecture. The prediction method maps static features through a linear decoding layer, fuses static features and dynamic time series features, extracts multi-scale time-dependent features through a multi-scale time convolutional network layer, and sends them into an encoder to obtain the output of a complementary attention mechanism, and predicts the complementary attention output through a lightweight decoder. Through the design of the complementary attention mechanism, multi-scale time convolutional network, and lightweight decoder, this application effectively improves the accuracy and efficiency of dynamic workload prediction, and has broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of edge cloud devices, and specifically relates to a dynamic workload prediction method based on a data stream batch processing integrated architecture. Background Art

[0002] With the rapid development of big data and distributed computing, data stream processing and batch processing have gradually become the core technologies supporting modern computing systems. The integrated computing mode of data stream batch processing can process massive real-time stream data and historical batch data simultaneously, providing an efficient solution for tasks such as stream analysis, event detection, and batch prediction. This technology is widely used in complex systems such as cloud computing, edge computing, and the Internet of Things to support dynamic resource scheduling, performance optimization, and cost control.

[0003] However, in data stream and batch processing tasks, the dynamic changes in workload pose huge challenges. The workloads in cloud computing data centers and edge computing environments are jointly affected by user behavior, application characteristics, and infrastructure changes, showing complex dynamic characteristics. For example, bursty loads such as real-time peak access, concept drift such as pattern changes during long-term use, and cold start problems such as new applications or infrastructure are key factors affecting prediction performance. These characteristics make it difficult for existing resource allocation strategies to effectively balance system performance and efficiency, especially in dynamic environments.

[0004] Workload prediction is one of the core technologies for resource management in data stream batch processing systems. Through accurate workload prediction, the system can implement proactive scaling strategies, thereby optimizing virtual machine (VM) resource allocation, load balancing, and power consumption management, and avoiding over-provisioning or under-provisioning of resources. Traditional workload prediction methods mainly rely on statistical modeling and classical machine learning techniques, such as ARIMA and VARMA, which can handle linear time series but have significant deficiencies in dealing with non-linear dependencies and complex dynamic patterns. In recent years, deep learning methods, such as LSTM, GRU, and Transformer, have made significant progress and can enhance prediction performance by extracting complex patterns. However, these methods have deficiencies in aspects such as global and local dependency modeling, multi-scale time feature modeling ability, and the balance between computational efficiency and prediction performance when facing large-scale dynamic data.

[0005] Chinese Patent Application No. CN2023102296659 discloses a method and system for predicting the load waveform of edge cloud devices, which is a load waveform prediction method based on an information pool and a Transformer architecture. This method decomposes the historical workload sequence through the STL algorithm, extracts the seasonal cycle part, and constructs an information pool using VaDE clustering to capture the global application pattern. At the same time, this method combines the encoder and decoder modules of the Transformer architecture, fuses the information pool with local load information, and realizes the interaction between global and local information through the GP Layer information merging layer, thereby achieving accurate prediction of the load waveform of edge cloud devices.

[0006] Although the load waveform prediction methods proposed in the prior art can effectively capture the dynamic changes of the load and provide global application pattern information in theory, they still face some challenges in practical applications. (1) Insufficient global and local dependence modeling: Existing methods mainly focus on capturing global patterns and ignore local sudden changes, resulting in poor robustness of prediction results in dynamic scenarios. (2) Limited multi-scale time feature modeling ability: Traditional models are difficult to capture long-term and short-term dependencies simultaneously, especially when dealing with multi-scale features in complex time series. (3) Balance problem between computational efficiency and prediction performance: In large-scale data stream scenarios, how to maintain high prediction accuracy while reducing computational costs remains a key challenge. (4) Insufficient adaptability to dynamic environments: Traditional workload prediction methods are difficult to effectively balance system performance and efficiency when the workload in cloud computing data centers and edge computing environments is jointly affected by user behavior, application characteristics, and infrastructure changes. Summary of the Invention

[0007] To solve the above technical problems, this application provides a dynamic workload prediction method based on a data stream batch processing integrated architecture. Through an innovative complementary attention mechanism, a multi-scale time convolutional network, and a lightweight decoder design, this application solves the problems of insufficient global and local dependence modeling, limited multi-scale time feature modeling ability, balance problem between computational efficiency and prediction performance, and insufficient adaptability to dynamic environments in the prior art, effectively improving the accuracy and efficiency of dynamic workload prediction and having broad application prospects.

[0008] To achieve the above object, this application is implemented through the following technical solutions:

[0009] This application is a dynamic workload prediction method based on a data stream batch processing integrated architecture, including an information embedding layer, a multi-scale time convolutional network layer, an encoding layer, and a linear decoding layer. The dynamic workload prediction method specifically includes the following steps:

[0010] Step 1: Input the static features in the bandwidth workload dataset into the linear decoding layer for mapping to obtain a dense vector space with the same dimension as the dynamic time series features in the bandwidth workload dataset;

[0011] Step 2: Fuse the static features and the dynamic time series features to generate the final input features;

[0012] Step 3: Process the final input features generated in Step 2 through the multi-scale time convolutional network layer to extract multi-scale time-dependent features, where the multi-scale time-dependent features include long-term dependence relationship features and short-term dependence relationship features;

[0013] Step 4: Feed the multi-scale time-dependent features extracted in Step 3 into the encoder. The encoder captures the long-term dependence relationship features and the short-term dependence relationship features through the complementary attention mechanism to obtain the output of the complementary attention mechanism, where the complementary attention mechanism includes sparse attention and local attention;

[0014] Step 5: Predict through the lightweight decoder the complementary attention output obtained in Step 4.

[0015] A further improvement of this application lies in: In Step 1, the static features are processed through static embedding to map the static features to a dense vector space with the same dimension as the dynamic time series features. The specific operation is as follows:

[0016]

[0017] Among them, is the static feature, is the result of mapping to the dimension of the dynamic time series features, is the index of the time step, () is the linear transformation.

[0018] A further improvement of this application lies in: In Step 2, the fusion of the static features and the dynamic time series features specifically includes the following steps:

[0019] Step 2.1: Value embedding: Map the dynamic time series features to a high-dimensional feature representation consistent with the dimension of the dynamic workload prediction model through a one-dimensional convolutional operation to obtain the value embedding. The specific is as follows:

[0020]

[0021] Among them, is the dynamic time series feature, is the value embedding after one-dimensional convolutional processing;

[0022] Step 2.2, Position Embedding: Encode the position information of the dynamic time series features through sine and cosine functions, and add the encoded position information of the dynamic time series features element-wise to the value embedding to encode the position information for each dynamic time series feature and maintain the time order. Specifically: For,

[0023]

[0024]

[0025] where, is the time step index, is the dimension index of the position embedding, is the dimension size of the position embedding, is the position information encoded by the sine function for the dynamic time series features, is the position information encoded by the cosine function for the dynamic time series features. Concatenate and to obtain the position embedding :

[0026] = ;

[0027] Step 2.3, Dynamic Time Series Feature Embedding: Encode periodic dynamic time series features such as hours, dates, etc. Specifically:

[0028]

[0029] where, is the dynamic time series feature embedding, is the embedding of the th dynamic time series feature, is the number of dynamic time series features;

[0030] In step 2.4, the information embedding layer of the dynamic workload prediction model adds the value embedding in step 2.1, the position embedding in step 2.2, and the dynamic time series feature embedding in step 2.3 to form the final input feature :

[0031]

[0032] where, is the final input feature.

[0033] A further improvement of this application is that in the above step 3, the multi-scale time convolutional network layer adopts causal convolutional layers and dilated convolutional layers to extract multi-scale time-dependent features through multiple convolutional modules, specifically including the following steps:

[0034] Step 3.1. The causal convolutional layer captures short-term dependency features by restricting the convolutional kernel to depend only on the current and previous time-step inputs. The operation of the causal convolutional layer is expressed as:

[0035]

[0036] where, is the output at time-step index , is the input at time-step index , is the weight of the convolutional kernel of the causal convolutional layer, is the size of the convolutional kernel of the causal convolutional layer, is the position of the convolutional kernel of the causal convolutional layer in the time dimension, is the bias term;

[0037] Step 3.2. The dilated convolutional layer captures long-term dependency features by inserting holes between the elements of the convolutional kernel to expand the receptive field of the convolution. The operation of the dilated convolutional layer is expressed as:

[0038]

[0039] where, is the input at time-step index , is the dilation rate, representing the interval between the elements of the convolutional kernel of the dilated convolutional layer;

[0040] Step 3.3. The multi-scale temporal convolutional network layer includes multiple stacked TemporalBlock convolutional blocks. Each TemporalBlock convolutional block contains a causal convolutional layer, a dilated convolutional layer, an activation function, and a residual connection. The output calculation formula for each TemporalBlock convolutional block is:

[0041]

[0042] where, is the input, is the convolutional operation, is the output after being processed by the activation function;

[0043] Step 3.4. To prevent overfitting, a regularization layer is added after each TemporalBlock convolutional block as a regularization means:

[0044] .

[0045] A further improvement of this application lies in: in step 4, the encoder captures long-term and short-term dependency features through a complementary attention mechanism, which specifically includes the following steps:

[0046] Step 4.1, Sparse Attention restricts the calculation range during the calculation of attention weights through a strategy, and only focuses on the first time steps to capture long-term dependency features;

[0047] Step 4.2, Local Attention adjusts the first time steps to be focused through a dynamic window mechanism, capturing short-term dependency features and sudden changes;

[0048] Step 4.3, Combine the sparse attention and local attention mechanisms to capture long-term and short-term dependency features, obtaining the output of the complementary attention mechanism.

[0049] A further improvement of this application lies in: step 4.1 specifically includes the following steps:

[0050] Step 4.1.1, The sparse attention mechanism calculates the dot product between the query vector and the key vector to obtain the original attention score matrix :

[0051]

[0052] Among them, is the query vector, is the key vector, is the dimension of the key vector;

[0053] Step 4.1.2, Perform screening on the attention score matrix obtained in step 4.1.1, and the calculation method is as follows:

[0054]

[0055] Step 4.1.3, The attention scores after being screened by the strategy are used to weight the value vector , obtaining the output after the sparse attention mechanism:

[0056] .

[0057] A further improvement of this application lies in: step 4.2 specifically includes the following steps:

[0058] Step 4.2.1: For each query vector The content is dynamically adjusted to the window size , so that the attention range is adaptive according to the actual data changes. Specifically, the window size Based on the current query vector The relationship between the historical time step is adjusted to focus on the most relevant time step, which is calculated as:

[0059]

[0060] in, is a learning function, based on the current query vector and the set of adjacent key vectors ( ) Calculate the appropriate dynamic window size ;

[0061] Step 4.2.2: Based on dynamic window size , the local attention mechanism is used for each current query vector Calculate the current query vector With the corresponding key vector The similarity is calculated only within the dynamic window range to obtain the attention score , the specific calculation process is as follows:

[0062]

[0063] in, and is the time step index determined by the dynamic window size, indicating the query vector The corresponding key vector The effective range of

[0064] Step 4.2.3: Attention score Normalize and then the attention score With value vector Perform weighted summation to obtain the final complementary attention mechanism output, which is calculated as follows:

[0065] .

[0066] A further improvement of the present application is that the step 4.3 specifically includes the following steps:

[0067] Step 4.3.1: In the forward propagation process of the dynamic workload prediction model, the dynamic time series features of the input are first calculated through the sparse attention mechanism. long-term dependency features, and obtain sparse attention output , and then perform regularization and normalization processing, specifically as follows:

[0068] ;

[0069] The sparse attention output obtained in Step 4.3.2 and Step 4.3.1 is passed to the local attention mechanism, and after calculation, the local attention output is obtained , specifically as follows:

[0070] ;

[0071] Step 4.3.3, add the sparse attention output and the local attention output, and then perform regularization and normalization processing to obtain the final combined output :

[0072]

[0073] Step 4.3.4, the combined output after fusion is fed into the feed-forward neural network of the dynamic workload prediction model to obtain the final complementary attention mechanism output , specifically as follows:

[0074] .

[0075] A further improvement of the present invention lies in that: the lightweight decoder in Step 5 adopts a linear decoding layer to reduce the computational complexity, specifically as follows:

[0076]

[0077] wherein, is the output after being processed by the linear decoding layer, is the prediction result.

[0078] The beneficial effects of this application are:

[0079] While improving the prediction accuracy, this application significantly improves the computational efficiency and is particularly suitable for processing complex time series data in a dynamic environment.

[0080] The complementary attention mechanism of this application combines sparse attention and local attention, and has significant advantages in capturing long-term dependence relationship features and short-term dependence relationship features, improving the performance of the model in processing dynamic load prediction.

[0081] In this application, a multi-scale temporal convolutional network (TCN) module is introduced before the encoder to extract multi-scale temporal features, effectively enhancing the model's ability to capture long-term and short-term dependency features. Especially when dealing with local sudden changes and concept drift, the robustness and generalization ability of the model are improved.

[0082] The lightweight decoder of this application replaces the traditional multi-layer decoder structure with a single-layer linear decoder, significantly reducing the computational complexity and inference time, and improving the computational efficiency of the model on large-scale datasets.

[0083] Through the combination of complementary attention mechanism, multi-scale temporal convolutional network and lightweight decoder, this application effectively improves the accuracy and computational efficiency of dynamic workload prediction, and is particularly suitable for complex dynamic data processing in distributed environments, cloud computing and edge computing environments. Thus, it helps the system optimize resource allocation, improve performance and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 is a flowchart of this application.

[0085] Figure 2 is a schematic diagram of the dynamic workload prediction model of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] The following will disclose the embodiments of this application in the form of diagrams. For the sake of clarity, many practical details will be described together in the following narrative. However, it should be understood that these practical details are not used to limit this application. That is to say, in some embodiments of this application, these practical details are not necessary.

[0087] As Figure 1 shown, this application is a dynamic workload prediction method based on a data flow batch processing integrated architecture. This method is implemented through a dynamic workload prediction model, and the dynamic workload prediction model includes an information embedding layer, a multi-scale temporal convolutional network layer, an encoding layer, and a linear decoding layer. The dynamic workload prediction method specifically includes the following steps:

[0088] Step 1: Input the static features in the bandwidth workload dataset into the linear decoding layer for mapping to obtain a dense vector space with the same dimension as the dynamic time series features in the bandwidth workload dataset, and effectively combine it with the subsequent dynamic time series features. The specific operation is as follows:

[0089]

[0090] Among them, is the static feature, is the result mapped to the dimension of the dynamic time series feature, that is, the static feature is mapped to the same dimension size as the dynamic feature, is the index of the time step, () is a linear transformation. Among them, the static features include server location, number of CPUs, memory size, etc.

[0091] Step 2: Integrate the static features and the dynamic time series features to generate the final input features. Specifically, it includes the following steps:

[0092] Step 2.1: Value embedding: Map the dynamic time series features to a high-dimensional feature representation consistent with the dimension of the dynamic workload prediction model through a one-dimensional convolution operation to obtain value embedding, specifically:

[0093]

[0094] Among them, is the dynamic time series feature, is the value embedding after one-dimensional convolution processing;

[0095] Step 2.2: Position embedding: Encode the position information of the dynamic time series features through sine and cosine functions, and add the encoded position information of the dynamic time series features to the value embedding element-wise to encode the position information for each dynamic time series feature and maintain the time order, specifically:

[0096]

[0097]

[0098] Among them, is the time step index, is the dimension index of the position embedding, is the dimension size of the position embedding, is the position information encoded by the sine function for the dynamic time series features, is the position information encoded by the cosine function for the dynamic time series features, and is concatenated with to obtain the position embedding :

[0099] = ;

[0100] In this way, the position embedding can effectively provide position information for each time step, ensuring that the model can distinguish the inputs at different time steps.

[0101] Step 2.3: Embedding of dynamic time series features: Encode periodic dynamic time series features such as hours, dates, etc., to enhance the dynamic workload prediction model's perception ability of changes in dynamic time series features, specifically:

[0102]

[0103] Among them, is the dynamic time series feature embedding, is the embedding of the th dynamic time series feature, is the number of dynamic time series features;

[0104] Step 2.4: The information embedding layer of the dynamic workload prediction model sums up the value embedding in Step 2.1, the position embedding in Step 2.2, and the dynamic time series feature embedding in Step 2.3 to form the final input feature :

[0105]

[0106] Among them, is the final input feature.

[0107] Step 3: Process the final input feature generated in Step 2 through a multi-scale temporal convolutional network layer to extract multi-scale temporal dependence features, where the multi-scale temporal dependence features include long-term dependence relationship features and short-term dependence relationship features.

[0108] Specifically, extracting the multi-scale temporal dependence features specifically includes the following steps:

[0109] Step 3.1: The causal convolutional layer ensures that the dynamic workload prediction model does not leak future information in the time series prediction task by restricting the convolutional kernel to depend only on the current and previous time step inputs, and captures short-term dependence relationship features. This design makes it particularly suitable for capturing short-term dependence features in time series, such as local trends and instantaneous changes. The operation of the causal convolutional layer is expressed as:

[0110]

[0111] Among them, is the output at time step index , is the input at time step index , is the weight of the convolutional kernel of the causal convolutional layer, is the size of the convolutional kernel of the causal convolutional layer, is the position of the convolutional kernel of the causal convolutional layer in the time dimension, is the bias term;

[0112] Step 3.2: To further expand the receptive field, the multi-scale temporal convolutional network TCN uses dilated convolution. The dilated convolutional layer expands the receptive field of the convolution by inserting holes between the elements of the convolutional kernel, capturing long-term dependency features. The operation of the dilated convolutional layer is expressed as:

[0113]

[0114] where is the time step index of the input, is the dilation rate, indicating the interval between the elements of the convolutional kernel of the dilated convolutional layer; it can capture long-term dependency features at a relatively low computational cost.

[0115] Step 3.3: The multi-scale temporal convolutional network layer includes multiple stacked TemporalBlock convolutional blocks. Each TemporalBlock convolutional block contains a causal convolutional layer, a dilated convolutional layer, an activation function, and a residual connection to enhance the model's multi-scale feature extraction ability for time series data. The output calculation formula for each TemporalBlock convolutional block is:

[0116]

[0117] where is the input, is the convolution operation, is the output after being processed by the activation function; the residual connection adds the input to the output of the convolutional layer, helping to alleviate the vanishing gradient problem and promoting the flow of information.

[0118] Step 3.4: To prevent overfitting, a regularization layer is added after each TemporalBlock convolutional block as a regularization means:

[0119] .

[0120] Step 4: Feed the multi-scale time-dependent features extracted in Step 3 into the encoder. The encoder captures long-term and short-term dependency features through a complementary attention mechanism, obtaining the output of the complementary attention mechanism, where the complementary attention mechanism includes sparse attention and local attention. Specifically, it includes the following steps:

[0121] Step 4.1: Sparse Attention restricts the calculation range during the calculation of attention weights through the strategy, only focusing on the first For each time step, capture the features of long-term dependencies.

[0122] First, the sparse attention mechanism calculates the dot product between the query vector and the key vector to obtain the original attention score matrix :

[0123]

[0124] where is the query vector, is the key vector, is the dimension of the key vector;

[0125] Secondly, the obtained attention score matrix is filtered. According to similarity or correlation, the most important attention connections are selected, and only the top corresponding to each query vector maximum values are retained, ignoring the irrelevant parts, that is, the remaining attention scores are set to zero. This can effectively reduce the computational complexity and focus on the most informative time steps, thereby improving the efficiency of the model in dealing with long-term dependencies. The calculation method is as follows:

[0126]

[0127] Finally, the attention scores filtered by the strategy are used to weight the value vector , obtaining the output after the sparse attention mechanism:

[0128] .

[0129] Step 4.2, Local Attention adjusts the first time steps to be attended to through a dynamic window mechanism, capturing short-term dependency features and sudden changes. Local Attention can quickly capture these sudden changes by limiting its attention range and only calculating between local time steps. Specifically:

[0130] First, by dynamically adjusting the window size for the content of each query vector , the attention range adapts to the actual data changes. Specifically, the window size will be adjusted according to the relationship between the current query vector and the historical time steps, so that the attention is concentrated on the most relevant time steps. The calculation method is:

[0131]

[0132] Among them, is a learning function that calculates an appropriate dynamic window size based on the current query vector and the set of adjacent key vectors ( ); ;

[0133] Secondly, based on the dynamic window size , the local attention mechanism calculates the similarity between each current query vector and the corresponding key vector only within the dynamic window range to obtain the attention score . The specific calculation process is as follows:

[0134]

[0135] Among them, and are time step indices determined by the dynamic window size, indicating the effective range of the key vector corresponding to the th query vector ;

[0136] Finally, the attention scores are normalized, and then the attention scores are weighted and summed with the value vectors to obtain the final output of the complementary attention mechanism. The calculation method is as follows:

[0137] .

[0138] Step 4.3: Combine the sparse attention and the local attention mechanism to capture the long-term and short-term dependency features, and obtain the output of the complementary attention mechanism. The combination of the sparse attention and the local attention mechanism can not only capture the long-term time dependency features but also quickly respond to short-term fluctuations. It specifically includes the following steps:

[0139] Step 4.3.1: During the forward propagation of the dynamic workload prediction model, first calculate the long-term dependency features of the input dynamic time series features through the sparse attention mechanism to obtain the sparse attention output , and then perform regularization and normalization processing. Specifically:

[0140] ;

[0141] Step 4.3.2, the sparse attention output obtained in Step 4.3.1 is passed to the local attention mechanism, and after calculation, the local attention output is obtained , specifically:

[0142] ;

[0143] Step 4.3.3, add the sparse attention output and the local attention output, and then perform regularization and normalization processing to obtain the final combined output :

[0144]

[0145] Step 4.3.4, the combined output after fusion is fed into the feed-forward neural network of the dynamic workload prediction model to obtain the final complementary attention mechanism output , specifically:

[0146] .

[0147] Step 5, predict the complementary attention output obtained in Step 4 through a lightweight decoder. Among them, the lightweight decoder uses a linear decoding layer to reduce the computational complexity, specifically:

[0148]

[0149] wherein, is the output after being processed by the linear decoding layer, is the prediction result.

[0150] To verify the advantages of this application in dynamic load prediction, the following experiments are provided.

[0151] 1. Experimental setup

[0152] Dataset: This application uses multiple datasets for experimental evaluation, including:

[0153] ECW: From the edge computing environment, containing the upload bandwidth workload changes of 797 edge servers in August 2022.

[0154] ECW-App-Switch: Records the workload changes of application switches from August 25, 2022 to August 30, 2022.

[0155] ECW-New-App: Contains the workload changes of newly added applications from August 30, 2022 to September 6, 2022.

[0156] In addition, static content data in 12 dimensions, such as maximum bandwidth, number of CPUs, location, etc., are also provided in the dataset to assist model prediction.

[0157] Experimental details: The workload sequences of all ECW datasets are divided by hour, and the maximum value per hour is selected to match the time granularity of application scheduling and billing rules. The input sequence length is set to T = 48, and the prediction length is L = 24. To train and evaluate the dynamic workload prediction model, the dataset is divided into a training set, a validation set, and a test set in chronological order, with a ratio of 6:2:2. All experiments are implemented in a PyTorch environment and conducted on a machine with an NVIDIA GeForce RTX 3090 GPU. This application uses the ADAM optimizer for training, with an initial learning rate of 10^-4, a batch size of 256, and 50 training epochs.

[0158] 2. Experimental Results

[0159] Evaluation metrics: To comprehensively evaluate the prediction effect of this application, this application uses the mean squared error (MSE) and the mean absolute error (MAE) as evaluation metrics. The lower the values of MSE and MAE, the better the prediction effect of this application.

[0160] Benchmark models: This application selects five widely used time series prediction methods as benchmark models for comparison, including: Deep Transformer, Informer, Autoformer (Transformer-based models); MQRNN (RNN-based model); and VaRDE-LSTM (cluster-based model).

[0161] Experimental results: As can be seen from Table 1, this application shows significant superiority on multiple datasets. Especially on the ECW dataset, this application reduces the MSE by an average of 9.2%, and performs better compared with other Transformer-based models. In the comparison with the MQRNN model based on RNN, the MSE of this application on the ECW dataset is reduced by 19.5% and 94.8% respectively, demonstrating its advantages in dealing with long-term dependencies and complex dynamic patterns.

[0162] Table 1 Comparison table of average prediction performance of workloads

[0163]

[0164] In Table 1, lower MSE or MAE indicates better prediction effect. The best performance on each dataset is shown in bold.

[0165] The superiority of the present application in multiple complex dynamic workload prediction tasks is proved by the above experimental results, especially in capturing long-term dependence features and local burst patterns, i.e., short-term dependence features. Compared with the existing mainstream methods, the present application has achieved lower MSE and MAE values on multiple data sets, demonstrating its powerful ability in multi-dimensional data modeling and capturing complex dependence relationships, and proving its great potential in practical applications.

[0166] In summary, through the innovative complementary attention mechanism, multi-scale temporal convolutional network, and lightweight decoder design, the present application solves the problems in the prior art, such as insufficient global and local dependence modeling, limited multi-scale temporal feature modeling ability, the balance problem between computational efficiency and prediction performance, and insufficient adaptability to dynamic environments. It effectively improves the accuracy and efficiency of dynamic workload prediction and has broad application prospects.

[0167] The above description is only for the implementation manners of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A dynamic workload prediction method based on a data stream batch processing integrated architecture, characterized in that: The dynamic workload prediction method is implemented by a dynamic workload prediction model, which includes an information embedding layer, a multi-scale time convolutional network layer, a coding layer, and a linear decoding layer. The dynamic workload prediction method specifically includes the following steps: Step 1: Input the static features in the bandwidth workload dataset into the linear decoding layer for mapping to obtain a dense vector space with the same dimension as the dynamic time series features in the bandwidth workload dataset; Step 2: Fuse the static features and dynamic time series features to generate the final input features; Step 3: Process the final input features generated in step 2 through a multi-scale time convolutional network layer to extract multi-scale time-dependent features, wherein the multi-scale time-dependent features include long-term dependency features and short-term dependency features; Step 4: The multi-scale time-dependent features extracted in step 3 are fed into the encoder. The encoder captures the long-term dependency features and the short-term dependency features through the complementary attention mechanism to obtain the complementary attention mechanism output. The complementary attention mechanism includes sparse attention and local attention. Step 5: Predict the complementary attention output obtained in step 4 through a lightweight decoder; wherein: In step 4, the encoder captures long-term dependency features and short-term dependency features through a complementary attention mechanism, which specifically includes the following steps: Step 4.1: Sparse attention uses the Top-k strategy to limit the calculation scope of attention weight calculation, focusing only on the first k time steps to capture long-term dependency features; Step 4.2: Local attention adjusts the first k time steps through a dynamic window mechanism to capture short-term dependency features; Step 4.3: Combine sparse attention and local attention to capture long-term dependency features and short-term dependency features, and obtain the output of the complementary attention mechanism. Step 4.3 specifically includes the following steps: Step 4.3.1: In the forward propagation process of the dynamic workload prediction model, the long-term dependency features of the input dynamic time series features X are first calculated through the sparse attention mechanism to obtain the sparse attention output SparseOut, and then regularized Dropout and normalized LayerNorm are performed, specifically: SparseOut=Dropout(LayerNorm(SparseAttention(X,X,X)+X)); The sparse attention output SparseOut obtained in step 4.3.2 and step 4.3.1 is passed to the local attention mechanism, and the local attention output LocalOut is obtained after calculation, which is specifically: LocalOut=LocalAttention(SparseOut, SparseOut, SparseOut); Step 4.3.3, add the sparse attention output and the local attention output, and then perform regularized Dropout and normalized LayerNorm processing to obtain the final combined output CombinedOut: CombinedOut=Dropout(LayerNorm(LocalOut+SparseOut)) Step 4.3.4: The combined output CombinedOut after fusion is fed into the feedforward neural network of the dynamic workload prediction model to obtain the final complementary attention mechanism output FinalOut, which is: FinalOut=PositionWise FeedForward(CombinedOut).

2. The dynamic workload prediction method based on the data stream batch processing integrated architecture according to claim 1 is characterized in that: In step 1, the static features are mapped through the static embedding input linear decoding layer to map the static features to a dense vector space of the same dimension as the dynamic time series features. The specific operation is: S′ t =Linear(S t ) Among them, S t is the static characteristic, S′ t It is the result of mapping to the feature dimension of dynamic time series, t is the time step index, and Linear() is the linear transformation.

3. The dynamic workload prediction method based on data stream batch processing integrated architecture according to claim 1 is characterized in that: In step 2, the static features and the dynamic time series features are fused, which specifically includes the following steps: Step 2.1, value embedding: Map the dynamic time series features to a high-dimensional feature representation consistent with the dimension of the dynamic workload prediction model through a one-dimensional convolution operation to obtain value embedding, specifically: X token =Conv1D(X) Among them, X is the dynamic time series feature, X token is the value embedding after one-dimensional convolution; Step 2.2, position embedding: Encode the position information of dynamic time series features through sine and cosine functions, and embed the encoded position information and value of dynamic time series features into X token Add element by element to encode the position information for each dynamic time series feature and keep the time order, specifically: Among them, t is the time step index, i is the dimension index of the position embedding, and d model is the dimension size of position embedding, PE(t.2i) is the sine function encoding the position information of dynamic time series features, PE(t.2i+1) is the cosine function encoding the position information of dynamic time series features, and the position embedding X is obtained by concatenating PE(t.2i) and PE(t.2i+1) position : X position =PE(t.2i)+PE(t.2i+1); Step 2.3: Dynamic time series feature embedding: Encode the periodic dynamic time series features, specifically: Among them, X temporal Embed is a dynamic time series feature embedding. j (t) is the embedding of the jth dynamic time series feature, and m is the number of dynamic time series features; Step 2.4: The information embedding layer of the dynamic workload prediction model adds the value embedding of step 2.1, the position embedding of step 2.2, and the dynamic time series feature embedding of step 2.3 to form the final input feature X embedded : X embedded =X token +X position +X temporal Among them, X embedded is the final input feature.

4. The dynamic workload prediction method based on data stream batch processing integrated architecture according to claim 1 is characterized in that: In step 3, the multi-scale temporal convolutional network layer uses a causal convolutional layer and a dilated convolutional layer to extract multi-scale time-dependent features through a multi-layer convolutional module, specifically including the following steps: Step 3.1, the causal convolution layer captures short-term dependency features by limiting the convolution kernel to only depend on the current and previous time step inputs. The operation of the causal convolution layer is expressed as: Among them, y t is the output at time step index t, x t-m is the input of the time step index tm, w m is the weight of the convolution kernel of the causal convolution layer, a is the size of the convolution kernel of the causal convolution layer, m is the position of the convolution kernel of the causal convolution layer in the time dimension, and b is the bias term; Step 3.2: The atrous convolution layer expands the receptive field of the convolution and captures the long-term dependency features by inserting holes between the elements of the convolution kernel. The operation of the atrous convolution layer is expressed as: Among them, x t-d·m is the input of the time step index td·m, d is the dilation rate, which represents the interval between the convolution kernel elements of the dilated convolution layer; Step 3.3, the multi-scale temporal convolutional network layer includes multiple stacked TemporalBlock convolution blocks, each TemporalBlock convolution block includes a causal convolution layer and a hole convolution layer, an activation function and a residual connection, and the output calculation formula of each TemporalBlock convolution block is: y block =ReLU(Conv1D(x)+x) Where x is the input, Conv1D(x) is the convolution operation, and y block is the output after being processed by the activation function; Step 3.4, a regularized Dropout layer is added after each TemporalBlock convolution block as a regularization method: y dropout =Dropout(y block )。 5. The method for dynamic workload prediction based on data stream batch processing integrated architecture according to claim 1, characterized in that: The step 4.1 specifically includes the following steps: Step 4.1.1, the sparse attention mechanism calculates the dot product between the query vector Q and the key vector K to obtain the original attention score matrix A: Among them, Q is the query vector, K is the key vector, and d k is the dimension of the key vector; Step 4.1.2: Perform Top-k screening on the attention score matrix A obtained in step 4.1.

1. The calculation method is as follows: A top-k =Top-k(A) Step 4.1.3: Attention score A after Top-k strategy screening top-k is used to weight the value vector V and obtain the output after the sparse attention mechanism: Sparse Attention(Q,K,V)=Softmax(A top-k )·V。 6. The method for dynamic workload prediction based on data stream batch processing integrated architecture according to claim 1, characterized in that: Step 4.2 specifically includes the following steps: Step 4.2.1: By dynamically adjusting the window size w for each query vector Q, the attention range is adaptive according to the actual data changes. The calculation method is: w=f(Q t ,{K t-n ,...,K t+n }) Among them, f(·) is a learning function, according to the current query vector Q t and the set of adjacent key vectors (K t-n , ..., K t+n ) Calculate the dynamic window size w; Step 4.2.2: Based on the dynamic window size w, the local attention mechanism is used for each current query vector Q t Calculate the current query vector Q t The similarity with the corresponding key vector K is calculated only within the dynamic window range to obtain the attention score A i , the specific calculation process is as follows: Among them, start and end are time step indices determined by the dynamic window size, representing the s-th query vector Q s The valid range of the corresponding key vector K; Step 4.2.3: Attention score A i Normalized, then the attention score A i The weighted sum is performed with the value vector V to obtain the final complementary attention mechanism output, which is calculated as follows: Local Attention(Q,K,V)=Softmax(A i )·V。 7. The method for dynamic workload prediction based on data stream batch processing integrated architecture according to claim 1, characterized in that: The lightweight decoder in step 5 adopts a linear decoding layer to reduce the computational complexity, specifically: Y t =Linear(V t ) Among them, V t is the output after the linear decoding layer, Y t is the prediction result.

Citation Information

Patent Citations

  • Highway traffic flow prediction method based on Transform and graph attention network

    CN116092294A

  • Long sequence knowledge tracking method based on Informer

    CN118350418A