Dynamic workload prediction method based on data flow batch processing integrated architecture
By adopting complementary attention mechanism, multi-scale time convolution network and lightweight decoder in the integrated data flow batch processing architecture, multiple shortcomings in workload prediction in the prior art are solved, and more efficient and accurate dynamic load prediction is achieved.
Patent Information
- Application Number
- CN202510443422.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-10
AI Technical Summary
When handling workload prediction in data flow and batch tasks, the prior art faces problems such as insufficient global and local dependency modeling, limited multi-scale time feature modeling capabilities, balance of computing efficiency and prediction performance, and insufficient adaptability of dynamic environments.
Using a dynamic workload prediction method based on an integrated data flow batch processing architecture, multi-scale time-dependent features are extracted and predicted through the combination of complementary attention mechanism, multi-scale time-dependent features.
It significantly improves the accuracy and computing efficiency of dynamic workload prediction, and can more effectively capture long-term and short-term dependency characteristics, adapt to complex dynamic environments, optimize resource allocation and improve system performance.
Smart Images

Figure CN119961010A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of edge cloud device technology, and specifically relates to a dynamic workload prediction method based on a data stream batch processing integrated architecture. Background Art
[0002] With the rapid development of big data and distributed computing, data stream processing and batch processing have gradually become the core technologies supporting modern computing systems. The integrated computing mode of data stream and batch processing can simultaneously process massive real-time streaming data and historical batch data, providing efficient solutions for tasks such as streaming analysis, event detection, and batch prediction. This technology is widely used in complex systems such as cloud computing, edge computing, and the Internet of Things to support dynamic resource scheduling, performance optimization, and cost control.
[0003] However, the dynamic changes of workloads in data streams and batch processing tasks pose great challenges. Workloads in cloud computing data centers and edge computing environments are jointly affected by user behavior, application characteristics, and infrastructure changes, and exhibit complex dynamic characteristics. For example, burst loads such as real-time peak access, concept drift such as pattern changes in long-term use, and cold start problems such as new applications or infrastructure are key factors affecting prediction performance. These characteristics make it difficult for existing resource allocation strategies to effectively balance system performance and efficiency, especially in dynamic environments.
[0004] Workload prediction is one of the core technologies for resource management in data stream batch processing systems. With accurate workload prediction, the system can implement active expansion strategies to optimize virtual machine (VM) resource allocation, load balancing and power management, and avoid over- or under-configuration of resources. Traditional workload prediction methods mainly rely on statistical modeling and classical machine learning techniques, such as ARIMA and VARMA, which can handle linear time series but have significant deficiencies in handling nonlinear dependencies and complex dynamic patterns. In recent years, deep learning methods, such as LSTM, GRU and Transformer, have made significant progress and can enhance prediction performance by extracting complex patterns. However, when faced with large-scale dynamic data, these methods have deficiencies in global and local dependency modeling, multi-scale temporal feature modeling capabilities, and the balance between computational efficiency and prediction performance.
[0005] Chinese patent application number CN2023102296659 discloses a method and system for predicting the load waveform of edge cloud devices, which is a load waveform prediction method based on an information pool and a Transformer architecture. This method decomposes the historical workload sequence through the STL algorithm, extracts the seasonal cycle part, and uses VaDE clustering to construct an information pool to capture global application patterns. At the same time, this method combines the encoder and decoder modules of the Transformer architecture, fuses the information pool with the local load information, and realizes the interaction of global and local information through the GP Layer information merging layer, thereby realizing accurate prediction of the load waveform of edge cloud devices.
[0006] Although the load waveform prediction methods proposed by the existing technologies can theoretically effectively capture the dynamic changes of the load and provide global application pattern information, they still face some challenges in practical applications. (1) Insufficient modeling of global and local dependencies: Existing methods mainly focus on capturing global patterns and ignore local sudden changes, resulting in poor robustness of prediction results in dynamic scenarios. (2) Limited ability to model multi-scale temporal features: Traditional models are difficult to capture long-term and short-term dependencies at the same time, especially when dealing with multi-scale features in complex time series. (3) The problem of balancing computational efficiency and prediction performance: In large-scale data stream scenarios, how to reduce computing costs while maintaining high prediction accuracy remains a key challenge. (4) Insufficient adaptability to dynamic environments: Traditional workload prediction methods are difficult to effectively balance system performance and efficiency when faced with workloads in cloud computing data centers and edge computing environments that are jointly affected by user behavior, application characteristics, and infrastructure changes. Summary of the invention
[0007] In order to solve the above technical problems, the present application provides a dynamic workload prediction method based on a data stream batch processing integrated architecture. The present application solves the problems of insufficient global and local dependency modeling, limited multi-scale temporal feature modeling capabilities, balance between computational efficiency and prediction performance, and insufficient adaptability to dynamic environments in the prior art through innovative complementary attention mechanisms, multi-scale temporal convolutional networks, and lightweight decoder designs. It effectively improves the accuracy and efficiency of dynamic workload prediction and has broad application prospects.
[0008] In order to achieve the above objectives, this application is implemented through the following technical solutions:
[0009] The present application is a dynamic workload prediction method based on a data stream batch processing integrated architecture, including an information embedding layer, a multi-scale time convolutional network layer, an encoding layer, and a linear decoding layer. The dynamic workload prediction method specifically includes the following steps:
[0010] Step 1: Input the static features in the bandwidth workload dataset into the linear decoding layer for mapping to obtain a dense vector space with the same dimension as the dynamic time series features in the bandwidth workload dataset;
[0011] Step 2: Fuse the static features and dynamic time series features to generate the final input features;
[0012] Step 3: Process the final input features generated in step 2 through a multi-scale time convolutional network layer to extract multi-scale time-dependent features, wherein the multi-scale time-dependent features include long-term dependency features and short-term dependency features;
[0013] Step 4: Send the multi-scale time-dependent features extracted in step 3 to the encoder. The encoder captures the long-term dependency features and the short-term dependency features through the complementary attention mechanism to obtain the complementary attention mechanism output, where the complementary attention mechanism includes sparse attention and local attention.
[0014] Step 5: Predict the complementary attention output obtained in step 4 through a lightweight decoder.
[0015] A further improvement of the present application is that in step 1, the static features are processed by static embedding, and the static features are mapped to a dense vector space of the same dimension as the dynamic time series features. The specific operation is:
[0016]
[0017] in, is a static feature, As a result of mapping to the feature dimension of dynamic time series, is the index of the time step, () is a linear transformation.
[0018] The further improvement of the present application is that: In step 2, the static features and the dynamic time series features are fused, which specifically includes the following steps:
[0019] Step 2.1, value embedding: Map the dynamic time series features to a high-dimensional feature representation consistent with the dimension of the dynamic workload prediction model through a one-dimensional convolution operation to obtain value embedding, specifically:
[0020]
[0021] in, is the dynamic time series feature, is the value embedding after one-dimensional convolution;
[0022] Step 2.2, position embedding: Encode the position information of dynamic time series features through sine and cosine functions, and embed the encoded position information of dynamic time series features with values Add element by element to encode the position information for each dynamic time series feature and keep the time order, specifically:
[0023]
[0024]
[0025] in, is the time step index, is the dimension index of the position embedding, is the dimension size of the position embedding, The sine function encodes the position information of the dynamic time series features. The cosine function encodes the position information of the dynamic time series features. and Concatenate to get position embedding :
[0026] = ;
[0027] Step 2.3, dynamic time series feature embedding: Encode periodic dynamic time series features such as hours, dates, etc., specifically:
[0028]
[0029] in, is the dynamic time series feature embedding, For the Embedding of dynamic time series features, is the number of dynamic time series features;
[0030] Step 2.4: The information embedding layer of the dynamic workload prediction model adds the value embedding of step 2.1, the position embedding of step 2.2, and the dynamic time series feature embedding of step 2.3 to form the final input feature :
[0031]
[0032] in, is the final input feature.
[0033] A further improvement of the present application is that in step 3, the multi-scale temporal convolutional network layer adopts a causal convolutional layer and a hole convolutional layer, and extracts multi-scale time-dependent features through a multi-layer convolutional module, specifically including the following steps:
[0034] Step 3.1, the causal convolution layer captures short-term dependency features by limiting the convolution kernel to only depend on the current and previous time step inputs. The operation of the causal convolution layer is expressed as:
[0035]
[0036] in, is the time step index The output, is the time step index Input, is the weight of the convolution kernel of the causal convolution layer, is the size of the convolution kernel of the causal convolution layer, is the position of the convolution kernel of the causal convolution layer in the time dimension, is the bias term;
[0037] Step 3.2: The atrous convolution layer expands the receptive field of the convolution and captures the long-term dependency features by inserting holes between the elements of the convolution kernel. The operation of the atrous convolution layer is expressed as:
[0038]
[0039] in, is the time step index Input, is the dilation rate, which indicates the spacing between the convolution kernel elements of the dilated convolution layer;
[0040] Step 3.3, the multi-scale temporal convolutional network layer includes multiple stacked TemporalBlock convolution blocks, each TemporalBlock convolution block includes a causal convolution layer and a hole convolution layer, an activation function and a residual connection, and the output calculation formula of each TemporalBlock convolution block is:
[0041]
[0042] in, is the input, is the convolution operation, is the output after being processed by the activation function;
[0043] Step 3.4: To prevent overfitting, regularization is added after each TemporalBlock convolution block. Layers as a means of regularization:
[0044] .
[0045] A further improvement of the present application is that in step 4, the encoder captures long-term dependency features and short-term dependency features through a complementary attention mechanism, specifically including the following steps:
[0046] Step 4.1: Sparse Attention The strategy limits the calculation scope of attention weight calculation and only focuses on the front time steps to capture long-term dependency features;
[0047] Step 4.2: Local Attention adjusts the focus through a dynamic window mechanism time steps to capture short-term dependency characteristics and sudden changes;
[0048] Step 4.3: Combine sparse attention and local attention mechanisms to capture long-term dependency features and short-term dependency features to obtain complementary attention mechanism output.
[0049] A further improvement of the present application is that step 4.1 specifically includes the following steps:
[0050] Step 4.1.1: The sparse attention mechanism calculates the query vector and key vector The dot product between them gives the original attention score matrix :
[0051]
[0052] in, is the query vector, is the key vector, is the dimension of the key vector;
[0053] Step 4.1.2: The attention score matrix obtained in step 4.1.1 conduct Screening, the calculation method is as follows:
[0054]
[0055] Step 4.1.3, after Attention score after policy selection Used to weight the value vector , and get the output after the sparse attention mechanism:
[0056] .
[0057] A further improvement of the present application is that step 4.2 specifically includes the following steps:
[0058] Step 4.2.1: For each query vector The content is dynamically adjusted to the window size , so that the attention range is adaptive according to the actual data changes. Specifically, the window size Based on the current query vector The relationship between the historical time step is adjusted to focus on the most relevant time step, which is calculated as:
[0059]
[0060] in, is a learning function, based on the current query vector and the set of adjacent key vectors ( ) Calculate the appropriate dynamic window size ;
[0061] Step 4.2.2: Based on dynamic window size , the local attention mechanism is used for each current query vector Calculate the current query vector With the corresponding key vector The similarity is calculated only within the dynamic window range to obtain the attention score , the specific calculation process is as follows:
[0062]
[0063] in, and is the time step index determined by the dynamic window size, indicating the query vector The corresponding key vector The effective range of
[0064] Step 4.2.3: Attention score Normalize and then the attention score With value vector Perform weighted summation to obtain the final complementary attention mechanism output, which is calculated as follows:
[0065] .
[0066] A further improvement of the present application is that the step 4.3 specifically includes the following steps:
[0067] Step 4.3.1: In the forward propagation process of the dynamic workload prediction model, the dynamic time series features of the input are first calculated through the sparse attention mechanism. long-term dependency features, and obtain sparse attention output , and then regularize and normalization Processing, specifically:
[0068] ;
[0069] Sparse attention output obtained in step 4.3.2 and step 4.3.1 It is passed to the local attention mechanism and the local attention output is obtained after calculation , specifically:
[0070] ;
[0071] Step 4.3.3: Add the sparse attention output and the local attention output, and then regularize and normalization Processing to obtain the final combined output :
[0072]
[0073] Step 4.3.4: Combined output after fusion Feed the feedforward neural network into the dynamic workload prediction model to get the final complementary attention mechanism output , specifically:
[0074] .
[0075] A further improvement of the present invention is that the lightweight decoder in step 5 adopts a linear decoding layer to reduce the computational complexity, specifically:
[0076]
[0077] in, is the output after the linear decoding layer, is the prediction result.
[0078] The beneficial effects of this application are:
[0079] This application not only improves prediction accuracy, but also significantly improves computational efficiency, and is particularly suitable for processing complex time series data in dynamic environments.
[0080] The complementary attention mechanism of this application combines sparse attention and local attention, and has significant advantages in capturing long-term dependency features and short-term dependency features, thereby improving the performance of the model in processing dynamic load forecasting.
[0081] This application introduces a multi-scale temporal convolutional network (TCN) module before the encoder to extract multi-scale temporal features, which effectively enhances the model's ability to capture long-term and short-term dependency features, especially when dealing with local sudden changes and concept drift, improving the robustness and generalization ability of the model.
[0082] The lightweight decoder of the present application replaces the traditional multi-layer decoder structure with a single-layer linear decoder, significantly reducing the computational complexity and inference time, and improving the computational efficiency of the model on large-scale data sets.
[0083] This application effectively improves the accuracy and computational efficiency of dynamic workload prediction by combining complementary attention mechanisms, multi-scale temporal convolutional networks, and lightweight decoders, and is particularly suitable for complex dynamic data processing in distributed environments, cloud computing, and edge computing environments, thereby helping the system optimize resource allocation and improve performance and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Figure 1 It is a flow chart of this application.
[0085] Figure 2 It is a schematic diagram of the dynamic workload prediction model of this application. DETAILED DESCRIPTION
[0086] The following will disclose the implementation of the present application with drawings. For the purpose of clear description, many practical details will be described together in the following description. However, it should be understood that these practical details should not be used to limit the present application. In other words, in some implementations of the present application, these practical details are not necessary.
[0087] like Figure 1 As shown, the present application is a dynamic workload prediction method based on a data stream batch processing integrated architecture, which is implemented by a dynamic workload prediction model, and the dynamic workload prediction model includes an information embedding layer, a multi-scale time convolutional network layer, an encoding layer, and a linear decoding layer. The dynamic workload prediction method specifically includes the following steps:
[0088] Step 1: Input the static features in the bandwidth workload dataset into the linear decoding layer for mapping, and obtain a dense vector space with the same dimension as the dynamic time series features in the bandwidth workload dataset, which is effectively combined with the subsequent dynamic time series features. The specific operations are as follows:
[0089]
[0090] in, is a static feature, To map to the same dimension size as the dynamic time series feature dimension, that is, the static feature is mapped to the dynamic feature. is the index of the time step, () is a linear transformation. Among them, static features include server location, CPU number, memory size, etc.
[0091] Step 2: Fuse the static features and dynamic time series features to generate the final input features. The specific steps are as follows:
[0092] Step 2.1, value embedding: Map the dynamic time series features to a high-dimensional feature representation consistent with the dimension of the dynamic workload prediction model through a one-dimensional convolution operation to obtain value embedding, specifically:
[0093]
[0094] in, is the dynamic time series feature, is the value embedding after one-dimensional convolution;
[0095] Step 2.2, position embedding: Encode the position information of dynamic time series features through sine and cosine functions, and embed the encoded position information of dynamic time series features with values Add element by element to encode the position information for each dynamic time series feature and keep the time order, specifically:
[0096]
[0097]
[0098] in, is the time step index, is the dimension index of the position embedding, is the dimension size of the position embedding, The sine function encodes the position information of the dynamic time series features. The cosine function encodes the position information of the dynamic time series features. and Concatenate to get position embedding :
[0099] = ;
[0100] In this way, position embeddings can effectively provide position information for each time step, ensuring that the model can distinguish between inputs at different time steps.
[0101] Step 2.3, dynamic time series feature embedding: Encode periodic dynamic time series features such as hours and dates to enhance the dynamic workload prediction model's ability to perceive changes in dynamic time series features. Specifically:
[0102]
[0103] in, is the dynamic time series feature embedding, For the Embedding of dynamic time series features, is the number of dynamic time series features;
[0104] Step 2.4: The information embedding layer of the dynamic workload prediction model adds the value embedding of step 2.1, the position embedding of step 2.2, and the dynamic time series feature embedding of step 2.3 to form the final input feature :
[0105]
[0106] in, is the final input feature.
[0107] Step 3: Process the final input features generated in step 2 through a multi-scale time convolutional network layer to extract multi-scale time-dependent features, where the multi-scale time-dependent features include long-term dependency features and short-term dependency features.
[0108] Specifically, extracting multi-scale time-dependent features includes the following steps:
[0109] Step 3.1, the causal convolution layer ensures that the dynamic workload prediction model does not leak future information in the time series prediction task by limiting the convolution kernel to only depend on the current and previous time step inputs, and captures short-term dependency features. This design makes it particularly suitable for capturing short-term dependency features in time series, such as local trends and instantaneous changes. The operation of the causal convolution layer is expressed as:
[0110]
[0111] in, is the time step index The output, is the time step index Input, is the weight of the convolution kernel of the causal convolution layer, is the size of the convolution kernel of the causal convolution layer, is the position of the convolution kernel of the causal convolution layer in the time dimension, is the bias term;
[0112] Step 3.2: In order to further expand the receptive field, the multi-scale temporal convolutional network (TCN) uses dilated convolution. The dilated convolution layer expands the receptive field of the convolution and captures the long-term dependency features by inserting holes between the elements of the convolution kernel. The operation of the dilated convolution layer is expressed as:
[0113]
[0114] in, is the time step index Input, It is the dilation rate, which indicates the interval between the convolution kernel elements of the dilated convolution layer; it can capture long-term dependency features at a lower computational cost.
[0115] Step 3.3, the multi-scale temporal convolutional network layer includes multiple stacked TemporalBlock convolution blocks, each TemporalBlock convolution block contains a causal convolution layer and a hole convolution layer, an activation function and a residual connection to enhance the model's ability to extract multi-scale features of time series data; the output calculation formula of each TemporalBlock convolution block is:
[0116]
[0117] in, is the input, is the convolution operation, is the output after the activation function; the residual connection connects the input Adding it to the output of the convolutional layer helps alleviate the gradient vanishing problem and promotes the flow of information.
[0118] Step 3.4: To prevent overfitting, regularization is added after each TemporalBlock convolution block. Layers as a means of regularization:
[0119] .
[0120] Step 4: Send the multi-scale time-dependent features extracted in step 3 to the encoder, which captures long-term dependency features and short-term dependency features through a complementary attention mechanism to obtain a complementary attention mechanism output, where the complementary attention mechanism includes sparse attention and local attention. Specifically, the following steps are included:
[0121] Step 4.1: Sparse Attention The strategy limits the calculation scope of attention weight calculation and only focuses on the front time steps to capture long-term dependency features.
[0122] First, the sparse attention mechanism calculates the query vector and key vector The dot product between them gives the original attention score matrix :
[0123]
[0124] in, is the query vector, is the key vector, is the dimension of the key vector;
[0125] Secondly, the obtained attention score matrix conduct Filter, select the most important ones based on similarity or relevance attention connections, only retaining each query vector The corresponding front The maximum value is obtained, and irrelevant parts are ignored, that is, the remaining attention scores are set to zero. This can effectively reduce the complexity of the calculation and focus on the most informative time step, thereby improving the efficiency of the model in dealing with long-term dependencies. The calculation method is as follows:
[0126]
[0127] Finally, after Attention score after policy selection Used to weight the value vector , and get the output after the sparse attention mechanism:
[0128] .
[0129] Step 4.2: Local Attention adjusts the focus through a dynamic window mechanism time steps to capture short-term dependency features and sudden changes. Local attention can quickly capture these sudden changes by limiting its attention range and only calculating between local time steps. Specifically:
[0130] First, for each query vector The content is dynamically adjusted to the window size , so that the attention range is adaptive according to the actual data changes. Specifically, the window size Based on the current query vector The relationship between the historical time steps is adjusted to focus on the most relevant time steps. The calculation method is:
[0131]
[0132] in, is a learning function, based on the current query vector and the set of adjacent key vectors ( ) Calculate the appropriate dynamic window size ;
[0133] Secondly, based on the dynamic window size , the local attention mechanism is used for each current query vector Calculate the current query vector With the corresponding key vector The similarity is calculated only within the dynamic window range to obtain the attention score , the specific calculation process is as follows:
[0134]
[0135] in, and is the time step index determined by the dynamic window size, indicating the query vector The corresponding key vector The effective range of
[0136] Finally, the attention score Normalize and then the attention score With value vector Perform weighted summation to obtain the final complementary attention mechanism output, which is calculated as follows:
[0137] .
[0138] Step 4.3: Combine sparse attention and local attention mechanisms to capture long-term dependency features and short-term dependency features, and obtain complementary attention mechanism output. The combination of sparse attention and local attention mechanisms can capture long-term temporal dependency features and respond quickly to short-term fluctuations. The specific steps include the following:
[0139] Step 4.3.1: In the forward propagation process of the dynamic workload prediction model, the dynamic time series features of the input are first calculated through the sparse attention mechanism. long-term dependency features, and obtain sparse attention output , and then regularize and normalization Processing, specifically:
[0140] ;
[0141] Sparse attention output obtained in step 4.3.2 and step 4.3.1 It is passed to the local attention mechanism and the local attention output is obtained after calculation , specifically:
[0142] ;
[0143] Step 4.3.3: Add the sparse attention output and the local attention output, and then regularize and normalization Processing to obtain the final combined output :
[0144]
[0145] Step 4.3.4: Combined output after fusion Feed the feedforward neural network into the dynamic workload prediction model to get the final complementary attention mechanism output , specifically:
[0146] .
[0147] Step 5: Predict the complementary attention output obtained in step 4 through a lightweight decoder, where the lightweight decoder uses a linear decoding layer to reduce computational complexity, specifically:
[0148]
[0149] in, is the output after the linear decoding layer, is the prediction result.
[0150] In order to verify the advantages of this application in dynamic load prediction, the following experiment is provided.
[0151] 1. Experimental Setup
[0152] Datasets: This application uses multiple datasets for experimental evaluation, including:
[0153] ECW: From the edge computing environment, including the upload bandwidth workload changes of 797 edge servers in August 2022.
[0154] ECW-App-Switch: Records the workload changes of application switching from August 25, 2022 to August 30, 2022.
[0155] ECW-New-App: Contains workload changes for new applications from August 30, 2022 to September 6, 2022.
[0156] In addition, the dataset also provides static content data in 12 dimensions, such as maximum bandwidth, number of CPUs, location, etc., to assist model prediction.
[0157] Experimental details: The workload sequences of all ECW datasets are divided by hour, and the maximum value per hour is selected to match the time granularity of the application scheduling and billing rules. The input sequence length is set to T=48 and the prediction length is L=24. In order to train and evaluate the dynamic workload prediction model, the dataset is divided into training set, validation set, and test set in chronological order with a ratio of 6:2:2. All experiments are implemented in the PyTorch environment and conducted on a machine with an NVIDIA GeForce RTX 3090GPU. This application uses the ADAM optimizer for training, with an initial learning rate of 10^-4, a batch size of 256, and a training cycle of 50.
[0158] 2. Experimental Results
[0159] Evaluation indicators: In order to comprehensively evaluate the prediction effect of this application, this application uses mean square error (MSE) and mean absolute error (MAE) as evaluation indicators. The lower the value of MSE and MAE, the better the prediction effect of this application.
[0160] Benchmark model: This application selects five widely used time series forecasting methods as benchmark models for comparison, including: Deep Transformer, Informer, Autoformer (Transformer-based models); MQRNN (RNN-based model); and VaRDE-LSTM (clustering-based model).
[0161] Experimental results: As can be seen from Table 1, this application shows significant superiority on multiple datasets. In particular, on the ECW dataset, this application reduces the MSE by an average of 9.2%, which is better than other Transformer-based models. In comparison with the RNN-based MQRNN model, this application reduces the MSE on the ECW dataset by 19.5% and 94.8%, respectively, demonstrating its advantages in dealing with long-term dependencies and complex dynamic patterns.
[0162] Table 1 Comparison of average prediction performance of workloads
[0163]
[0164] In Table 1, lower MSE or MAE indicates better prediction results. The best performance on each dataset is indicated in bold.
[0165] The above experimental results prove the superiority of this application in multiple complex dynamic workload prediction tasks, especially in capturing long-term dependency features and local burst patterns, i.e., short-term dependency features. Compared with existing mainstream methods, this application has achieved lower MSE and MAE values on multiple data sets, demonstrating its powerful capabilities in multidimensional data modeling and complex dependency capture, and proving its great potential in practical applications.
[0166] In summary, this application solves the problems of insufficient global and local dependency modeling, limited multi-scale temporal feature modeling capabilities, the balance between computational efficiency and prediction performance, and insufficient adaptability to dynamic environments in the prior art through innovative complementary attention mechanisms, multi-scale temporal convolutional networks, and lightweight decoder designs. It effectively improves the accuracy and efficiency of dynamic workload prediction and has broad application prospects.
[0167] The above is only the implementation mode of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A dynamic workload prediction method based on a data stream batch processing integrated architecture, characterized in that: The dynamic workload prediction method is implemented by a dynamic workload prediction model, which includes an information embedding layer, a multi-scale time convolutional network layer, a coding layer, and a linear decoding layer. The dynamic workload prediction method specifically includes the following steps: Step 1: Input the static features in the bandwidth workload dataset into the linear decoding layer for mapping to obtain a dense vector space with the same dimension as the dynamic time series features in the bandwidth workload dataset; Step 2: Fuse the static features and dynamic time series features to generate the final input features; Step 3: Process the final input features generated in step 2 through a multi-scale time convolutional network layer to extract multi-scale time-dependent features, wherein the multi-scale time-dependent features include long-term dependency features and short-term dependency features; Step 4: Send the multi-scale time-dependent features extracted in step 3 to the encoder. The encoder captures the long-term dependency features and the short-term dependency features through the complementary attention mechanism to obtain the complementary attention mechanism output, where the complementary attention mechanism includes sparse attention and local attention. Step 5: Predict the complementary attention output obtained in step 4 through a lightweight decoder.
2. The dynamic workload prediction method based on the data stream batch processing integrated architecture according to claim 1 is characterized in that: In step 1, the static features are mapped through the static embedding input linear decoding layer to map the static features to a dense vector space of the same dimension as the dynamic time series features. The specific operation is: , in, is a static feature, As a result of mapping to the feature dimension of dynamic time series, is the time step index, () is a linear transformation.
3. The dynamic workload prediction method based on data stream batch processing integrated architecture according to claim 1 is characterized in that: In step 2, the static features and the dynamic time series features are fused, which specifically includes the following steps: Step 2.1, value embedding: Map the dynamic time series features to a high-dimensional feature representation consistent with the dimension of the dynamic workload prediction model through a one-dimensional convolution operation to obtain value embedding, specifically: , in, is the dynamic time series feature, is the value embedding after one-dimensional convolution; Step 2.2, position embedding: Encode the position information of dynamic time series features through sine and cosine functions, and embed the encoded position information of dynamic time series features with values Add element by element to encode the position information for each dynamic time series feature and keep the time order, specifically: , , in, is the time step index, is the dimension index of the position embedding, is the dimension size of the position embedding, The sine function encodes the position information of the dynamic time series features. The cosine function encodes the position information of the dynamic time series features. and Concatenate to get position embedding : = ; Step 2.3: Dynamic time series feature embedding: Encode the periodic dynamic time series features, specifically: , in, is the dynamic time series feature embedding, For the Embedding of dynamic time series features, is the number of dynamic time series features; Step 2.4: The information embedding layer of the dynamic workload prediction model adds the value embedding of step 2.1, the position embedding of step 2.2, and the dynamic time series feature embedding of step 2.3 to form the final input feature : , in, is the final input feature.
4. The dynamic workload prediction method based on data stream batch processing integrated architecture according to claim 1 is characterized in that: In step 3, the multi-scale temporal convolutional network layer uses a causal convolutional layer and a dilated convolutional layer to extract multi-scale time-dependent features through a multi-layer convolutional module, specifically including the following steps: Step 3.1, the causal convolution layer captures short-term dependency features by limiting the convolution kernel to only depend on the current and previous time step inputs. The operation of the causal convolution layer is expressed as: , in, is the time step index The output, is the time step index Input, is the weight of the convolution kernel of the causal convolution layer, is the size of the convolution kernel of the causal convolution layer, is the position of the convolution kernel of the causal convolution layer in the time dimension, is the bias term; Step 3.2: The atrous convolution layer expands the receptive field of the convolution and captures the long-term dependency features by inserting holes between the elements of the convolution kernel. The operation of the atrous convolution layer is expressed as: , in, is the time step index Input, is the dilation rate, which indicates the spacing between the convolution kernel elements of the dilated convolution layer; Step 3.3, the multi-scale temporal convolutional network layer includes multiple stacked TemporalBlock convolution blocks, each TemporalBlock convolution block includes a causal convolution layer and a hole convolution layer, an activation function and a residual connection, and the output calculation formula of each TemporalBlock convolution block is: , in, is the input, is the convolution operation, is the output after being processed by the activation function; Step 3.4: Regularization is added after each TemporalBlock convolution block. Layers as regularization means: 。 5. The method for dynamic workload prediction based on data stream batch processing integrated architecture according to claim 1, characterized in that: In step 4, the encoder captures long-term dependency features and short-term dependency features through a complementary attention mechanism, which specifically includes the following steps: Step 4.1: Sparse attention The strategy limits the calculation scope of attention weight calculation and only focuses on the front time steps to capture long-term dependency features; Step 4.2: Local attention adjusts the focus through the dynamic window mechanism time steps to capture short-term dependency features; Step 4.3: Combine sparse attention and local attention to capture long-term dependency features and short-term dependency features to obtain the output of the complementary attention mechanism.
6. The method for dynamic workload prediction based on data stream batch processing integrated architecture according to claim 5, characterized in that: The step 4.1 specifically includes the following steps: Step 4.1.1: The sparse attention mechanism calculates the query vector and key vector The dot product between them gives the original attention score matrix : , in, is the query vector, is the key vector, is the dimension of the key vector; Step 4.1.2: The attention score matrix obtained in step 4.1.1 conduct Screening, the calculation method is as follows: ; Step 4.1.3, after Attention score after policy selection Used to weight the value vector , and get the output after the sparse attention mechanism: 。 7. The method for dynamic workload prediction based on data stream batch processing integrated architecture according to claim 5, characterized in that: Step 4.2 specifically includes the following steps: Step 4.2.1: For each query vector The content is dynamically adjusted to the window size , so that the attention range adapts according to the actual data changes, and the calculation method is: , in, is a learning function, based on the current query vector and the set of adjacent key vectors ( ) Calculate dynamic window size ; Step 4.2.2: Based on dynamic window size , the local attention mechanism is used for each current query vector Calculate the current query vector With the corresponding key vector The similarity is calculated only within the dynamic window range to obtain the attention score , the specific calculation process is as follows: , in, and is the time step index determined by the dynamic window size, indicating the query vector The corresponding key vector The effective range of Step 4.2.3: Attention score Normalize and then the attention score With value vector Perform weighted summation to obtain the final complementary attention mechanism output, which is calculated as follows: 。 8. The method for dynamic workload prediction based on data stream batch processing integrated architecture according to claim 5, characterized in that: The step 4.3 specifically includes the following steps: Step 4.3.1: In the forward propagation process of the dynamic workload prediction model, the dynamic time series features of the input are first calculated through the sparse attention mechanism. long-term dependency features, and obtain sparse attention output , and then regularize and normalization Processing, specifically: ; Sparse attention output obtained in step 4.3.2 and step 4.3.1 It is passed to the local attention mechanism and the local attention output is obtained after calculation , specifically: ; Step 4.3.3: Add the sparse attention output and the local attention output, and then regularize and normalization Processing to obtain the final combined output : ; Step 4.3.4: Combined output after fusion Feed the feedforward neural network into the dynamic workload prediction model to get the final complementary attention mechanism output , specifically: 。 9. The method for dynamic workload prediction based on data stream batch processing integrated architecture according to claim 1, characterized in that: The lightweight decoder in step 5 adopts a linear decoding layer to reduce the computational complexity, specifically: , in, is the output after the linear decoding layer, is the prediction result.
Citation Information
Patent Citations
Highway traffic flow prediction method based on Transform and graph attention network
CN116092294A
Edge cloud device load waveform prediction method and system
CN116225710A
Long sequence knowledge tracking method based on Informer
CN118350418A
Cited By
Efficient real-time big data stream processing method and system
CN120596268A
A Highly Efficient Real-Time Big Data Stream Processing Method and System
CN120596268B
Fault monitoring and alarming method and system for power equipment
CN121350710A