Random workflow load prediction method for data missing

By adopting the kernel density missing data filling method and Kalman filtering technology based on common features in random workflow scheduling, combined with the N-BEATS model for load prediction, the problem of data missing affecting workflow prediction is solved, and more efficient and accurate load prediction is achieved.

CN120067554AActive Publication Date: 2025-05-30SOUTHEAST UNIV

Patent Information

Application Number
CN202510219671.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In random workflow scheduling, the lack of data hinders the effective workload prediction of the workflow and reduces workflow scheduling performance, especially in the face of complex nonlinear patterns.

Method used

A method of filling the kernel density missing data based on common characteristics is proposed. Combined with Kalman filtering technology, the weights of the kernel density estimate are dynamically updated, and the N-BEATS model is used to predict the cluster load to improve the robustness and generalization of the prediction.

Benefits of technology

By filling in missing data and dynamically updating weights, the accuracy and efficiency of workflow load prediction is improved, and the fluctuations and periodicity in cluster load can be better captured, and the computing efficiency and the interpretability of predictions can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067554A_ABST
    Figure CN120067554A_ABST
Patent Text Reader

Abstract

The invention discloses a random workflow load prediction method for data missing, and the method comprises the steps: a workflow preprocessing stage: converting a workflow into an embedded vector form, and carrying out the similarity clustering according to historical records; a workflow missing value completion and load prediction stage: filling missing data by a kernel density missing data filling method based on common characteristics, and proposing a workflow load analysis and prediction method based on network cloud edge structure characteristics and historical information such as workflow calculation, storage, communication and the like; a cluster modeling and feature engineering stage; according to the cluster load prediction stage, the kernel method and the Kalman filtering technology are combined to predict the missing value and the workflow load, the prediction robustness and generalization are optimized, and the cluster load prediction method has wide application value and use prospects in the fields of load prediction and workflow scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field The present invention relates to a random workflow load prediction method for data missing, belonging to the technical fields of data analysis, deep learning, load prediction, workflow, etc. Background Art In random workflow scheduling, the missing data hinders the effective workload prediction of the workflow and reduces the workflow scheduling performance. For the problem of missing data, many filling prediction algorithms have been proposed to handle the missing values, mainly based on supervised learning, that is, by using the remaining complete data to construct a prediction model to fill the missing values. Ma et al. [1] studied the correlation between missing values and non-missing values using a deep denoising autoencoder. Yoon et al. [2] developed a generative adversarial imputation network (GAIN), in which the discriminator aims to distinguish whether the value is real or imputed. Fortuin et al. [3] used a deep variational autoencoder and Gaussian process to transform incomplete time series data into a low-dimensional space for better prediction. In addition, Spinelli et al. [4] used a graph imputation neural network to construct the graph similarity between complete and incomplete data sets. Pan [5] et al. proposed an adaptive learning median filling deep autoencoder (AM-DAE) aiming to fill the missing values of industrial time series data in an unsupervised manner. It continuously replaces the missing values with the median of the input data and its reconstruction, thus allowing the filling information to be transmitted together with the training process. These methods have demonstrated the excellent expressive ability of deep learning-based methods. Cluster load prediction is a key technology in cloud computing and data center management. Accurately predicting the usage of cluster resources has become the key to resource scheduling and optimization. Scholars at home and abroad have proposed various prediction methods for this problem, from traditional statistical methods to deep learning models, continuously improving the prediction accuracy and efficiency. Early research mainly used statistical methods and traditional machine learning models for load prediction. Martinez et al. [6] used the Kalman filter to predict the CPU usage of nodes in a serverless environment. Deepika et al. [7] used a Multilayer Perceptron (MLP) regression model to predict the power consumption of virtual machines. Iqbal et al. [8] proposed an adaptive model selector that dynamically selects the most suitable prediction model from Linear Regression (LR), Support Vector Machine (SVM), Gradient Boosting Decision Tree (GBDT), and Gaussian Process (GP) through Random Forest (RF) to estimate the CPU resource consumption in the data center, especially for time series and bursty behaviors. These traditional methods have achieved load prediction to a certain extent, but they perform poorly in the face of complex non-linear patterns.

[0001] Q. Ma, W. Lee, T. Fu, Y. Gu and G. Yu, "MIDIA: Exploring denoising autoencoders for missing data imputation", Data Min. Knowl. Disc., vol. 34, no. 6, pp. 1859 - 1897, 2020.

[0002] J. Yoon, J. Jordon and M. van der Schaar, "GAIN: Missing data imputation using generative adversarial nets", arxiv:1806.02920, 2018.

[0003] V. Fortuin, D. Baranchuk, G. and S. Mandt, "GP-VAE: Deep probabilistic time series imputation", arxiv:1907.04155, 2019.

[0004] I. Spinelli, S. Scardapane, M. Scarpiniti and A. Uncini, "Efficient data augmentation using graph imputation neural networks" in Progresses in Artificial Intelligence and Neural Systems, Singapore: Springer, pp. 57-66, 2021.

[0005] Z. Pan, Y. Wang, K. Wang, H. Chen, C. Yang and W. Gui, "Imputation of Missing Values in Time Series Using an Adaptive-Learned Median-Filled Deep Autoencoder," in IEEE Transactions on Cybernetics, vol. 53, no. 2, pp. 695-706, Feb. 2023, doi:10.1109 / TCYB.2022.3167995.

[0006] Martinez, M. M., & Pandey, S. R. (2022, March). Predictive function placement for distributed serverless environments. In 2022 25th Conference on Innovation in Clouds, Internet and Networks (ICIN) (pp. 86-90). IEEE.

[0007] Deepika, T., & Prakash, P. (2020). Power consumption prediction in cloud data center using machine learning. Int. J. Electr. Comput. Eng. (IJECE), 10(2), 1524-1532.

[0008] Iqbal, W., Berral, J. L., Erradi, A., & Carrera, D. (2019). Adaptive prediction models for data center resources utilization estimation. IEEE Transactions on Network and Service Management, 16(4), 1681 - 1693. Summary of the Invention Technical Problem: Facing the complex diversity of workflow tasks, in order to further improve the robustness and generalization of stochastic workflow load prediction, the present invention considers the situation where there are missing values in the stochastic workflow, uses its historical records, and proposes a kernel density missing data filling method based on common features. On this basis, workflow load prediction is carried out, and a cluster load prediction method for stochastic workflows facing data missing is provided for the dynamicity and resource competition problems in the shared GPU cluster environment. Technical Solution: To achieve the above object, the technical solution adopted by the present invention is: a stochastic workflow load prediction method facing data missing, including the following stages: A Workflow Preprocessing Stage: Convert the workflow into the form of an embedded vector, based on the historical data of the stochastic workflow of the computing power network, memorize and learn the time series data information, extract the time series features of the requests, and perform spectral clustering based on common features through three convolutional neural network models responsible for the head node, tail node, and relationship respectively. B Workflow Missing Value Completion and Load Prediction Stage: Construct a kernel density model for each type of request obtained by clustering, combine with the Kalman filter to dynamically update the weights of the kernel density estimation, and at the same time use the maximum likelihood estimation to calculate the recursive characteristics of the filter to obtain the kernel bandwidth and other dynamic control parameters, and predict the stochastic workflow load based on methods such as calculating the expectation of the probability density function. C Cluster Modeling and Feature Engineering Stage: Convert the cluster load x t into a high - dimensional feature vector x containing various time series features t for more accurate prediction and analysis. D Cluster Load Prediction Stage: Use the N - BEATS model to predict the cluster load. The N - BEATS model normalizes the input load data X, and then uses the STL decomposition technique to decompose the time series into a linear combination of basis functions, including three parts: trend, seasonality, and residual. Each basis function is learned and represented by a neural network block, and the prediction result can be formed by separately modeling the above components. The specific steps of the workflow preprocessing stage are as follows: A1. The directed acyclic graph G(V, E) represents a workflow, where V represents the nodes in the directed acyclic graph, E represents the adjacency matrix of the graph, and the graph information is transformed into an embedding vector W; A2. Based on the random workflow historical data T of the computing power network W , a neural network model is constructed to memorize the information of neurons at the current moment and learn the dependency relationship of time series data; A3. Extract the time series feature t of each request arrival W ; A4. The workflow demand similarity metric d based on information entropy W = d(t W , T W ), and spectral clustering is performed based on the one-dimensional time series feature t W to find tasks with high similarity Y = G p (t W , T W ). The specific steps of the workflow missing value completion and load prediction phase are as follows: B1. Construct three convolutional neural network models, which are responsible for completing the head node, tail node, and dependency relationship respectively. The convolutional neural network structure is as follows: B1.1 Link the two input embedding vectors Y i and Y j into a two-dimensional array or three-dimensional array, and use a fusion algorithm to make the two embedding vectors fully fuse with each other to facilitate information interaction under the action of the convolution kernel; B1.2 Output the edge embedding vector W E or the node embedding vector W G through a fully connected layer, and this vector is transformed into a workflow task node or an inter-task dependency relationship to complete the missing data information of the workflow; B1.3 When training the model, a workflow is processed by three convolutional neural network models at the same time, and the three models use the same loss function and are trained synchronously; B2. For each class of requests obtained by clustering, a kernel density model K is constructed respectively, and the original calculation sequence of the workflows in each class is used as a data record in the kernel density prediction model of this class; B3. By using techniques such as linear filtering, exponential weighted filtering, and Kalman filtering, the weights are automatically updated through the covariance matrix C at each time point. To estimate the time-varying density, Kalman filtering and kernel density estimation are combined, and the weights of the dynamic kernel density estimation are updated by the method of Kalman filtering; The application of maximum likelihood estimation is as follows: B4.1 Obtain the estimated values of the kernel bandwidth MaxL(h) and other control parameters according to the maximized likelihood function; B4.2 Calculate the estimated values of the probability density function, cumulative distribution function, and quantiles through these estimated values; B5 Predict the random workflow load E(w t ) based on methods such as calculating the expectation of the probability density function. The specific steps of the cluster modeling and feature engineering stage are as follows: C1. Define the cluster load x t as a sequence that changes over time, where t represents the timestamp: X = {x 1 , x 2 , x 3 ,..., x t}; C2. Extract features from the timestamp t, such as month, day, hour, minute, weekday, etc., which can be expressed as: T t = {month(t), day(t), hour(t), minute(t), weekday(t)}; C3. Use the sliding window technique to generate lag features. For example, for a window size of k, the lag features can be expressed as: L t = {x t-1 , x t-2 ,..., x t-k}; C4. Calculate the statistics within the sliding window, considering the mean μ t , median m t and standard deviation σ t : C5. Combine the original cluster load x t with the extracted time features T t , lag features L t and statistical features {μ t , m t , σ t} to form a high-dimensional feature vector X t : X t = {x t , T t , L t , μ t , m t , σ t} X = {X 1 , X 2 ,..., Xt} In this way, the cluster load x t is transformed into a high-dimensional feature vector X containing various time series features t , and the advantages are as follows: 1. Comprehensively capture the load dynamics of the system, avoid the limitations of relying only on a single indicator, and provide more accurate prediction and analysis for load prediction; 2. Learn more potential patterns, learn the impacts of seasonality, periodicity or sudden events, so as to capture more complex non-linear relationships and improve the prediction ability of the model and the intelligent level of the system; 3. By making different combinations and selections of high-dimensional features, it can help the model focus on the most important load features, reduce redundant features, and improve the training and inference efficiency. The specific steps in the cluster load prediction stage are as follows: D1. Normalize the input load data X: X = Normalization(X); D2. Use a feedforward neural network to extract features through a block structure, and use the STL decomposition technique to decompose the time series into a linear combination of basis functions, including three parts: trend, seasonality, and residual: trend,seasonal,residual = STL Decomposition(X); D3. Use specific activation functions and models to model the above components: For the prediction of the trend part trend pred It is expressed as: trend pred = W t2 (ELU(W t1 x + b t1 )) + b t2 , where ELU is the exponential linear unit, W t1 , b t1 are the weights and biases of ELU, and W t2 , b t2 are the weights and biases for trend prediction. For the prediction of the seasonal part seasonal pred It is expressed as: seasonal pred = W s2 (GLU(W s1 x + b s1 )) + b s2 , GLU is the gated linear unit, W s1 , bs1 are the weights and biases for GLU, W s2 , b s2 are the weights and biases for seasonal prediction. For the prediction of the residual part, residual pred is expressed as: residual pred = W r2 (ELU(W r1 x + b r1 )) + b r2 , GLU(x) = (σ(W g1 x + b g1 )) ⊙ (W g2 x + b g2 ), where σ is the Sigmoid activation function, ⊙ represents the Hadamard product, W r1 , b r1 are the weights and biases for the Sigmoid activation function, W r2 , b r2 are the weights and biases for residual prediction. D4. Predicting the time series is expressed as the sum of three: Decomposing the time series prediction into trend, seasonality, and residual and then processing and integrating them separately has several significant advantages: 1. Improving the interpretability of the model: The trend component reflects the long-term changes in the data, the seasonal component reveals the periodic fluctuations, and the residual component represents the random noise. Through this decomposition, the model can more clearly represent the dynamic characteristics of the time series. 2. Enhancing the prediction ability of the model: The trend part usually shows changes on a longer time scale, and seasonality reflects short-term periodic fluctuations. Modeling and optimizing these two components separately can enable the model to capture these characteristics more accurately. 3. Improving the robustness of the model: After decomposing the data, each component uses different modeling techniques, enabling each part to be optimized according to its characteristics, thereby improving the robustness and accuracy of the overall model. Beneficial effects: The random workflow load prediction method for data missing provided by the present invention has the following beneficial effects compared with the prior art: (1) The present invention predicts the future workflow itself and the cluster load through the designed time series prediction model, which not only considers the task uncertainty in the cluster workflow scheduling process but also reasonably refers to the historical load information for load prediction. The prediction results can be used as an important basis for task priority determination and scheduling strategies. (2) The present invention combines a kernel density missing data filling method based on common features and Kalman filtering technology to provide a powerful tool for estimating and predicting missing values in time series, providing accurate assistance for workflow load prediction. (3) The N-BEATS model used in the present invention can adaptively capture various fluctuations and periodicities in the clustered time series, and adopts a highly parallelized design, which can improve the computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 is the flowchart of the method of the present invention, Figure 2 is the overall process schematic diagram of the present invention. Figure 3 is the schematic diagram of the missing value filling network of the workflow of the method of the present invention, Figure 4 is the schematic diagram of the clustered load prediction stage of the method of the present invention, Figure 5 is the schematic diagram of the data set used for the load prediction problem, DETAILED DESCRIPTION OF THE EMBODIMENTS The present invention will be further clarified below in conjunction with the drawings and specific implementation cases. It should be understood that these examples are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications of the present invention by those skilled in the art all fall within the scope defined by the appended claims of this application. Example: As Figure 1 shown is the flowchart of the random workflow load prediction method for data missing, Figure 2 is the overall process schematic diagram of the present invention. When performing clustered load prediction, the N-BEATS model is used to predict the load in the future time period, enhancing the prediction accuracy and efficiency. A. Workflow preprocessing stage: Convert the information in the workflow into an embedded vector W, and then divide the request log records in the cloud center into historical data and prediction data according to the requirements of the workflow. Based on the historical data T of the random workflow of the computing power network W , construct a neural network model, record the information of the neurons at the current moment, and learn the temporal dependence. Extract the time features of each request, and for the one-dimensional temporal feature t of each request sequence W According to the similarity d W = d(t W , T W ) for clustering, and obtain the tasks Y = G p (t W , T W ) with high similarity in the temporal feature space. B. Workflow Missing Value Completion and Load Prediction Phase: B1. Figure 3 The following figure shows the schematic diagram of the workflow missing value filling network. Three convolutional neural network models are constructed, which are responsible for completing the head node, tail node, and dependency relationship respectively: The structure of the convolutional neural network is as follows: B1.1 links the two input embedding vectors Y i and Y j into a two-dimensional or three-dimensional array, and uses a fusion algorithm to fully fuse these two embedding vectors with each other to facilitate information interaction under the action of the convolution kernel; B1.2 outputs the edge embedding vector W E or the node embedding vector W G through a fully connected layer. This vector is transformed into a workflow task node or an inter-task dependency relationship to complete the missing data information of the workflow; B1.3 When training the model, a workflow is processed by three convolutional neural network models at the same time, and the three models use the same loss function and are trained synchronously; B2. For each type of request obtained by clustering, a kernel density model K is constructed respectively, and the original calculation sequence of the workflows in each type is used as a data record in the kernel density prediction model of this type; B3. By using techniques such as linear filtering, exponential weighted filtering, and Kalman filtering, the weights are automatically updated through the covariance matrix C at each time point. To estimate the time-varying density, Kalman filtering and kernel density estimation are combined, and the weights of the dynamic kernel density estimation are updated by the method of Kalman filtering: B4. The application of maximum likelihood estimation is as follows: B4.1 Obtain the estimated values of the kernel bandwidth MaxL(h) and other control parameters according to the maximized likelihood function; B4.2 Calculate the estimated values of the probability density function, cumulative distribution function, and quantile through these estimated values; B5 Predict the random workflow load E(w t ) based on methods such as calculating the expectation of the probability density function. C. Cluster Modeling and Feature Engineering Phase: Assume that the cluster load at time t, that is, the proportion of the number of busy GPUs to the total number of GPUs, is x t , and its value range is [0, 1]. Define x t as a sequence that changes with time, where t represents the timestamp: X = {x 1 , x 2 , x 3 ,..., x t ,}. At the same time, relevant time features are extracted from the timestamp t, such as month, day, hour, minute, week, etc., which are expressed as: T t= {month(t), day(t), hour(t), minute(t), weekday(t)}. After obtaining x t a sliding window technique needs to be used to generate lag features. For a window size of k, the lag features are represented as: L t = {x t-1 , x t-2 ,..., x t-k}. Considering the mean the median m t = median(x t-1 , x t-2 ,..., x t-k ) and the standard deviation Combining the original cluster load x t with the extracted time features T t , the lag features L t and the statistical features {μ t , m t , σ t} to form a high-dimensional feature vector X t = {x t , T t , L t , μ t , m t , σ t}. Then the high-dimensional feature vectors X at all times can be represented as X = {X 1 , X 2 ,..., X t}. D. Cluster load prediction stage: Figure 4 The schematic diagram of the cluster load prediction stage is shown as follows: D1. Normalize the input load data X: X = Normalization(X); D2. Use a feedforward neural network to extract features through a block structure, and use the STL decomposition technique to decompose the time series into a linear combination of basis functions, including three parts: trend, seasonal, and residual: trend, seasonal, residual = STL Decomposition(X); D3. Use specific activation functions and models to model the above components: trend pred = W t2 (ELU(W t1 x + b t1 )) + b t2 , seasonl pred = W s2 (GLU(W s1 x + b s1 )) + b s2, residual pred = W r2 (ELU(W r1 x + b r1 )) + b r2 , GLU(x) = (σ(W g1 x + b g1 )) ⊙ (W g2 x + b g2 ), where GLU is the gated linear unit, ELU is the exponential linear unit, σ is the Sigmoid activation function, and ⊙ represents the Hadamard product; D4. Predicting time series is expressed as the sum of three parts: Regarding the data faced in this example as Figure 5 shown, focusing on the dynamic scheduling problem of deep learning tasks in large enterprises or research institutions, using the workflow data and cluster trace data generated by them. It should be noted that the above embodiments are not used to limit the protection scope of the present invention. Equivalent transformations or substitutions made on the basis of the above technical solutions all fall within the protection scope of the claims of the present invention.

Claims

1. A random workflow load prediction method for data missing, characterized by: The following phases are included: A. Workflow preprocessing stage: The workflow is converted into an embedded vector. Based on the historical data of random workflows in the computing network, the time series data information is memorized and learned. The time series features of the requests are extracted and spectral clustering is performed based on common features through three convolutional neural network models responsible for the head node, tail node and relationship respectively. B. Workflow missing value completion and load prediction stage: For each type of request obtained by clustering, a kernel density model is constructed separately. Combined with Kalman filtering, the weights of the kernel density estimation are dynamically updated. At the same time, the recursive characteristics of the filter are calculated using maximum likelihood estimation to make the kernel bandwidth and other dynamic control parameters, and the random workflow load is predicted based on methods such as calculating the expected probability density function. C. Cluster modeling and feature engineering phase: Cluster load x t Transformed into a high-dimensional feature vector x containing multiple time series features t , in order to make more accurate predictions and analysis, D. Cluster load prediction stage: The N-BEATS model is used to predict the cluster load. The N-BEATS model normalizes the input load data X, and then uses the STL decomposition technology to decompose the time series into a linear combination of basis functions, including trend, seasonality and residual. Each basis function is learned and represented by a neural network block. The prediction results can be modeled by the above components separately.

2. The method for predicting random workflow load with data loss according to claim 1, characterized in that: The specific steps of the workflow preprocessing stage are as follows: A1.Convert workflow information into embedded vector form; A2. Based on the historical data of random workflows in the computing network, a neural network model is constructed to memorize the information of neurons at the current moment and learn the dependency relationship of time series data; A3. Extract the timing characteristics of each request arrival; A4. Workflow requirement similarity measurement based on information entropy and spectral clustering based on one-dimensional time series characteristics.

3. The method for predicting random workflow load with data loss according to claim 2, characterized in that: The specific steps of the workflow missing value completion and load prediction phase are as follows: B1. Build three convolutional neural network models, which are responsible for completing the head node, tail node and relationship respectively; B2. Build a kernel density model for each type of request obtained by clustering, and the original calculation sequence of the workflow in each type is used as a data record in the kernel density prediction model of that type; B3. Automatically update weights through the covariance matrix at each time point by using linear filtering, exponential weighted filtering, Kalman filtering and other techniques; B4. Calculate kernel bandwidth and other dynamic control parameters by maximum likelihood estimation; B5. Predict random workflow loads based on methods such as calculating the expected probability density function.

4. The method for predicting random workflow load with missing data according to claim 3, characterized in that: The specific steps of the cluster modeling and feature engineering stage are as follows: C1. Define time series, C2. Extract time features, C3. Create hysteresis features, C4. Calculate statistical features, C5. Construct high-dimensional features.

5. The method for predicting random workflow load with data loss according to claim 4, characterized in that: The specific steps of the cluster load prediction stage are as follows: D1. Normalize the input load data X; D2. Use feedforward neural network to extract features through block structure and use STL decomposition technique to decompose time series into linear combinations of basis functions, including trend, seasonality and residual. D3. Model the above components using specific activation functions and models; D4. Forecasting time series It is expressed as the sum of the three.

Citation Information

Patent Citations

  • Load prediction method for cloud computation cluster tasks based on cluster characteristic extraction

    CN108415777A

  • Data processing method for intelligent gateway protocol

    CN118740953A

  • Transformation method and device of novel protective transformer

    CN118801394A

  • Systems for predictive data analytics, and related methods and apparatus

    WO2018075995A1

Cited By

  • Fusion device state monitoring method, fusion device state monitoring system, fusion device state monitoring equipment and storage medium

    CN121167435A