Multi-task monitoring and scheduling method and system

Through deep learning technology combined with kernel density estimation and Markov decision model, the shortcomings of task execution time prediction and anomaly detection in the existing technology are solved, the task scheduling strategy is optimized, and the system prediction accuracy and scheduling efficiency are improved.

CN119576505BActive Publication Date: 2025-05-23北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510131703.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-23
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

The prior art has shortcomings in task execution time prediction, abnormal detection and task scheduling, and it is difficult to accurately capture the dynamic characteristics and resource competition relationships during task execution, resulting in low prediction accuracy and low scheduling efficiency.

Method used

The multi-task monitoring and scheduling method based on deep learning is adopted to achieve accurate prediction of task execution time through kernel density estimation and dual convolutional networks, and the accuracy of abnormal detection is improved by using hierarchical Markov decision model and spectral clustering algorithm. The task rescheduling strategy is optimized based on task resource competition and improved minimum spanning tree algorithm, and a prediction error feedback mechanism is established to achieve dynamic optimization of the model.

Benefits of technology

It improves the accuracy of task execution time prediction and task scheduling efficiency, enhances the system's fault tolerance and reliability, and realizes accurate identification and timely discovery of task abnormal states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119576505B_ABST
    Figure CN119576505B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-task monitoring and scheduling method and system, which relates to the technical field of task scheduling, including receiving a task submission request, building a task execution time prediction model based on a historical execution time series and a bidirectional gated recurrent neural network; monitoring the task execution process, and identifying abnormal tasks using a hierarchical Markov decision and spectral clustering algorithm; for abnormal tasks, building a task dependency graph based on resource competition and calculating a critical path to determine the optimal computing node for re-execution; the present invention can accurately predict the task execution time, timely discover abnormal tasks and perform resource optimization scheduling, thereby improving system operation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of task scheduling, and in particular to a multi-task monitoring and scheduling method and system. Background Art

[0002] With the widespread application of large-scale distributed computing systems, the number and complexity of tasks running in the system are constantly increasing, which puts higher requirements on the monitoring and scheduling of task execution. Existing multi-task monitoring systems usually use fixed thresholds and rule matching to detect anomalies, and perform task scheduling based on static load balancing strategies. At the same time, in terms of task execution time prediction, traditional methods mainly rely on statistical analysis of historical data, which makes it difficult to accurately capture the dynamic characteristics and resource competition relationships during task execution.

[0003] However, there are several problems with the existing technology. The fixed anomaly detection threshold is difficult to adapt to the execution characteristics of different types of tasks, and is prone to false positives or negatives. The existing task execution time prediction method does not fully consider the noise interference in the time series data, and the resource competition relationship between tasks is not considered enough, resulting in low prediction accuracy. The traditional task scheduling method often only considers the system load balancing, ignores the resource competition relationship and execution dependency between tasks, and it is difficult to find the optimal task rescheduling scheme. There is a lack of effective feedback mechanism for the abnormal task re-execution process, and it is impossible to dynamically optimize the prediction model and scheduling strategy according to the actual execution situation.

[0004] In summary, there is an urgent need for a multi-task monitoring and scheduling method based on deep learning. Through kernel density estimation and dual convolutional networks, accurate prediction of task execution time is achieved. The hierarchical Markov decision model and spectral clustering algorithm are used to improve the accuracy of anomaly detection. The task rescheduling strategy is optimized based on task resource competition and an improved minimum spanning tree algorithm. A prediction error feedback mechanism is established to realize dynamic optimization of the model, thereby solving the above-mentioned technical problems existing in the prior art. Summary of the invention

[0005] The embodiment of the present invention provides a multi-task monitoring and scheduling method and system, which can solve the problems in the prior art.

[0006] According to a first aspect of the embodiments of the present invention,

[0007] A multi-task monitoring and scheduling method is provided, comprising:

[0008] Receive a task submission request carrying task identification information, divide the task group according to the task identification information, collect the historical execution time series of the same task group, execute an adaptive filtering algorithm based on kernel density estimation, perform noise reduction processing on the historical execution time series to obtain an initial feature sequence, input the initial feature sequence into a temporal feature extraction network including a causal convolution layer and a hole convolution layer, extract the task feature vector, and based on the task feature vector, construct a task execution time prediction model for predicting a maximum execution time threshold through a bidirectional gated recurrent neural network;

[0009] Monitor the execution process of the task, collect the task execution time and real-time operation data, when the task execution time exceeds the maximum execution time threshold, construct the task operation state space based on the real-time operation data, use the hierarchical Markov decision model to model the task state, convert the real-time operation data into a task state sequence according to the time series, use the Bayesian inference method to calculate the state transition probability, based on the state transition probability, use the spectral clustering algorithm to segment the task state sequence, and combine the dynamic time warping algorithm to identify abnormal state segments, and when the abnormal state segment is identified, determine the task as an abnormal task;

[0010] For the abnormal task, a task resource competition matrix is ​​constructed based on the real-time operation data, a weighted directed graph is constructed based on the task resource competition as a task dependency graph, the resource critical path is calculated using an improved minimum spanning tree algorithm, and the optimal computing node for re-executing the abnormal task is determined by a task sorting algorithm based on the resource critical path; the re-execution trajectory information of the abnormal task is recorded, the Pearson correlation coefficient of the predicted value and the actual value of the task execution time prediction model is calculated to determine the prediction error, and the prediction error is fed back to the different network layers of the task execution time prediction model by a time series decomposition algorithm for optimization; when the number of re-executions of the abnormal task reaches a preset maximum number of retries, the abnormal task is terminated and a failure analysis report is returned.

[0011] In an optional embodiment,

[0012] Executing an adaptive filtering algorithm based on kernel density estimation to perform noise reduction processing on the historical execution time series to obtain an initial feature sequence including:

[0013] Divide the historical execution time series into multiple discrete time points, each of the discrete time points corresponds to a time series data value in the historical execution time series, and the time series data value and the discrete time point together constitute a time series data point;

[0014] Performing kernel density estimation on the historical execution time series, calculating the probability density value of each time series data point at the discrete time point by using a Gaussian kernel function, determining the bandwidth parameter of the Gaussian kernel function based on the standard deviation and interquartile range of the historical execution time series, and obtaining the probability density distribution of the historical execution time series;

[0015] Based on the probability density distribution, the probability density value is used as a density estimation weight to calculate the local mean of the historical execution time series at each discrete time point, and the local variance value of the historical execution time series is calculated based on the local mean and the probability density value;

[0016] Normalizing the local variance value based on the probability density value, calculating the ratio of the normalized local variance value to the global standard deviation of the historical execution time series, determining an adjustment coefficient according to the ratio, and dynamically determining a filter window size for each discrete time point based on a preset basic window size and in combination with the adjustment coefficient;

[0017] For each of the discrete time points, within the corresponding filtering window, the time series data points are selected to construct a cubic polynomial fitting model, and a regression weight function is constructed based on the probability density value, wherein the regression weight function is the product of a hyperbolic function form and the probability density value, and the weight value of the regression weight function decreases cubically as the distance between the time series data point and the corresponding discrete time point increases;

[0018] Using the regression weight function, weighted processing is performed on the time series data points within the filter window, coefficients of the cubic polynomial fitting model are solved by weighted least squares method, and a correction value at each discrete time point is obtained according to the coefficients;

[0019] The local standard deviation of the time series data points in the filtering window is calculated in combination with the probability density value, and the time series data points whose time series data values ​​deviate from the local mean by more than twice the local standard deviation are determined as abnormal data points. The time series data values ​​of the abnormal data points are replaced with the corresponding correction values ​​to obtain the initial feature sequence.

[0020] In an optional embodiment,

[0021] Inputting the initial feature sequence into a temporal feature extraction network including a causal convolution layer and a dilated convolution layer to extract a task feature vector, and constructing a task execution time prediction model for predicting a maximum execution time threshold through a bidirectional gated recurrent neural network based on the task feature vector, including:

[0022] Inputting the initial feature sequence into the causal convolution layer, the causal convolution layer adopts a one-dimensional convolution structure, the convolution kernel size of the one-dimensional convolution structure is three, the step size is one, two zero-value data are filled at the beginning position of the initial feature sequence, and the convolution kernel of the causal convolution layer is used to perform a convolution operation with the initial feature sequence to obtain a local feature sequence that retains the temporal causal relationship;

[0023] Input the local feature sequence into a multi-layer dilated convolutional layer, wherein the multi-layer dilated convolutional layer adopts a progressive expansion rate, the expansion rate of the first dilated convolutional layer is two, and the expansion rate of each additional dilated convolutional layer is multiplied by two, and the convolution kernel of each dilated convolutional layer performs interval sampling on the local feature sequence and performs convolution operation to obtain feature sequences of different time scales;

[0024] The feature sequences of different time scales are spliced ​​and fused in the feature dimension to obtain a multi-scale fused feature sequence; the multi-scale fused feature sequence is input into a bidirectional gated recurrent neural network, the bidirectional gated recurrent neural network includes a forward gated recurrent unit and a backward gated recurrent unit, the forward gated recurrent unit and the backward gated recurrent unit respectively include an update gate and a reset gate; based on the gating state of the update gate and the reset gate, the input feature of the multi-scale fused feature sequence at the current moment is combined with the hidden state at the previous moment to obtain a candidate hidden state, and the candidate hidden state and the hidden state at the previous moment are weightedly fused according to the gating coefficient of the update gate to obtain an output hidden state at the current moment;

[0025] The output hidden states of the forward gated recurrent unit and the backward gated recurrent unit are concatenated in the feature dimension, and the concatenated features are mapped to a prediction threshold of the task execution time through a fully connected layer.

[0026] In an optional embodiment,

[0027] Constructing a task operation state space based on the real-time operation data, modeling the task state using a hierarchical Markov decision model, converting the real-time operation data into a task state sequence according to a time series, and calculating the state transition probability using a Bayesian inference method include:

[0028] Based on the real-time operation data, determine the task execution phase data, system resource usage data and system call sequence data; map the task execution phase data to the task layer state space to form a task layer state node, map the system resource usage data to the resource layer state space to form a resource layer state node, map the system call sequence data to the operation layer state space to form an operation layer state node; construct a task operation state space based on the task layer state node, the resource layer state node and the operation layer state node;

[0029] In accordance with the time series order, the task layer state nodes are connected to form a task layer state sequence, the resource layer state nodes are connected to form a resource layer state sequence, and the operation layer state nodes are connected to form an operation layer state sequence; the Bayesian inference algorithm is applied to calculate the task layer state transition probability matrix, the resource layer state transition probability matrix and the operation layer state transition probability matrix;

[0030] The task layer state transition probability matrix is ​​set as a constraint condition, and the resource layer state transition probability matrix is ​​modified; the modified resource layer state transition probability matrix is ​​set as a constraint condition, and the operation layer state transition probability matrix is ​​modified;

[0031] Performing a state clustering operation on the operation layer state sequence to obtain an operation layer clustering result, and recalculating the resource layer state transition probability matrix based on the operation layer clustering result as observation data; performing a state clustering operation on the resource layer state sequence to obtain a resource layer clustering result, and recalculating the task layer state transition probability matrix based on the resource layer clustering result as observation data;

[0032] A hierarchical Markov decision model is constructed based on the recalculated task layer state transition probability matrix, the resource layer state transition probability matrix and the operation layer state transition probability matrix.

[0033] In an optional embodiment,

[0034] Based on the state transition probability, the task state sequence is segmented using a spectral clustering algorithm, and the abnormal state segments are identified in combination with a dynamic time warping algorithm, including:

[0035] Obtain a state transition probability distribution vector between any two state nodes in the task state sequence, calculate the similarity between corresponding state nodes based on the state transition probability distribution vector, and construct a state node similarity matrix; construct a diagonal matrix based on the sum of row elements of the state node similarity matrix; calculate a normalized Laplace matrix using the state node similarity matrix and the diagonal matrix;

[0036] Calculate the eigenvalues ​​and eigenvectors of the normalized Laplace matrix, select the eigenvectors corresponding to the minimum non-zero eigenvalues ​​of the corresponding number according to a preset number, construct a feature matrix, perform clustering operation on the row vectors of the feature matrix to obtain cluster labels of state nodes; group state nodes with the same cluster label and continuous in time into state segments;

[0037] Extracting multidimensional state transition features from the state fragments, performing wavelet transform on the multidimensional state transition features to obtain time-frequency feature coefficients, calculating local energy distribution of the time-frequency feature coefficients, determining core feature coefficients based on the local energy distribution, and constructing a state transition feature sequence;

[0038] Construct an empty state transfer feature cumulative distance matrix, calculate the time series autocorrelation coefficient of the state transfer feature sequence, determine the initial window size based on the time series autocorrelation coefficient, dynamically adjust the initial window size according to the local fluctuation degree of the state transfer feature sequence, obtain an adaptive window constraint, calculate the direction angle difference of adjacent state transfer feature vectors in the state transfer feature sequence, map the direction angle difference to a state transfer direction penalty factor, and calculate the weighted distance between the current state transfer feature vector and three adjacent cumulative distance points within the range of the adaptive window constraint;

[0039] Adding the Euclidean distance between the state transfer feature vectors to the minimum value of the weighted distance, and multiplying it by the state transfer direction penalty factor to obtain the corresponding element value of the state transfer feature cumulative distance matrix;

[0040] Backtracking in the state transition feature cumulative distance matrix to obtain a minimum cumulative distance path, and determining an optimal alignment method and distance value between the state segment and a preset reference sequence template;

[0041] When the distance value exceeds a preset distance threshold, the state segment is determined to be an abnormal state segment.

[0042] In an optional embodiment,

[0043] For the abnormal task, a task resource competition matrix is ​​constructed based on the real-time operation data, a weighted directed graph is constructed based on the task resource competition as a task dependency graph, an improved minimum spanning tree algorithm is used to calculate the resource critical path, and a task sorting algorithm based on the resource critical path is used to determine the optimal computing node for re-executing the abnormal task, including:

[0044] Based on the real-time operation data, for any two computing nodes, the ratio of the minimum value to the maximum value of each type of resource demand is calculated, the sum of the ratios is used as the task resource competition degree between the two computing nodes, and a task resource competition degree matrix is ​​constructed;

[0045] Based on the task resource competition degree matrix, a weighted directed graph is constructed as a task dependency graph, the computing nodes are used as vertices in the task dependency graph, and the weighted sum of the task resource competition degree and the communication overhead between nodes is used as the initial weight of the edge in the task dependency graph;

[0046] Calculate the ratio of resource usage to resource capacity of each computing node in the task dependency graph to obtain a resource load factor, multiply the initial weight of the edge in the task dependency graph by the average value of the resource load factors of two connected computing nodes to obtain an improved edge weight; construct a minimum spanning tree based on the improved edge weight to obtain a resource critical path;

[0047] The inverse of the resource load factor, the inverse of the improved edge weight sum of the connected edges, and the node processing capacity index are weighted to obtain a comprehensive score of each computing node on the resource critical path; the computing nodes on the resource critical path are sorted from high to low according to the comprehensive score, and the computing node with the highest comprehensive score is selected as a candidate execution node;

[0048] Calculate the difference in task resource contention before and after the abnormal task is migrated to the candidate execution node to obtain a task resource contention change value; determine whether the task resource contention change value exceeds a preset change threshold; if so, select the next computing node in the sorting as the candidate execution node and repeat the calculation until it is less than the preset change threshold to determine the optimal computing node.

[0049] In an optional embodiment,

[0050] Recording the re-execution trajectory information of the abnormal task, calculating the Pearson correlation coefficient between the predicted value and the actual value of the task execution time prediction model to determine the prediction error, and using the time series decomposition algorithm to feed back the prediction error to different network layers of the task execution time prediction model for optimization; when the number of re-executions of the abnormal task reaches a preset maximum number of retries, terminating the abnormal task and returning a failure analysis report includes:

[0051] Recording the re-execution trajectory information of the abnormal task; acquiring the predicted time series and the actual execution time series of the task execution time prediction model based on the re-execution trajectory information;

[0052] Calculate the covariance between the predicted time series and the actual execution time series divided by the product of the corresponding standard deviations to obtain the Pearson correlation coefficient as the prediction error; decompose the prediction error into a trend error term, a periodic error term, and a random error term through a time series decomposition algorithm;

[0053] Adjusting the prediction bias parameter in the task execution time prediction model based on the trend error term, adjusting the timing feature parameter in the task execution time prediction model based on the periodic error term, adjusting the weight parameter in the task execution time prediction model based on the random error term, and recording the adjustment information of the task execution time prediction model;

[0054] Counting the number of re-executions of the abnormal task, and when the number of re-executions reaches a preset maximum number of retries, terminating the execution of the abnormal task;

[0055] generating execution statistics based on the re-execution trajectory information, generating error analysis data based on the prediction error, and generating model optimization data based on the adjustment information of the task execution time prediction model;

[0056] The execution statistics data, the error analysis data and the model optimization data are integrated to generate a failure analysis report and output it.

[0057] According to a second aspect of the embodiments of the present invention,

[0058] A multi-task monitoring and scheduling system is provided, comprising:

[0059] The first unit is used to receive a task submission request carrying task identification information, divide the task group according to the task identification information, collect the historical execution time series of the same task group, execute an adaptive filtering algorithm based on kernel density estimation, perform noise reduction processing on the historical execution time series to obtain an initial feature sequence, input the initial feature sequence into a time series feature extraction network including a causal convolution layer and a hole convolution layer, extract the task feature vector, and construct a task execution time prediction model for predicting a maximum execution time threshold based on the task feature vector through a bidirectional gated recurrent neural network;

[0060] The second unit is used to monitor the execution process of the task, collect the task execution time and real-time operation data, and when the task execution time exceeds the maximum execution time threshold, construct the task operation state space based on the real-time operation data, use the hierarchical Markov decision model to model the task state, convert the real-time operation data into a task state sequence according to the time series, use the Bayesian inference method to calculate the state transition probability, and based on the state transition probability, use the spectral clustering algorithm to segment the task state sequence, and combine the dynamic time warping algorithm to identify abnormal state segments, and when the abnormal state segment is identified, the task is determined as an abnormal task;

[0061] The third unit is used to construct a task resource competition matrix for the abnormal task based on the real-time operation data, construct a weighted directed graph as a task dependency graph based on the task resource competition, calculate the resource critical path using an improved minimum spanning tree algorithm, and determine the optimal computing node for re-executing the abnormal task using a task sorting algorithm based on the resource critical path; record the re-execution trajectory information of the abnormal task, calculate the Pearson correlation coefficient between the predicted value and the actual value of the task execution time prediction model to determine the prediction error, and use a time series decomposition algorithm to feed back the prediction error to different network layers of the task execution time prediction model for optimization; when the number of re-executions of the abnormal task reaches a preset maximum number of retries, terminate the abnormal task and return a failure analysis report.

[0062] According to a third aspect of the embodiments of the present invention,

[0063] An electronic device is provided, comprising:

[0064] processor;

[0065] a memory for storing processor-executable instructions;

[0066] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0067] A fourth aspect of the embodiments of the present invention is:

[0068] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0069] In an embodiment of the present invention, a historical execution time series is subjected to noise reduction processing by an adaptive filtering algorithm based on kernel density estimation, and a task feature vector is extracted by combining a time series feature extraction network of a causal convolution layer and a hole convolution layer, and then a prediction model is constructed by using a bidirectional gated recurrent neural network, which can accurately predict the maximum execution time threshold of the task, thereby improving the accuracy and reliability of task execution time prediction; a hierarchical Markov decision model is used to model the task state, and a Bayesian inference method is combined to calculate the state transition probability, and a spectral clustering algorithm and a dynamic time warping algorithm are used to identify abnormal state fragments, thereby achieving accurate identification and timely discovery of abnormal task states, and improving the real-time and accuracy of task monitoring; a weighted directed graph is constructed as a task dependency graph based on task resource competition, and a resource critical path is calculated by an improved minimum spanning tree algorithm, and an optimal computing node is determined by using a task sorting algorithm based on a resource critical path, and a time series decomposition algorithm is used to feed back the prediction error to the prediction model for optimization, thereby improving the rescheduling efficiency and execution success rate of abnormal tasks, and enhancing the fault tolerance and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 A flowchart of a multi-task monitoring and scheduling method according to an embodiment of the present invention;

[0071] Figure 2 It is a structural diagram of a multi-task monitoring and scheduling system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0072] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0073] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0074] Figure 1 FIG. 1 is a flow chart of a multi-task monitoring and scheduling method according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0075] S101. Receive a task submission request carrying task identification information, divide the task group according to the task identification information, collect the historical execution time series of the same task group, execute an adaptive filtering algorithm based on kernel density estimation, perform noise reduction processing on the historical execution time series to obtain an initial feature sequence, input the initial feature sequence into a temporal feature extraction network including a causal convolution layer and a hole convolution layer, extract the task feature vector, and construct a task execution time prediction model for predicting the maximum execution time threshold based on the task feature vector through a bidirectional gated recurrent neural network;

[0076] In the task execution time threshold prediction phase, tasks are first grouped by task identification information, which includes characteristic attributes such as task type and resource requirements. Historical execution data is collected for each task group to construct an execution time series.

[0077] The historical execution time series is then preprocessed, and an adaptive filter is constructed using the kernel density estimation method. The filter dynamically adjusts the kernel function bandwidth parameter to filter out noise and outliers in the sequence, and obtains a smoother initial feature sequence. These processed sequences will serve as the input of the subsequent deep learning model.

[0078] The preprocessed feature sequence is input into the composite convolutional network structure. The network consists of causal convolutional layers and dilated convolutional layers, which are responsible for extracting short-term and long-term temporal dependency features respectively. The extracted feature vector is then input into the bidirectional gated recurrent neural network, and the maximum execution time threshold of the task is finally predicted through forward and reverse feature learning. This multi-level feature extraction and bidirectional learning architecture not only ensures the temporal causality of the prediction, but also improves the accuracy of threshold prediction.

[0079] In this embodiment, the adaptive filtering algorithm of kernel density estimation effectively reduces noise, provides high-quality historical execution time data, and improves the accuracy of the time prediction model; the causal convolution layer and the hole convolution layer are combined to capture short-term dependencies and retain long-term trends, thereby improving the expression effect of task features; the bidirectional gated recurrent neural network models the before and after dependencies, adapts to the nonlinear changes of task execution time, and enhances the robustness of the model; the maximum execution time threshold is predicted to provide a basis for task allocation and resource optimization, thereby reducing the risk of timeout or resource waste.

[0080] S102. Monitor the execution process of the task, collect the task execution time and real-time operation data, and when the task execution time exceeds the maximum execution time threshold, construct the task operation state space based on the real-time operation data, use the hierarchical Markov decision model to model the task state, convert the real-time operation data into a task state sequence according to the time series, use the Bayesian inference method to calculate the state transition probability, and based on the state transition probability, use the spectral clustering algorithm to segment the task state sequence, and combine the dynamic time warping algorithm to identify abnormal state segments, and when the abnormal state segment is identified, determine the task as an abnormal task;

[0081] In the task anomaly detection phase, the system continuously monitors the execution status of the task and collects real-time operating indicator data including CPU usage, memory usage, I / O status, etc. When it is detected that the task execution time exceeds the predicted maximum time threshold, an in-depth anomaly analysis process is triggered.

[0082] Based on the collected real-time operation data, a multi-dimensional task state space is constructed, and the hierarchical Markov decision model is used for state modeling. This model regards the task execution process as a state transition sequence, and calculates the transition probability between different states through the Bayesian inference method, thereby describing the dynamic characteristics of task operation.

[0083] After obtaining the state transition sequence, the spectral clustering algorithm is used to segment the sequence. By analyzing the eigenvalues ​​and eigenvectors, the state sequence is divided into multiple subsequences. Then, the dynamic time warping algorithm is used to compare these subsequences with the normal execution mode to identify abnormal state fragments with significant deviations. When it is confirmed that there is an abnormal state fragment, the system marks the task as an abnormal task and triggers the subsequent processing mechanism.

[0084] In this embodiment, based on real-time operation data and dynamic time warping algorithm, it is possible to quickly identify abnormal state segments of tasks and achieve efficient abnormal task judgment; the hierarchical Markov decision model and Bayesian inference method are used to calculate the state transition probability and accurately describe the law of change of the operation state of the task; the spectral clustering algorithm segments the task state sequence and clearly divides the different operation stages, which is helpful for in-depth analysis of the behavioral characteristics of the task state; timely intervention is carried out when the maximum execution time threshold is exceeded, and the abnormal task is accurately located, effectively reducing the impact scope and subsequent costs of operation failures.

[0085] S103. For the abnormal task, a task resource competition matrix is ​​constructed based on the real-time operation data, a weighted directed graph is constructed based on the task resource competition as a task dependency graph, the resource critical path is calculated using an improved minimum spanning tree algorithm, and the optimal computing node for re-executing the abnormal task is determined by a task sorting algorithm based on the resource critical path; the re-execution trajectory information of the abnormal task is recorded, the Pearson correlation coefficient between the predicted value and the actual value of the task execution time prediction model is calculated to determine the prediction error, and the prediction error is fed back to the different network layers of the task execution time prediction model by a time series decomposition algorithm for optimization; when the number of re-executions of the abnormal task reaches the preset maximum number of retries, the abnormal task is terminated and a failure analysis report is returned.

[0086] In the rescheduling phase of abnormal tasks, the system first analyzes the resource competition relationship between tasks based on real-time operation data and constructs a resource competition matrix. This matrix reflects the degree of contention of different tasks for computing resources, storage resources, etc., and constructs a weighted directed graph based on this as a representation of task dependencies.

[0087] Through the improved minimum spanning tree algorithm, the system calculates the resource critical path on the task dependency graph, which represents the task execution chain where resource competition is most critical. Based on this critical path and the load status of each computing node, the system determines the optimal target node for re-execution of abnormal tasks to minimize the impact of resource competition.

[0088] During the task re-execution process, the system records the execution trajectory information and calculates the error of the prediction model. The Pearson correlation coefficient is used to evaluate the degree of deviation between the predicted value and the actual value, and the time series decomposition algorithm is used to distribute the error to different network layers of the prediction model for targeted optimization. If the number of retries reaches the preset upper limit and still fails, the task is terminated and an analysis report containing the failure cause, resource competition and execution trajectory is generated.

[0089] In this embodiment, a task dependency graph is constructed based on the task resource competition degree, and an improved minimum spanning tree algorithm is used to optimize resource allocation to ensure that the optimal computing node is selected for the re-execution of abnormal tasks, thereby reducing resource waste and execution delays; the prediction error is fed back to the different network layers of the task execution time prediction model through a timing decomposition algorithm for optimization, and the time prediction accuracy of the model for abnormal tasks is gradually improved; the maximum number of retries is set and the failed task is terminated, and a failure analysis report is returned at the same time, which effectively avoids repeated resource occupation and improves the overall stability of the system; the re-execution trajectory is recorded and the critical path of resources is analyzed, which helps to identify the bottleneck of abnormal tasks and provide data support for subsequent system optimization and fault diagnosis.

[0090] In an optional implementation, executing an adaptive filtering algorithm based on kernel density estimation to perform noise reduction processing on the historical execution time series to obtain an initial feature sequence includes:

[0091] Divide the historical execution time series into multiple discrete time points, each of the discrete time points corresponds to a time series data value in the historical execution time series, and the time series data value and the discrete time point together constitute a time series data point;

[0092] Performing kernel density estimation on the historical execution time series, calculating the probability density value of each time series data point at the discrete time point by using a Gaussian kernel function, determining the bandwidth parameter of the Gaussian kernel function based on the standard deviation and interquartile range of the historical execution time series, and obtaining the probability density distribution of the historical execution time series;

[0093] Based on the probability density distribution, the probability density value is used as a density estimation weight to calculate the local mean of the historical execution time series at each discrete time point, and the local variance value of the historical execution time series is calculated based on the local mean and the probability density value;

[0094] Normalizing the local variance value based on the probability density value, calculating the ratio of the normalized local variance value to the global standard deviation of the historical execution time series, determining an adjustment coefficient according to the ratio, and dynamically determining a filter window size for each discrete time point based on a preset basic window size and in combination with the adjustment coefficient;

[0095] For each of the discrete time points, within the corresponding filtering window, the time series data points are selected to construct a cubic polynomial fitting model, and a regression weight function is constructed based on the probability density value, wherein the regression weight function is the product of a hyperbolic function form and the probability density value, and the weight value of the regression weight function decreases cubically as the distance between the time series data point and the corresponding discrete time point increases;

[0096] Using the regression weight function, weighted processing is performed on the time series data points within the filter window, coefficients of the cubic polynomial fitting model are solved by weighted least squares method, and a correction value at each discrete time point is obtained according to the coefficients;

[0097] The local standard deviation of the time series data points in the filtering window is calculated in combination with the probability density value, and the time series data points whose time series data values ​​deviate from the local mean by more than twice the local standard deviation are determined as abnormal data points. The time series data values ​​of the abnormal data points are replaced with the corresponding correction values ​​to obtain the initial feature sequence.

[0098] In a specific implementation, the kernel density estimation-based adaptive filtering algorithm first needs to preprocess the input historical execution time series. Assume that the input historical execution time series contains 1000 data points with a time span of 10 minutes. These data points are evenly distributed on the time axis in chronological order, and each data point contains a timestamp and a corresponding execution time value.

[0099] When performing kernel density estimation, the Gaussian kernel function is selected as the basic kernel function. For each time series data point, its probability density contribution at each discrete time point is calculated. The selection of the bandwidth parameter is crucial and is determined by calculating the standard deviation and interquartile range of the historical execution time series. Specifically, when the sequence fluctuates greatly, the smoothing effect is improved by increasing the bandwidth parameter; when the sequence fluctuates less, the bandwidth parameter is reduced to retain more detailed features.

[0100] After obtaining the probability density distribution, the probability density value is used as the weight to calculate the local statistical features. For each discrete time point, within its neighborhood, the weighted average value is calculated based on the probability density value as the local mean, and the weighted variance is calculated as the local variance. These local statistical features can reflect the distribution characteristics of the data at different time points.

[0101] In order to achieve adaptive filtering, the filter window size needs to be determined dynamically. First, the local variance is normalized by dividing it by the probability density value, and then compared with the global standard deviation to obtain the adjustment coefficient. The basic window size is set to 60 data points, which is approximately 36 seconds in time span. When the local fluctuation is large, the window size is increased by the adjustment coefficient to enhance the smoothing effect; when the local fluctuation is small, the window size is reduced to retain more detailed features.

[0102] A cubic polynomial fitting model is constructed within a certain filtering window. The regression weight function is in the form of the product of a hyperbolic function and a probability density value, ensuring that the farther the data point is from the target time point, the smaller the weight. For example, for a data point with a distance of d, its weight decreases as the cube of d increases, which can better protect local features.

[0103] The polynomial coefficients are solved by weighted least squares method to obtain the correction value for each time point. At the same time, the local standard deviation of the data points in the filter window is calculated, and the data points that deviate from the local mean by more than twice the local standard deviation are marked as abnormal points. For example, if the execution time of a time point is 100ms, the local mean is 50ms, and the local standard deviation is 20ms, then this point will be judged as an abnormal point, and its value will be replaced by the correction value obtained by polynomial fitting.

[0104] In this embodiment, kernel density estimation is used to adaptively extract data distribution features, thus avoiding the over-sensitivity of traditional fixed parameter methods to abnormal fluctuations and improving the robustness and accuracy of time series processing. An adaptive filtering strategy with a dynamic window size is used, which can automatically adjust processing parameters according to local data characteristics, effectively suppress noise interference while retaining important patterns, and improve the adaptability and processing effect of the algorithm. The anomaly detection method based on weighted polynomial fitting makes full use of the local correlation of the data, can accurately identify and correct abnormal fluctuations, and at the same time maintains the overall trend characteristics of the time series, ensuring the reliability and practicality of the processing results.

[0105] In an optional implementation, the initial feature sequence is input into a temporal feature extraction network including a causal convolution layer and a dilated convolution layer to extract a task feature vector, and based on the task feature vector, a task execution time prediction model for predicting a maximum execution time threshold is constructed through a bidirectional gated recurrent neural network, including:

[0106] Inputting the initial feature sequence into the causal convolution layer, the causal convolution layer adopts a one-dimensional convolution structure, the convolution kernel size of the one-dimensional convolution structure is three, the step size is one, two zero-value data are filled at the beginning position of the initial feature sequence, and the convolution kernel of the causal convolution layer is used to perform a convolution operation with the initial feature sequence to obtain a local feature sequence that retains the temporal causal relationship;

[0107] Input the local feature sequence into a multi-layer dilated convolutional layer, wherein the multi-layer dilated convolutional layer adopts a progressive expansion rate, the expansion rate of the first dilated convolutional layer is two, and the expansion rate of each additional dilated convolutional layer is multiplied by two, and the convolution kernel of each dilated convolutional layer performs interval sampling on the local feature sequence and performs convolution operation to obtain feature sequences of different time scales;

[0108] The feature sequences of different time scales are spliced ​​and fused in the feature dimension to obtain a multi-scale fused feature sequence; the multi-scale fused feature sequence is input into a bidirectional gated recurrent neural network, the bidirectional gated recurrent neural network includes a forward gated recurrent unit and a backward gated recurrent unit, the forward gated recurrent unit and the backward gated recurrent unit respectively include an update gate and a reset gate; based on the gating state of the update gate and the reset gate, the input feature of the multi-scale fused feature sequence at the current moment is combined with the hidden state at the previous moment to obtain a candidate hidden state, and the candidate hidden state and the hidden state at the previous moment are weightedly fused according to the gating coefficient of the update gate to obtain an output hidden state at the current moment;

[0109] The output hidden states of the forward gated recurrent unit and the backward gated recurrent unit are concatenated in the feature dimension, and the concatenated features are mapped to a prediction threshold of the task execution time through a fully connected layer.

[0110] In a specific implementation, first, the initial feature sequence is preprocessed and feature extracted. In practical applications, the initial feature sequence contains multiple attribute information of the task, such as task priority, estimated resource demand, historical execution time, etc. These features are organized in the form of a time series, and each time step contains a multi-dimensional feature vector.

[0111] In the causal convolution layer processing stage, a one-dimensional convolution structure is used to process the initial feature sequence. In the specific implementation, the convolution kernel size is set to three, that is, the convolution operation is performed on the features of three consecutive time steps each time. In order to keep the output sequence length the same as the input sequence, two zero-valued data are padded at the beginning of the sequence. The convolution step size is set to one to ensure that no time step information is missed. For example, for an initial sequence containing ten time steps, each time step contains an eight-dimensional feature vector, and a local feature sequence of the same length can be obtained through the causal convolution layer.

[0112] Next, the local feature sequence obtained is input into a multi-layer dilated convolutional network. Specifically, four dilated convolutional layers are used, with dilation rates of 2, 4, 8, and 16, respectively. Taking the first layer as an example, a dilation rate of 2 means that the convolution kernel samples once every other time step when extracting features. This progressive dilation design enables the network to capture dependencies at different time scales. The output of each layer of dilated convolution keeps the sequence length unchanged, but obtains feature representations of different receptive fields.

[0113] Then, the output features of the four dilated convolutional layers are concatenated in terms of feature dimensions. Assuming that the feature dimension of each dilated convolutional layer output is 16, a 64-dimensional multi-scale fusion feature sequence is obtained after concatenation. This feature fusion method retains information at different time scales.

[0114] The fused feature sequence is input into the bidirectional gated recurrent neural network for further processing. In the forward processing, the sequence is processed from left to right; in the backward processing, the sequence is processed from right to left. The processing in each direction includes two control units: the update gate and the reset gate. The update gate controls how much historical information needs to be retained at the current moment, and the reset gate controls how much historical information needs to be forgotten.

[0115] Finally, the results of the bidirectional processing are concatenated in the feature dimension and mapped to the final time prediction value through the fully connected layer. Assuming that the output of the bidirectional processing is 32-dimensional features, 64-dimensional features are obtained after concatenation, and a scalar value is finally output through the fully connected layer, which is the predicted task execution time threshold.

[0116] In this embodiment, through the combined application of causal convolution and dilated convolution, the long-term and short-term dependencies in the task sequence can be effectively captured, the accuracy of time prediction can be improved, and the problem of information loss in the traditional method can be avoided; the multi-layer dilated convolution structure with a progressive expansion rate is adopted to significantly expand the receptive field of the model, so that the model can make full use of historical information at different time scales, and enhance the comprehensiveness and robustness of feature extraction; the timing modeling scheme based on the bidirectional gated recurrent neural network, through the forward and backward bidirectional information flow, combined with the dynamic adjustment mechanism of the update gate and the reset gate, realizes the selective retention and forgetting of historical information, and improves the model's ability to understand timing patterns and generalization performance.

[0117] In an optional implementation, constructing a task operation state space based on the real-time operation data, modeling the task state using a hierarchical Markov decision model, converting the real-time operation data into a task state sequence according to a time series, and calculating the state transition probability using a Bayesian inference method include:

[0118] Based on the real-time operation data, determine the task execution phase data, system resource usage data and system call sequence data; map the task execution phase data to the task layer state space to form a task layer state node, map the system resource usage data to the resource layer state space to form a resource layer state node, map the system call sequence data to the operation layer state space to form an operation layer state node; construct a task operation state space based on the task layer state node, the resource layer state node and the operation layer state node;

[0119] In accordance with the time series order, the task layer state nodes are connected to form a task layer state sequence, the resource layer state nodes are connected to form a resource layer state sequence, and the operation layer state nodes are connected to form an operation layer state sequence; the Bayesian inference algorithm is applied to calculate the task layer state transition probability matrix, the resource layer state transition probability matrix and the operation layer state transition probability matrix;

[0120] The task layer state transition probability matrix is ​​set as a constraint condition, and the resource layer state transition probability matrix is ​​modified; the modified resource layer state transition probability matrix is ​​set as a constraint condition, and the operation layer state transition probability matrix is ​​modified;

[0121] Performing a state clustering operation on the operation layer state sequence to obtain an operation layer clustering result, and recalculating the resource layer state transition probability matrix based on the operation layer clustering result as observation data; performing a state clustering operation on the resource layer state sequence to obtain a resource layer clustering result, and recalculating the task layer state transition probability matrix based on the resource layer clustering result as observation data;

[0122] A hierarchical Markov decision model is constructed based on the recalculated task layer state transition probability matrix, the resource layer state transition probability matrix and the operation layer state transition probability matrix.

[0123] In a specific implementation, firstly, the real-time operation data of the task is obtained during the execution of the task, which includes the task execution stage information, system resource usage and system call sequence. For the task execution stage information, each stage identifier of the task from start, run to end is recorded; the system resource usage includes indicators such as CPU usage, memory usage, disk read and write rate; the system call sequence records all system calls that occur during the task execution.

[0124] The raw data is converted into state space through data preprocessing. At the task level, the task execution stages such as initialization, calculation, data transmission, etc. are mapped to different state nodes; at the resource level, the resource utilization rate is divided into high load, medium load, low load and other state nodes according to the preset threshold; at the operation level, the continuous system call sequence is aggregated into operation state nodes according to functional similarity. For example, multiple file reading and writing related system calls can be aggregated into a file operation state node.

[0125] Next, we construct a time-series state sequence. Using timestamps as indexes, we connect the state nodes at the same level in chronological order. For example, the task layer may form a state sequence of "initialization-data reading-computation-data writing-end"; the resource layer may form a state sequence of "low load-medium load-high load-medium load-low load"; the operation layer may form a state sequence of "file operation-memory operation-computation operation-network operation".

[0126] The Bayesian inference method is used to calculate the state transition probability. The transition probability between states is calculated by counting the number of transitions between adjacent states in the state sequence. For example, at the task level, if it is observed that there are 80 transitions from the "data reading" state to the "computing" state and 20 transitions from "data reading" to "data writing", then the corresponding transition probabilities are 0.8 and 0.2 respectively.

[0127] After obtaining the preliminary transition probability matrix, it is necessary to correct the probabilities between the layers. Using the task layer transition probability as a constraint, adjust the resource layer transition probability to ensure that the transition of resource status is consistent with the transition of task status. Similarly, using the corrected resource layer transition probability as a constraint, adjust the operation layer transition probability.

[0128] In order to improve the accuracy of the model, cluster analysis is performed on the state sequence. At the operation layer, similar operation sequence patterns are clustered to obtain typical operation patterns; these clustering results are used as new observation data to recalculate the state transition probability of the resource layer. Clustering is also performed at the resource layer, and the clustering results are used to update the transition probability of the task layer.

[0129] Finally, the state spaces and transition probability matrices of the three levels are integrated to construct a complete hierarchical Markov decision model, which can describe the evolution law and mutual influence relationship of the states at each level during task execution.

[0130] In this embodiment, a multi-dimensional characterization of the task running status is achieved through hierarchical modeling, which improves the completeness and accuracy of the state expression; a method combining Bayesian inference and state clustering is adopted to improve the calculation accuracy of the state transition probability, so that the model can better reflect the dynamic characteristics of the actual system; a probability correction mechanism based on hierarchical constraints ensures the consistency of state transfers at different levels, enhancing the interpretability and reliability of the model.

[0131] In an optional implementation, based on the state transition probability, segmenting the task state sequence using a spectral clustering algorithm, and identifying abnormal state segments in combination with a dynamic time warping algorithm includes:

[0132] Obtain a state transition probability distribution vector between any two state nodes in the task state sequence, calculate the similarity between corresponding state nodes based on the state transition probability distribution vector, and construct a state node similarity matrix; construct a diagonal matrix based on the sum of row elements of the state node similarity matrix; calculate a normalized Laplace matrix using the state node similarity matrix and the diagonal matrix;

[0133] Calculate the eigenvalues ​​and eigenvectors of the normalized Laplace matrix, select the eigenvectors corresponding to the minimum non-zero eigenvalues ​​of the corresponding number according to a preset number, construct a feature matrix, perform clustering operation on the row vectors of the feature matrix to obtain cluster labels of state nodes; group state nodes with the same cluster label and continuous in time into state segments;

[0134] Extracting multidimensional state transition features from the state fragments, performing wavelet transform on the multidimensional state transition features to obtain time-frequency feature coefficients, calculating local energy distribution of the time-frequency feature coefficients, determining core feature coefficients based on the local energy distribution, and constructing a state transition feature sequence;

[0135] Construct an empty state transfer feature cumulative distance matrix, calculate the time series autocorrelation coefficient of the state transfer feature sequence, determine the initial window size based on the time series autocorrelation coefficient, dynamically adjust the initial window size according to the local fluctuation degree of the state transfer feature sequence, obtain an adaptive window constraint, calculate the direction angle difference of adjacent state transfer feature vectors in the state transfer feature sequence, map the direction angle difference to a state transfer direction penalty factor, and calculate the weighted distance between the current state transfer feature vector and three adjacent cumulative distance points within the range of the adaptive window constraint;

[0136] Adding the Euclidean distance between the state transfer feature vectors to the minimum value of the weighted distance, and multiplying it by the state transfer direction penalty factor to obtain the corresponding element value of the state transfer feature cumulative distance matrix;

[0137] Backtracking in the state transition feature cumulative distance matrix to obtain a minimum cumulative distance path, and determining an optimal alignment method and distance value between the state segment and a preset reference sequence template;

[0138] When the distance value exceeds a preset distance threshold, the state segment is determined to be an abnormal state segment.

[0139] In a specific implementation, the abnormal state identification method based on state transition probability first needs to process the input task state sequence. For any two state nodes, by calculating the frequency of state transition between them, a state transition probability distribution vector can be obtained. For example, for adjacent states A and B in the state sequence, the probability of transitioning from A to all other possible states is counted to form a probability distribution vector.

[0140] After obtaining the state transition probability distribution vector, the cosine similarity method is used to calculate the similarity between any two state nodes. Specifically, the probability distribution vectors of the two state nodes are dot-producted and divided by the product of the module lengths of the two vectors to obtain a similarity value between zero and one. A larger similarity value indicates that the transition characteristics of the two state nodes are more similar. The similarity values ​​between all state nodes are combined into a square matrix, which is the state node similarity matrix.

[0141] Next, we sum each row of the state node similarity matrix and use these sums as diagonal elements to construct a diagonal matrix. Combining the state node similarity matrix and the diagonal matrix, we can get a normalized Laplace matrix. The eigenvalues ​​and eigenvectors of this matrix contain the clustering information of the state sequence.

[0142] Select the eigenvectors corresponding to the smallest non-zero eigenvalues ​​from the normalized Laplace matrix to form a feature matrix. Use the K-means clustering algorithm to cluster the row vectors of the feature matrix to obtain the cluster label of each state node. Combine state nodes that are continuous in time and have the same cluster label to form a state fragment.

[0143] For the obtained state fragments, extract its state transition features, including features of multiple dimensions such as state duration and transition frequency. Perform wavelet transform on these features to obtain coefficients reflecting different time scales and frequency characteristics. By calculating the energy distribution of these coefficients on the time-frequency plane, the most representative core feature coefficients can be determined to construct the state transition feature sequence.

[0144] In the dynamic time warping algorithm, an empty cumulative distance matrix is ​​first constructed. The initial matching window size is determined by calculating the temporal autocorrelation coefficient of the state transition feature sequence. The window size is dynamically adjusted according to the local fluctuation of the sequence to achieve adaptive window constraints. At the same time, the direction changes between adjacent feature vectors are calculated, and the angle differences are converted into penalty factors to constrain the direction of the matching path.

[0145] Within the adaptive window, the weighted distance between the current feature vector and the adjacent cumulative distance point is calculated. The Euclidean distance between feature vectors is added to the minimum value of the weighted distance, and multiplied by the directional penalty factor to obtain the element value of the cumulative distance matrix. By backtracking the minimum cumulative distance path, the optimal alignment and distance value between the state fragment and the preset reference template can be obtained. When the distance value exceeds the preset threshold, the state fragment can be determined to be an abnormal state.

[0146] In this embodiment, by introducing the state transition probability and spectral clustering algorithm, the structural features in the state sequence can be effectively captured, more accurate state segmentation can be achieved, and the accuracy of anomaly detection is improved; the dynamic time warping algorithm with adaptive window constraints and directional penalty mechanism can better handle the time scale changes and noise interference in the state sequence, thereby improving the robustness of the algorithm; combined with wavelet transform and energy distribution analysis, multi-scale state transition features can be extracted, and the most representative core features can be screened out, thereby reducing the computational complexity and improving the efficiency of the algorithm.

[0147] In an optional implementation, for the abnormal task, a task resource competition matrix is ​​constructed based on the real-time operation data, a weighted directed graph is constructed based on the task resource competition as a task dependency graph, a resource critical path is calculated using an improved minimum spanning tree algorithm, and a task sorting algorithm based on the resource critical path is used to determine the optimal computing node for re-executing the abnormal task, including:

[0148] Based on the real-time operation data, for any two computing nodes, the ratio of the minimum value to the maximum value of each type of resource demand is calculated, the sum of the ratios is used as the task resource competition degree between the two computing nodes, and a task resource competition degree matrix is ​​constructed;

[0149] Based on the task resource competition degree matrix, a weighted directed graph is constructed as a task dependency graph, the computing nodes are used as vertices in the task dependency graph, and the weighted sum of the task resource competition degree and the communication overhead between nodes is used as the initial weight of the edge in the task dependency graph;

[0150] Calculate the ratio of resource usage to resource capacity of each computing node in the task dependency graph to obtain a resource load factor, multiply the initial weight of the edge in the task dependency graph by the average value of the resource load factors of two connected computing nodes to obtain an improved edge weight; construct a minimum spanning tree based on the improved edge weight to obtain a resource critical path;

[0151] The inverse of the resource load factor, the inverse of the improved edge weight sum of the connected edges, and the node processing capacity index are weighted to obtain a comprehensive score of each computing node on the resource critical path; the computing nodes on the resource critical path are sorted from high to low according to the comprehensive score, and the computing node with the highest comprehensive score is selected as a candidate execution node;

[0152] Calculate the difference in task resource contention before and after the abnormal task is migrated to the candidate execution node to obtain a task resource contention change value; determine whether the task resource contention change value exceeds a preset change threshold; if so, select the next computing node in the sorting as the candidate execution node and repeat the calculation until it is less than the preset change threshold to determine the optimal computing node.

[0153] In a specific implementation, for rescheduling the abnormal task, it is first necessary to obtain the real-time operation data of each computing node in the system, including resource usage such as CPU usage, memory usage, network bandwidth, etc. Based on these real-time data, the resource competition relationship between any two computing nodes is calculated.

[0154] For any two computing nodes, their demands for various resources such as CPU, memory, and network bandwidth are counted respectively. By calculating the ratio of the minimum and maximum values ​​of these resource demands and adding up the ratios of all resource types, the task resource competition between the two nodes is obtained. For example, if the CPU usage ratio of node A and node B is 0.8, the memory usage ratio is 0.7, and the network bandwidth ratio is 0.9, then the task resource competition between them is 2.4. In this way, the competition between all node pairs is calculated and the task resource competition matrix is ​​constructed.

[0155] Next, we construct a weighted directed graph with computing nodes as vertices and the weighted sum of the task resource contention and communication overhead between nodes as the weight of the edge. The communication overhead can be measured by the network delay and data transmission volume between nodes. For example, if the resource contention of nodes A and B is 2.4, the communication overhead is 0.6, and the weight coefficients are 0.7 and 0.3 respectively, then the initial weight of the edge is 1.98.

[0156] Then calculate the resource load factor of each computing node. Taking the CPU as an example, if the CPU utilization rate of a node is 75% and the CPU capacity is 100%, then its CPU resource load factor is 0.75. Taking the load factors of various resources into consideration, the overall resource load status of the node can be obtained. Multiply the initial weight of the edge by the average resource load factor of the connected nodes to obtain the improved edge weight.

[0157] Based on the improved edge weights, the minimum spanning tree is constructed using the improved minimum spanning tree algorithm. Based on the traditional minimum spanning tree algorithm, this algorithm takes into account the resource load balance of the nodes to avoid excessive resource concentration. Through this algorithm, the resource critical path can be obtained, that is, the optimal path in terms of resource competition and load balancing.

[0158] Each node on the critical resource path is comprehensively scored. The score considers three aspects: the inverse of the resource load factor (indicating the degree of resource surplus), the inverse of the sum of the improved edge weights of the connected edges (indicating the degree of competition with other nodes), and the processing capacity index of the node (such as CPU main frequency, memory size, etc.). These three indicators are weighted and calculated to obtain the final score.

[0159] Sort the nodes in descending order according to the comprehensive score, and select the node with the highest score as the candidate execution node. Calculate the change value of the task resource competition degree after migrating the abnormal task to this node. If the change value exceeds the preset threshold (such as 20%), consider the next node with a higher score, and repeat this process until a suitable node is found.

[0160] In this embodiment, by constructing a task resource competition degree matrix and a weighted directed graph, the resource competition relationship between computing nodes is accurately reflected, making the task scheduling more targeted and avoiding resource conflicts; an improved minimum spanning tree algorithm and a comprehensive scoring mechanism are adopted, and factors such as resource load, node performance, and communication overhead are comprehensively considered when selecting an execution node, improving the system resource utilization efficiency; a judgment mechanism for the change value of the task resource competition degree is introduced to ensure that task migration will not cause a drastic fluctuation in the system load, ensuring the stability and reliability of the system operation.

[0161] In an alternative embodiment, record the re-execution trajectory information of the abnormal task, calculate the Pearson correlation coefficient between the predicted value and the actual value of the task execution time prediction model to determine the prediction error, and use the time series decomposition algorithm to feedback the prediction error to different network layers of the task execution time prediction model for optimization; when the number of re-executions of the abnormal task reaches the preset maximum retry times, terminate the abnormal task and return a failure analysis report including:

[0162] Record the re-execution trajectory information of the abnormal task; obtain the predicted time series and the actual execution time series of the task execution time prediction model based on the re-execution trajectory information;

[0163] Calculate the product of the covariance between the predicted time series and the actual execution time series divided by the corresponding standard deviation to obtain the Pearson correlation coefficient as the prediction error; decompose the prediction error into a trend error term, a periodic error term, and a random error term through the time series decomposition algorithm;

[0164] Based on the trend error term, adjust the prediction bias parameter in the task execution time prediction model, based on the periodic error term, adjust the time series feature parameter in the task execution time prediction model, based on the random error term, adjust the weight parameter in the task execution time prediction model, and record the adjustment information of the task execution time prediction model;

[0165] Count the number of re-executions of the abnormal task. When the number of re-executions reaches the preset maximum retry times, terminate the execution of the abnormal task;

[0166] Generate execution statistical data based on the re-execution trajectory information, generate error analysis data based on the prediction error, and generate model optimization data based on the adjustment information of the task execution time prediction model;

[0167] The execution statistics data, the error analysis data and the model optimization data are integrated to generate a failure analysis report and output it.

[0168] In a specific implementation, first, when an abnormal task is detected, the system will record the re-execution trajectory information of the abnormal task. This information includes detailed data such as the start time, end time, execution status, resource consumption, etc. of the task. For example, for a data processing task, the following information may be recorded: the task ID is Task001, the start time is 2023-04-01 10:00:00, the end time is 2023-04-01 10:15:30, the execution status is "completed", the CPU usage is 75%, and the memory usage is 2GB.

[0169] Next, the system obtains the predicted time series and actual execution time series of the task execution time prediction model based on the recorded re-execution trajectory information. The predicted time series is the execution time predicted by the model before the task starts, while the actual execution time series is the time it actually takes to complete the task. For example, the predicted time series may be [900 seconds, 920 seconds, 910 seconds], while the actual execution time series may be [930 seconds, 915 seconds, 925 seconds].

[0170] The system then calculates the Pearson correlation coefficient between the predicted time series and the actual execution time series as a measure of the forecast error. The calculation of the Pearson correlation coefficient involves the covariance and standard deviation of the two series. For example, if the calculated Pearson correlation coefficient is 0.95, it means that there is a strong positive correlation between the predicted value and the actual value, but there is still a certain forecast error.

[0171] Next, the system uses a time series decomposition algorithm to decompose the forecast error into a trend error term, a period error term, and a random error term. The time series decomposition algorithm can use classic time series decomposition methods, such as the moving average method or the X-12-ARIMA method. For example, the trend error term may show that the forecast model has a 5% underestimated tendency in the long term, the period error term may show a 10% overestimate during the peak hours of each day, and the random error term may fluctuate within a range of ±3%.

[0172] Based on the decomposed error terms, the system optimizes the task execution time prediction model:

[0173] For trend error terms, the system adjusts the forecast bias parameter in the forecast model. For example, if the model is found to be chronically underestimating execution time, the bias parameter may be adjusted from 1.0 to 1.05 to compensate for this underestimation trend.

[0174] For the periodic error term, the system adjusts the time series feature parameters in the forecast model. For example, if the model is found to overestimate the execution time during the peak hours of the day, the weight coefficient for these hours may be reduced from 1.2 to 1.1.

[0175] For random error terms, the system adjusts the weight parameters in the prediction model. For example, the weight of certain key features may be increased from 0.3 to 0.35 to improve the model's adaptability to random fluctuations.

[0176] The system will record these adjustment information, including parameter values ​​before and after the adjustment, reasons for the adjustment, etc. This information is very important for subsequent model analysis and further optimization.

[0177] At the same time, the system will count the number of times the abnormal task is re-executed. When the number of re-executions reaches the preset maximum number of retries (for example, 5 times), the system will terminate the execution of the abnormal task to prevent excessive consumption of resources.

[0178] Finally, the system generates a failure analysis report based on the collected information:

[0179] Execution statistics: including the total number of task executions, average execution time, success rate, etc. For example, it may be shown that the task was executed 5 times in total, with an average execution time of 920 seconds and a success rate of 80%.

[0180] Error analysis data: including the mean, variance, and distribution of the forecast error. For example, it may show that the mean of the forecast error is 3%, the variance is 0.0001, and it is normally distributed.

[0181] Model optimization data: includes the adjustment of each model parameter and its effect. For example, it may show that after the bias parameter is adjusted from 1.0 to 1.05, the long-term forecast error is reduced by 2%.

[0182] The system integrates this data into a comprehensive failure analysis report, providing important basis for subsequent task optimization and system improvement.

[0183] In this embodiment, by recording the re-execution trajectory information of abnormal tasks and calculating the prediction error based on this information, the method can accurately evaluate the performance of the task execution time prediction model. This evaluation method based on actual execution data can better reflect the performance of the model in practical applications than simply relying on historical data, thereby providing a more reliable basis for model optimization; the prediction error is decomposed into trend error terms, periodic error terms and random error terms by using a time series decomposition algorithm, and the different parameters of the prediction model are adjusted in a targeted manner, making the model optimization more accurate and efficient. This fine-grained optimization method can effectively improve the prediction accuracy of the model while maintaining the stability of the model; by generating a failure analysis report containing execution statistics, error analysis data and model optimization data, the method provides system administrators and developers with a comprehensive task execution and model performance analysis. These detailed analysis results not only help to quickly locate and solve problems, but also provide important references for long-term system optimization and resource planning.

[0184] Figure 2 FIG. 1 is a schematic diagram of the structure of a multi-task monitoring and scheduling system according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0185] The first unit is used to receive a task submission request carrying task identification information, divide the task group according to the task identification information, collect the historical execution time series of the same task group, execute an adaptive filtering algorithm based on kernel density estimation, perform noise reduction processing on the historical execution time series to obtain an initial feature sequence, input the initial feature sequence into a time series feature extraction network including a causal convolution layer and a hole convolution layer, extract the task feature vector, and construct a task execution time prediction model for predicting a maximum execution time threshold based on the task feature vector through a bidirectional gated recurrent neural network;

[0186] The second unit is used to monitor the execution process of the task, collect the task execution time and real-time operation data, and when the task execution time exceeds the maximum execution time threshold, construct the task operation state space based on the real-time operation data, use the hierarchical Markov decision model to model the task state, convert the real-time operation data into a task state sequence according to the time series, use the Bayesian inference method to calculate the state transition probability, and based on the state transition probability, use the spectral clustering algorithm to segment the task state sequence, and combine the dynamic time warping algorithm to identify abnormal state segments, and when the abnormal state segment is identified, the task is determined as an abnormal task;

[0187] The third unit is used to construct a task resource competition matrix for the abnormal task based on the real-time operation data, construct a weighted directed graph as a task dependency graph based on the task resource competition, calculate the resource critical path using an improved minimum spanning tree algorithm, and determine the optimal computing node for re-executing the abnormal task using a task sorting algorithm based on the resource critical path; record the re-execution trajectory information of the abnormal task, calculate the Pearson correlation coefficient between the predicted value and the actual value of the task execution time prediction model to determine the prediction error, and use a time series decomposition algorithm to feed back the prediction error to different network layers of the task execution time prediction model for optimization; when the number of re-executions of the abnormal task reaches a preset maximum number of retries, terminate the abnormal task and return a failure analysis report.

[0188] According to a third aspect of the embodiments of the present invention,

[0189] An electronic device is provided, comprising:

[0190] processor;

[0191] a memory for storing processor-executable instructions;

[0192] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0193] According to a fourth aspect of the embodiments of the present invention,

[0194] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0195] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-task monitoring and scheduling method, characterized in that: include: Receive a task submission request carrying task identification information, divide the task group according to the task identification information, collect the historical execution time series of the same task group, execute an adaptive filtering algorithm based on kernel density estimation, perform noise reduction processing on the historical execution time series to obtain an initial feature sequence, input the initial feature sequence into a temporal feature extraction network including a causal convolution layer and a hole convolution layer, extract the task feature vector, and based on the task feature vector, construct a task execution time prediction model for predicting a maximum execution time threshold through a bidirectional gated recurrent neural network; Monitor the execution process of the task, collect the task execution time and real-time operation data, when the task execution time exceeds the maximum execution time threshold, construct the task operation state space based on the real-time operation data, use the hierarchical Markov decision model to model the task state, convert the real-time operation data into a task state sequence according to the time series, use the Bayesian inference method to calculate the state transition probability, based on the state transition probability, use the spectral clustering algorithm to segment the task state sequence, and combine the dynamic time warping algorithm to identify abnormal state segments, and when the abnormal state segment is identified, determine the task as an abnormal task; For the abnormal task, a task resource competition matrix is ​​constructed based on the real-time operation data, a weighted directed graph is constructed based on the task resource competition as a task dependency graph, the resource critical path is calculated using an improved minimum spanning tree algorithm, and the optimal computing node for re-executing the abnormal task is determined by a task sorting algorithm based on the resource critical path; the re-execution trajectory information of the abnormal task is recorded, the Pearson correlation coefficient of the predicted value and the actual value of the task execution time prediction model is calculated to determine the prediction error, and the prediction error is fed back to the different network layers of the task execution time prediction model by a time series decomposition algorithm for optimization; when the number of re-executions of the abnormal task reaches a preset maximum number of retries, the abnormal task is terminated and a failure analysis report is returned.

2. The method according to claim 1, characterized in that Executing an adaptive filtering algorithm based on kernel density estimation to perform noise reduction processing on the historical execution time series to obtain an initial feature sequence including: Divide the historical execution time series into multiple discrete time points, each of the discrete time points corresponds to a time series data value in the historical execution time series, and the time series data value and the discrete time point together constitute a time series data point; Performing kernel density estimation on the historical execution time series, calculating the probability density value of each time series data point at the discrete time point by using a Gaussian kernel function, determining the bandwidth parameter of the Gaussian kernel function based on the standard deviation and interquartile range of the historical execution time series, and obtaining the probability density distribution of the historical execution time series; Based on the probability density distribution, the probability density value is used as a density estimation weight to calculate the local mean of the historical execution time series at each discrete time point, and the local variance value of the historical execution time series is calculated based on the local mean and the probability density value; Normalizing the local variance value based on the probability density value, calculating the ratio of the normalized local variance value to the global standard deviation of the historical execution time series, determining an adjustment coefficient according to the ratio, and dynamically determining a filter window size for each discrete time point based on a preset basic window size and in combination with the adjustment coefficient; For each of the discrete time points, within the corresponding filtering window, the time series data points are selected to construct a cubic polynomial fitting model, and a regression weight function is constructed based on the probability density value, wherein the regression weight function is the product of a hyperbolic function form and the probability density value, and the weight value of the regression weight function decreases cubically as the distance between the time series data point and the corresponding discrete time point increases; Using the regression weight function, weighted processing is performed on the time series data points within the filter window, coefficients of the cubic polynomial fitting model are solved by weighted least squares method, and a correction value at each discrete time point is obtained according to the coefficients; The local standard deviation of the time series data points in the filtering window is calculated in combination with the probability density value, and the time series data points whose time series data values ​​deviate from the local mean by more than twice the local standard deviation are determined as abnormal data points. The time series data values ​​of the abnormal data points are replaced with the corresponding correction values ​​to obtain the initial feature sequence.

3. The method according to claim 1, characterized in that Inputting the initial feature sequence into a temporal feature extraction network including a causal convolution layer and a dilated convolution layer to extract a task feature vector, and constructing a task execution time prediction model for predicting a maximum execution time threshold through a bidirectional gated recurrent neural network based on the task feature vector, including: Inputting the initial feature sequence into the causal convolution layer, the causal convolution layer adopts a one-dimensional convolution structure, the convolution kernel size of the one-dimensional convolution structure is three, the step size is one, two zero-value data are filled at the beginning position of the initial feature sequence, and the convolution kernel of the causal convolution layer is used to perform a convolution operation with the initial feature sequence to obtain a local feature sequence that retains the temporal causal relationship; Input the local feature sequence into a multi-layer dilated convolutional layer, wherein the multi-layer dilated convolutional layer adopts a progressive expansion rate, the expansion rate of the first dilated convolutional layer is two, and the expansion rate of each additional dilated convolutional layer is multiplied by two, and the convolution kernel of each dilated convolutional layer performs interval sampling on the local feature sequence and performs convolution operation to obtain feature sequences of different time scales; The feature sequences of different time scales are spliced ​​and fused in the feature dimension to obtain a multi-scale fused feature sequence; the multi-scale fused feature sequence is input into a bidirectional gated recurrent neural network, the bidirectional gated recurrent neural network includes a forward gated recurrent unit and a backward gated recurrent unit, the forward gated recurrent unit and the backward gated recurrent unit respectively include an update gate and a reset gate; based on the gating state of the update gate and the reset gate, the input feature of the multi-scale fused feature sequence at the current moment is combined with the hidden state at the previous moment to obtain a candidate hidden state, and the candidate hidden state and the hidden state at the previous moment are weightedly fused according to the gating coefficient of the update gate to obtain an output hidden state at the current moment; The output hidden states of the forward gated recurrent unit and the backward gated recurrent unit are concatenated in the feature dimension, and the concatenated features are mapped to a prediction threshold of the task execution time through a fully connected layer.

4. The method according to claim 1, characterized in that: Constructing a task operation state space based on the real-time operation data, modeling the task state using a hierarchical Markov decision model, converting the real-time operation data into a task state sequence according to a time series, and calculating the state transition probability using a Bayesian inference method include: Based on the real-time operation data, determine the task execution phase data, system resource usage data and system call sequence data; map the task execution phase data to the task layer state space to form a task layer state node, map the system resource usage data to the resource layer state space to form a resource layer state node, map the system call sequence data to the operation layer state space to form an operation layer state node; construct a task operation state space based on the task layer state node, the resource layer state node and the operation layer state node; In accordance with the time series order, the task layer state nodes are connected to form a task layer state sequence, the resource layer state nodes are connected to form a resource layer state sequence, and the operation layer state nodes are connected to form an operation layer state sequence; the Bayesian inference algorithm is applied to calculate the task layer state transition probability matrix, the resource layer state transition probability matrix and the operation layer state transition probability matrix; The task layer state transition probability matrix is ​​set as a constraint condition, and the resource layer state transition probability matrix is ​​modified; the modified resource layer state transition probability matrix is ​​set as a constraint condition, and the operation layer state transition probability matrix is ​​modified; Performing a state clustering operation on the operation layer state sequence to obtain an operation layer clustering result, and recalculating the resource layer state transition probability matrix based on the operation layer clustering result as observation data; performing a state clustering operation on the resource layer state sequence to obtain a resource layer clustering result, and recalculating the task layer state transition probability matrix based on the resource layer clustering result as observation data; A hierarchical Markov decision model is constructed based on the recalculated task layer state transition probability matrix, the resource layer state transition probability matrix and the operation layer state transition probability matrix.

5. The method according to claim 1, characterized in that: Based on the state transition probability, the task state sequence is segmented using a spectral clustering algorithm, and the abnormal state segments are identified in combination with a dynamic time warping algorithm, including: Obtain a state transition probability distribution vector between any two state nodes in the task state sequence, calculate the similarity between corresponding state nodes based on the state transition probability distribution vector, and construct a state node similarity matrix; construct a diagonal matrix based on the sum of row elements of the state node similarity matrix; calculate a normalized Laplace matrix using the state node similarity matrix and the diagonal matrix; Calculate the eigenvalues ​​and eigenvectors of the normalized Laplace matrix, select the eigenvectors corresponding to the minimum non-zero eigenvalues ​​of the corresponding number according to a preset number, construct a feature matrix, perform clustering operation on the row vectors of the feature matrix to obtain cluster labels of state nodes; group state nodes with the same cluster label and continuous in time into state segments; Extracting multidimensional state transition features from the state fragments, performing wavelet transform on the multidimensional state transition features to obtain time-frequency feature coefficients, calculating local energy distribution of the time-frequency feature coefficients, determining core feature coefficients based on the local energy distribution, and constructing a state transition feature sequence; Construct an empty state transfer feature cumulative distance matrix, calculate the time series autocorrelation coefficient of the state transfer feature sequence, determine the initial window size based on the time series autocorrelation coefficient, dynamically adjust the initial window size according to the local fluctuation degree of the state transfer feature sequence, obtain an adaptive window constraint, calculate the direction angle difference of adjacent state transfer feature vectors in the state transfer feature sequence, map the direction angle difference to a state transfer direction penalty factor, and calculate the weighted distance between the current state transfer feature vector and three adjacent cumulative distance points within the range of the adaptive window constraint; Adding the Euclidean distance between the state transfer feature vectors to the minimum value of the weighted distance, and multiplying it by the state transfer direction penalty factor to obtain the corresponding element value of the state transfer feature cumulative distance matrix; Backtracking in the state transition feature cumulative distance matrix to obtain a minimum cumulative distance path, and determining an optimal alignment method and distance value between the state segment and a preset reference sequence template; When the distance value exceeds a preset distance threshold, the state segment is determined to be an abnormal state segment.

6. The method according to claim 1, characterized in that For the abnormal task, a task resource competition matrix is ​​constructed based on the real-time operation data, a weighted directed graph is constructed based on the task resource competition as a task dependency graph, an improved minimum spanning tree algorithm is used to calculate the resource critical path, and a task sorting algorithm based on the resource critical path is used to determine the optimal computing node for re-executing the abnormal task, including: Based on the real-time operation data, for any two computing nodes, the ratio of the minimum value to the maximum value of each type of resource demand is calculated, the sum of the ratios is used as the task resource competition degree between the two computing nodes, and a task resource competition degree matrix is ​​constructed; Based on the task resource competition degree matrix, a weighted directed graph is constructed as a task dependency graph, the computing nodes are used as vertices in the task dependency graph, and the weighted sum of the task resource competition degree and the communication overhead between nodes is used as the initial weight of the edge in the task dependency graph; Calculate the ratio of resource usage to resource capacity of each computing node in the task dependency graph to obtain a resource load factor, multiply the initial weight of the edge in the task dependency graph by the average value of the resource load factors of two connected computing nodes to obtain an improved edge weight; construct a minimum spanning tree based on the improved edge weight to obtain a resource critical path; The inverse of the resource load factor, the inverse of the improved edge weight sum of the connected edges, and the node processing capacity index are weighted to obtain a comprehensive score of each computing node on the resource critical path; the computing nodes on the resource critical path are sorted from high to low according to the comprehensive score, and the computing node with the highest comprehensive score is selected as a candidate execution node; Calculate the difference in task resource contention before and after the abnormal task is migrated to the candidate execution node to obtain a task resource contention change value; determine whether the task resource contention change value exceeds a preset change threshold; if so, select the next computing node in the sorting as the candidate execution node and repeat the calculation until it is less than the preset change threshold to determine the optimal computing node.

7. The method according to claim 1, characterized in that Record the re-execution trajectory information of the abnormal task, calculate the Pearson correlation coefficient between the predicted value and the actual value of the task execution time prediction model to determine the prediction error, and use a time series decomposition algorithm to feed back the prediction error to different network layers of the task execution time prediction model for optimization; When the number of re-executions of the abnormal task reaches a preset maximum number of retries, terminating the abnormal task and returning a failure analysis report includes: Recording the re-execution trajectory information of the abnormal task; obtaining the predicted time series and the actual execution time series of the task execution time prediction model based on the re-execution trajectory information; Calculate the covariance between the predicted time series and the actual execution time series divided by the product of the corresponding standard deviations to obtain the Pearson correlation coefficient as the prediction error; decompose the prediction error into a trend error term, a periodic error term, and a random error term through a time series decomposition algorithm; Adjusting the prediction bias parameter in the task execution time prediction model based on the trend error term, adjusting the timing feature parameter in the task execution time prediction model based on the periodic error term, adjusting the weight parameter in the task execution time prediction model based on the random error term, and recording the adjustment information of the task execution time prediction model; Counting the number of re-executions of the abnormal task, and when the number of re-executions reaches a preset maximum number of retries, terminating the execution of the abnormal task; generating execution statistics based on the re-execution trajectory information, generating error analysis data based on the prediction error, and generating model optimization data based on the adjustment information of the task execution time prediction model; The execution statistics data, the error analysis data and the model optimization data are integrated to generate a failure analysis report and output it.

8. A multi-task monitoring and scheduling system, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to receive a task submission request carrying task identification information, divide the task group according to the task identification information, collect the historical execution time series of the same task group, execute an adaptive filtering algorithm based on kernel density estimation, perform noise reduction processing on the historical execution time series to obtain an initial feature sequence, input the initial feature sequence into a time series feature extraction network including a causal convolution layer and a hole convolution layer, extract the task feature vector, and construct a task execution time prediction model for predicting a maximum execution time threshold based on the task feature vector through a bidirectional gated recurrent neural network; The second unit is used to monitor the execution process of the task, collect the task execution time and real-time operation data, and when the task execution time exceeds the maximum execution time threshold, construct the task operation state space based on the real-time operation data, use the hierarchical Markov decision model to model the task state, convert the real-time operation data into a task state sequence according to the time series, use the Bayesian inference method to calculate the state transition probability, and based on the state transition probability, use the spectral clustering algorithm to segment the task state sequence, and combine the dynamic time warping algorithm to identify abnormal state segments, and when the abnormal state segment is identified, the task is determined as an abnormal task; The third unit is used to construct a task resource competition matrix for the abnormal task based on the real-time operation data, construct a weighted directed graph as a task dependency graph based on the task resource competition, calculate the resource critical path using an improved minimum spanning tree algorithm, and determine the optimal computing node for re-executing the abnormal task using a task sorting algorithm based on the resource critical path; record the re-execution trajectory information of the abnormal task, calculate the Pearson correlation coefficient between the predicted value and the actual value of the task execution time prediction model to determine the prediction error, and use a time series decomposition algorithm to feed back the prediction error to different network layers of the task execution time prediction model for optimization; when the number of re-executions of the abnormal task reaches a preset maximum number of retries, terminate the abnormal task and return a failure analysis report.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Therapy control using motion prediction based on cyclic motion model

    CN109152926A

  • Self-adaptive task scheduling execution unit management method and system

    CN119376903A