Execution unit management method and system for adaptive task scheduling

Through federated learning and differential privacy technology, and task scheduling is carried out in combination with deep learning and reinforcement learning frameworks, the problem of failure to fully consider the relationship between the execution unit in the existing technology is solved, and efficient and flexible task scheduling and resource optimization are achieved.

CN119376903BActive Publication Date: 2025-05-16北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510000923.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-16
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

The existing adaptive task scheduling methods fail to fully consider the complex relationship between execution units, cannot accurately capture the overall state of the system, cannot adapt to dynamically changing tasks and execution environments, and cannot effectively utilize historical execution data.

Method used

Federated learning and differential privacy technology are used to process the load and state data of the execution unit, and timing execution features and topological associations are extracted through two-way long and short-term memory networks and graph structure attention networks, capability evaluation is performed in combination with deep reinforcement learning frameworks, and state prediction is performed through variational autoencoders, and task scheduling is finally realized through multi-level task allocation strategies and global scheduling optimization.

Benefits of technology

It improves the forward-looking and flexible task scheduling, can more accurately capture the state of the execution unit and the dynamic changes in tasks, optimizes resource utilization and load balancing, and improves the robustness and long-term operation effect of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119376903B_ABST
    Figure CN119376903B_ABST
Patent Text Reader

Abstract

The present invention provides an execution unit management method and system for adaptive task scheduling, which relates to the technical field of computer task scheduling, including: collecting execution unit data, extracting features through federated learning and hybrid neural networks, evaluating capabilities and predicting states based on deep reinforcement learning, generating allocation plans in combination with improved algorithms, performing spectral clustering and hierarchical analysis to optimize decomposition, building a multidimensional dependency model, optimizing scheduling through Monte Carlo search and multi-armed bandit algorithms, and realizing adaptive task allocation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer task scheduling, and in particular to an execution unit management method and system for adaptive task scheduling. Background Art

[0002] The management of the execution unit of adaptive task scheduling is a key issue in modern complex systems. As the scale and complexity of the system continue to increase, the traditional static scheduling method can no longer meet the needs of dynamically changing environments;

[0003] In recent years, the development of machine learning and artificial intelligence technology has provided new ideas for solving this problem. Researchers have proposed various adaptive scheduling methods based on machine learning, such as reinforcement learning and deep learning, to improve the flexibility and efficiency of task scheduling.

[0004] However, existing adaptive task scheduling methods still have problems such as not fully considering the complex relationships between execution units, not being able to accurately capture the overall state of the system, not being able to adapt to dynamically changing tasks and execution environments, and not being able to effectively utilize historical execution data.

[0005] Therefore, a solution is urgently needed to solve the problems existing in the prior art. Summary of the invention

[0006] The embodiment of the present invention provides an execution unit management method and system for adaptive task scheduling, which can at least solve some of the problems existing in the prior art.

[0007] A first aspect of an embodiment of the present invention provides an execution unit management method for adaptive task scheduling, comprising:

[0008] Collect the load data and operation status data of the execution unit, establish a distributed data processing node to perform federated learning operations, perform differential privacy desensitization and standardization on the collected data to obtain a standardized data set, build a hybrid neural network of a bidirectional long short-term memory network and a graph structure attention network based on the standardized data set, use the bidirectional long short-term memory network to extract timing execution features, calculate the topological associations between the execution units through the graph structure attention network to obtain a state dependency matrix, input the timing execution features and the state dependency matrix into a deep reinforcement learning framework for proximal policy optimization, obtain the capability evaluation result of the execution unit through alternating training of the policy network and the value network, construct a digital mapping model based on the capability evaluation result to encode the state of the execution unit into the latent space of the variational autoencoder, and output the predicted probability distribution of the future state of the execution unit;

[0009] Based on the predicted probability distribution, task allocation control is performed, and the initial allocation strategy is calculated in combination with the improved soft actor critic algorithm and added to the preset two-way adversarial generative network, multiple candidate allocation schemes are generated and the quality of the scheme is determined in combination with the discriminator, and the candidate allocation scheme with the highest quality score is selected as the initial task allocation scheme, and the spectral clustering operation is performed on the initial task allocation scheme and an adaptive similarity matrix is ​​constructed in combination with the load characteristic network topology of the execution unit, and the number of clusters is determined in combination with the dynamic density peak algorithm and the execution unit is divided into different load level groups, and a hierarchical spatiotemporal graph attention network analysis is performed for each load level group, and the task is decomposed in combination with the hierarchical deep deterministic policy gradient algorithm and the decomposition scheme is optimized according to the adaptive crossover mutation operator to obtain the final decomposition result;

[0010] Based on the final decomposition result, a spatiotemporal heterogeneous graph neural network is constructed and the temporal dependency and spatial constraint of the task are integrated to obtain a multidimensional dependency model, the grouping state of the execution group is added to the multidimensional dependency model, and the suboptimal solution is pruned by Monte Carlo search to obtain a task scheduling plan, the multi-armed bandit algorithm of Thompson sampling is executed on the task scheduling plan to calculate the importance weights of different scheduling goals and adjust the direction, the adjusted scheduling plan is added to a pre-set hierarchical reinforcement learning framework to plan the global scheduling strategy and realize task allocation by executing scheduling instructions, the task execution status is monitored in real time, if there is a performance anomaly, the samples with performance anomaly are learned through the priority experience playback mechanism, the learning results are added to the graph structure neural network and a knowledge graph is constructed to store the execution experience and optimize the global scheduling strategy, and the optimization is repeated to obtain the global optimal scheduling strategy.

[0011] In an optional embodiment,

[0012] Collect the load data and operation status data of the execution unit, establish a distributed data processing node to perform federated learning operations, perform differential privacy desensitization and standardization on the collected data to obtain a standardized data set, build a hybrid neural network of a bidirectional long short-term memory network and a graph structure attention network based on the standardized data set, and use the bidirectional long short-term memory network to extract timing execution features, including:

[0013] Deploy a hierarchical data collection system to collect load data and operating status data of the execution unit, the load data includes CPU usage, memory occupancy, network bandwidth utilization, disk read and write rate and task queue length, and the operating status data includes device temperature, power consumption, vibration frequency and noise;

[0014] Construct a distributed storage cluster to store the load data and operation status data of the execution group, deploy federated learning service nodes to build a hierarchical aggregation architecture, perform local data processing through the bottom nodes in the hierarchical aggregation architecture, perform regional data aggregation through the middle nodes in the hierarchical aggregation architecture, perform global model updates through the top nodes in the hierarchical aggregation architecture, perform differential privacy protection on the load data and operation status data of the execution group based on the Laplace mechanism, and set a privacy budget according to the sensitivity of the data;

[0015] Calculate the noise range according to the set privacy budget, limit the noise range of the load data and the operating status data within the ratio range of the corresponding data sensitivity and the privacy budget, sample Laplace noise from the noise range, add the sampled Laplace noise to the original data, verify whether the data after adding the noise meets the differential privacy protection requirements by combining the privacy budget, and perform piecewise normalization processing on the data that meets the differential privacy protection requirements to obtain a standardized data set;

[0016] A two-layer hybrid neural network model including a time series feature extraction layer and a topological relationship modeling layer is constructed, and the standardized data set is added to the two-layer hybrid neural network model. In the time series feature extraction layer, the standardized data set is segmented by a sliding window to generate a time series data sequence, and forward and reverse processing is performed through a first bidirectional long short-term memory network to obtain a first feature output. The first feature output is subjected to deep feature extraction through a second long short-term memory network to obtain a bidirectional context feature and a time series pooling operation is performed to obtain a time series execution feature.

[0017] In an optional embodiment,

[0018] The topological association between the execution units is calculated through the graph structure attention network to obtain a state dependency matrix, the temporal execution features and the state dependency matrix are input into the deep reinforcement learning framework of proximal policy optimization, the ability evaluation result of the execution unit is obtained by alternately training the policy network and the value network, and a digital mapping model is constructed according to the ability evaluation result to encode the state of the execution unit into the latent space of the variational autoencoder, and the predicted probability distribution of the future state of the output execution unit includes:

[0019] Construct a graph structure attention network in the pre-constructed topological relationship modeling layer, map the execution units into graph structure nodes with attributes, map the physical connection relationships between the execution units into graph structure edges, initialize the weights of the graph structure edges based on the physical distance, construct a multi-head attention mechanism to calculate the attention distribution between nodes, perform normalization on the attention scores to obtain the attention weights, and construct a state dependency matrix based on the attention weights;

[0020] Input the timing execution feature and the state dependency matrix into a proximal policy optimization framework, construct a policy network and a value network in the proximal policy optimization framework, wherein both the policy network and the value network adopt a multi-layer fully connected structure, construct an experience replay buffer to store training samples, sample training batches from the experience replay buffer, and calculate the policy gradient loss of the policy network according to the sampled training batches, wherein the policy gradient loss includes an action probability ratio term, an entropy regularization term, and a value loss term;

[0021] updating the parameters of the policy network according to the policy gradient loss, limiting the difference between the updated policy distribution and the original policy distribution within the confidence domain, calculating the value loss of the value network based on the temporal difference error, wherein the value loss represents the deviation between the predicted value and the actual return, updating the parameters of the value network according to the value loss, using the updated value estimate to calculate the advantage function, and performing policy network optimization and value network optimization alternately until the policy distribution converges and the value estimate is stable;

[0022] Based on the trained proximal policy optimization framework, the execution unit capability is evaluated, and a digital mapping model of a variational autoencoder structure is constructed, wherein the variational autoencoder structure includes an encoder and a decoder. The encoder is used to map the state vector of the execution unit to a low-dimensional latent space, and reparameterized sampling is performed on the latent variables to generate multiple samples, and the predicted probability distribution of the future state of the execution unit is estimated and output.

[0023] In an optional embodiment,

[0024] Based on the predicted probability distribution, task allocation control is performed, and the initial allocation strategy is calculated in combination with the improved soft actor critic algorithm and added to the preset two-way adversarial generative network, multiple candidate allocation schemes are generated and the quality of the scheme is determined in combination with the discriminator, and the candidate allocation scheme with the highest quality score is selected as the initial task allocation scheme, and the spectral clustering operation is performed on the initial task allocation scheme and the adaptive similarity matrix is ​​constructed in combination with the load characteristics of the execution unit. The network topology includes:

[0025] An improved soft actor-critic model is constructed, in which the actor network adopts a multi-layer fully connected structure and is regularized by a Dropout layer, and the critic network adopts a dual-branch structure to process state information and action information respectively;

[0026] Based on the predicted probability distribution, the predicted distribution of processor usage, the predicted distribution of memory occupancy, and the predicted distribution of network bandwidth corresponding to the execution group are determined and added as input state information to the improved soft actor-critic model, and the initial allocation strategy is obtained by the probability distribution of the task allocation action output by the actor network. The critic network combines the state action to evaluate the computational value, constructs an experience replay buffer to store training samples and regularly updates sample data, calculates the temporal difference target value and policy gradient by sampling training batches, and updates network parameters based on the adaptive moment estimation optimizer;

[0027] The initial allocation strategy is added to a pre-set bidirectional adversarial generative network including a generator and a discriminator. The generator adopts a multi-layer transposed convolution structure and introduces a residual connection mechanism. Each layer of transposed convolution is connected to a batch normalization layer and a nonlinear activation function. The discriminator adopts a multi-scale structure including a global discriminator and a local discriminator. The global discriminator evaluates the overall allocation plan, and the local discriminator evaluates the detailed task allocation. The generator generates multiple candidate allocation plans based on randomly sampled latent variables. The candidate allocation plans include task mapping relationships, task priorities, and resource demand constraint information. The discriminator evaluates the quality score of each candidate allocation plan based on load balancing, network communication overhead, and task dependency.

[0028] The load characteristics, topology characteristics and historical performance characteristics of the execution groups are extracted, and based on the load characteristics, topology characteristics and historical performance characteristics, the similarity values ​​between the execution groups are calculated in combination with an adaptive Gaussian kernel function. A similarity matrix is ​​constructed based on the similarity values, and a spectral clustering operation is performed on the candidate scheme with the highest quality score and the similarity matrix to obtain an initial clustering result.

[0029] In an optional embodiment,

[0030] Combined with the dynamic density peak algorithm, the number of clusters is determined and the execution units are divided into different load level groups. For each load level group, a hierarchical spatiotemporal graph attention network analysis is performed. Combined with the hierarchical deep deterministic policy gradient algorithm, the task is decomposed and the decomposition scheme is optimized according to the adaptive crossover mutation operator. The final decomposition results include:

[0031] Based on the pre-acquired initial clustering results and the dynamic density peak algorithm, the local density value and relative distance value of the execution group are calculated. The local density is determined by counting the neighborhood samples through the cutoff kernel function. The influence range of the density peak point is determined based on the density reachability analysis. The cluster center is iteratively updated to divide the execution group into different load level groups.

[0032] Construct a hierarchical spatiotemporal graph attention network to process the load level group, use causal convolution layers with different expansion rates in the time dimension to obtain multi-scale temporal features, and construct a dynamic graph structure modeling execution unit dependency in the space dimension, wherein the spatiotemporal graph attention network includes parallel time channels and space channels, captures temporal dependencies through a self-attention mechanism, aggregates node information through a graph attention layer, constructs a hierarchical policy network to perform task decomposition, splits tasks into subtask sets based on computational complexity and data dependencies through a high-level policy network, assigns subtask priorities through a middle-level policy network, determines the execution scheduling order through a low-level policy network, and uses a deep deterministic policy gradient algorithm to train each layer of policy networks to obtain an initial decomposition plan;

[0033] Based on the adaptive crossover mutation operator, an adaptive evolutionary operator is constructed to optimize the task decomposition scheme. The crossover probability and crossover position are dynamically adjusted according to the fitness value and population diversity. The mutation intensity is adaptively adjusted based on the population convergence degree. The non-dominated sorting method is used to select the optimization scheme that meets the load balancing and execution efficiency constraints to obtain the final decomposition result.

[0034] In an optional embodiment,

[0035] Based on the final decomposition result, a spatiotemporal heterogeneous graph neural network is constructed and the temporal dependency and spatial constraint of the task are integrated to obtain a multidimensional dependency model. The grouping state of the execution group is added to the multidimensional dependency. The suboptimal solution is pruned by Monte Carlo search to obtain a task scheduling plan. The Thompson sampling multi-armed bandit algorithm is executed on the task scheduling plan to calculate the importance weights of different scheduling targets and adjust the direction. The adjusted scheduling plan is added to a pre-set hierarchical reinforcement learning framework to plan a global scheduling strategy and implement task allocation by executing scheduling instructions. The task execution status is monitored in real time. If there is a performance anomaly, the samples with performance anomalies are learned through a priority experience playback mechanism. The learning results are added to a graph structure neural network and a knowledge graph is constructed to store execution experience and optimize the global scheduling strategy. Repeated optimization to obtain the global optimal scheduling strategy includes:

[0036] Based on the final decomposition result, a spatiotemporal heterogeneous graph neural network is constructed, and task nodes are divided into computing-intensive nodes, storage-intensive nodes, and communication-intensive nodes, wherein the task nodes include CPU requirements, memory requirements, network bandwidth requirements, expected execution time, task priority, data dependencies, resource constraints, and deadline requirements, and the number of processor cores, memory capacity, network bandwidth, and current load of the execution unit node are constructed as node attributes;

[0037] Extract serial dependencies, parallel dependencies, and semi-dependencies between adjacent tasks to construct timing dependency edges. The serial dependency indicates that tasks must be executed in a strict order, the parallel dependency indicates that tasks can be executed simultaneously, and the semi-dependency indicates that only part of the data dependency needs to be satisfied. Different weights are assigned to each timing dependency edge based on the dependency, where the serial dependency has the highest weight and the parallel dependency has the lowest weight.

[0038] The cosine similarity between the resource requirement vector of the computing task and the resource capacity vector of the execution unit is used to construct a spatial constraint edge, and a high-weight edge, a medium-weight edge, and a low-weight edge are constructed according to the size of the cosine similarity, wherein the similarity corresponds to different levels of weights from high to low, and when the similarity is lower than a preset similarity threshold, no edge connection is constructed;

[0039] Collect the group status information of the execution units within a preset time period to construct dynamic attributes, wherein the group status information includes the CPU utilization, task queue length, and average response time of the computing load group, the memory occupancy, disk IO rate, and cache hit rate of the storage load group, and the network throughput, message loss rate, and end-to-end delay of the communication load group. Based on the Monte Carlo tree search tree method, multiple sampling sequences are expanded from the initial state, and the sampling depth is set according to the total number of tasks. The sequence score is calculated based on the completion time, resource utilization, and load balancing. The sequence with a standardized score result that is higher than a preset score threshold is used as the task scheduling plan;

[0040] A multi-armed bandit model is constructed, and a multi-armed bandit algorithm of Thompson sampling is executed on the task scheduling scheme. The multi-armed bandit algorithm sets minimizing completion time, maximizing resource utilization, and optimizing load balancing as scheduling arms, and calculates the importance weights of different scheduling objectives through Thompson sampling, and adjusts the direction of the task scheduling scheme based on the importance weights;

[0041] Construct a hierarchical reinforcement learning framework including a global policy network and a local execution network, wherein the global policy network and the local execution network are both composed of multiple layers of fully connected layers, input the adjusted task scheduling scheme into the hierarchical reinforcement learning framework, the global policy network receives the system state vector and outputs a global scheduling policy, the local execution network converts the scheduling decision into a scheduling instruction, collects performance data on task execution progress, resource usage, and system response time at preset time intervals, divides anomalies into multiple levels according to the degree to which performance indicators deviate from expectations, and uses abnormal execution sequences as key training samples;

[0042] A three-layer knowledge graph is constructed to store execution experience, wherein the knowledge graph includes a task feature layer, an execution environment layer, and a scheduling strategy layer. The information between layers is connected by associative edges and the edge weights are set based on the association strength. The global scheduling strategy is optimized based on the knowledge graph, and the network parameters are updated through the policy gradient method to make the network output biased towards the scheduling mode that has been historically verified to be effective, until the scheduling performance tends to stabilize and the global optimal scheduling strategy is obtained.

[0043] In an optional embodiment,

[0044] The multi-armed bandit algorithm of Thompson sampling is executed on the task scheduling scheme, wherein the multi-armed bandit algorithm sets minimizing completion time, maximizing resource utilization, and optimizing load balancing as scheduling arms, and calculates the importance weights of different scheduling objectives through Thompson sampling. Adjusting the direction of the task scheduling scheme based on the importance weights includes:

[0045] Construct a multi-armed bandit model, set minimizing completion time, maximizing resource utilization, and optimizing load balancing as the scheduling arm, create a Beta distribution as a prior distribution for the scheduling arm, and initialize the number of success parameters, number of failure parameters, cumulative reward value, and number of selections of the scheduling arm;

[0046] Performing Thompson sampling on each scheduling arm, extracting sample values ​​from the Beta distribution of the scheduling arm, the sample values ​​representing the expected rate of return of the scheduling arm, normalizing the sample values ​​to obtain importance weights, the importance weights being obtained by dividing the sample values ​​of each scheduling arm by the sum of the sample values ​​of all scheduling arms;

[0047] Calculate a comprehensive score of the scheduling scheme based on the importance weight, the comprehensive score includes a completion time score, a resource utilization score, and a load balance score, the completion time score is obtained by multiplying the importance weight of the first scheduling arm by the inverse of the completion time, the resource utilization score is obtained by multiplying the importance weight of the second scheduling arm by the resource utilization, and the load balance score is obtained by multiplying the importance weight of the third scheduling arm by the inverse of the load variance, and the comprehensive score is obtained by adding the completion time score, the resource utilization score, and the load balance score;

[0048] The task scheduling scheme is adjusted according to the comprehensive score, including reordering task priorities, adjusting resource allocation ratios, and updating task execution order.

[0049] A second aspect of an embodiment of the present invention provides an execution unit management system for adaptive task scheduling, including:

[0050] The first unit is used to collect the load data and operation status data of the execution unit, establish a distributed data processing node to perform federated learning operations, perform differential privacy desensitization and standardization on the collected data to obtain a standardized data set, build a hybrid neural network of a bidirectional long short-term memory network and a graph structure attention network based on the standardized data set, use the bidirectional long short-term memory network to extract timing execution features, calculate the topological associations between the execution units through the graph structure attention network to obtain a state dependency matrix, input the timing execution features and the state dependency matrix into a deep reinforcement learning framework for proximal policy optimization, obtain the capability evaluation result of the execution unit through alternating training of the policy network and the value network, construct a digital mapping model based on the capability evaluation result to encode the state of the execution unit into the latent space of the variational autoencoder, and output the predicted probability distribution of the future state of the execution unit;

[0051] The second unit is used to perform task allocation control based on the predicted probability distribution, calculate the initial allocation strategy in combination with the improved soft actor critic algorithm and add it to a pre-set two-way adversarial generative network, generate multiple candidate allocation schemes and determine the quality of the scheme in combination with the discriminator, select the candidate allocation scheme with the highest quality score as the initial task allocation scheme, perform spectral clustering operation on the initial task allocation scheme and construct an adaptive similarity matrix in combination with the load characteristic network topology of the execution unit, determine the number of clusters in combination with the dynamic density peak algorithm and divide the execution unit into different load level groups, perform hierarchical spatiotemporal graph attention network analysis on each load level group, perform task decomposition in combination with the hierarchical deep deterministic policy gradient algorithm and optimize the decomposition scheme according to the adaptive crossover mutation operator to obtain the final decomposition result;

[0052] The third unit is used to construct a spatiotemporal heterogeneous graph neural network based on the final decomposition result and integrate the temporal dependency and spatial constraint of the task to obtain a multidimensional dependency model, add the grouping state of the execution group to the multidimensional dependency model, prune the suboptimal solution through Monte Carlo search to obtain a task scheduling plan, execute the Thompson sampling multi-armed bandit algorithm on the task scheduling plan to calculate the importance weights of different scheduling goals and adjust the direction, add the adjusted scheduling plan to a pre-set hierarchical reinforcement learning framework to plan the global scheduling strategy and realize task allocation by executing scheduling instructions, monitor the task execution status in real time, if there is a performance anomaly, learn the samples with performance anomaly through the priority experience playback mechanism, add the learning results to the graph structure neural network and construct a knowledge graph to store the execution experience and optimize the global scheduling strategy, and repeat the optimization to obtain the global optimal scheduling strategy.

[0053] According to a third aspect of the embodiments of the present invention,

[0054] An electronic device is provided, comprising:

[0055] processor;

[0056] a memory for storing processor-executable instructions;

[0057] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0058] A fourth aspect of the embodiments of the present invention is:

[0059] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0060] In the present invention, the secure processing of data is achieved through federated learning and differential privacy technology, and a hybrid neural network is constructed by combining a bidirectional long short-term memory network and a graph attention network, which effectively extracts the timing characteristics and topological associations of the execution group, laying a solid data foundation for subsequent task scheduling, and adopts a variational autoencoder to encode the state of the execution group, so as to achieve accurate prediction of the future state, greatly improving the foresight and flexibility of scheduling, combining the soft actor critic algorithm with the bidirectional adversarial generative network to generate a high-quality initial task allocation scheme, and realizing the adaptive grouping of the execution group through spectral clustering and dynamic density peak algorithm, and introducing a hierarchical spatiotemporal graph attention network for refined task decomposition, which greatly improves the pertinence and rationality of task allocation, and the multi-level and multi-angle task allocation strategy effectively balances the load of the execution group and improves the overall execution efficiency, and formulates the global scheduling strategy through the hierarchical reinforcement learning framework, and introduces the priority experience playback mechanism and the knowledge graph to continuously optimize the scheduling strategy, so as to realize the adaptive learning and continuous evolution of the scheduling system, and greatly improve the robustness and long-term operation effect of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 A schematic flow chart of an execution unit management method for adaptive task scheduling according to an embodiment of the present invention;

[0062] Figure 2 It is a structural diagram of an execution unit management system for adaptive task scheduling according to an embodiment of the present invention. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0064] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0065] Figure 1 FIG. 1 is a flow chart of an execution unit management method for adaptive task scheduling according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0066] S1. Collect the load data and operation status data of the execution unit, establish a distributed data processing node to perform federated learning operations, perform differential privacy desensitization and standardization on the collected data to obtain a standardized data set, build a hybrid neural network of a bidirectional long short-term memory network and a graph structure attention network based on the standardized data set, use the bidirectional long short-term memory network to extract timing execution features, calculate the topological associations between the execution units through the graph structure attention network to obtain a state dependency matrix, input the timing execution features and the state dependency matrix into a deep reinforcement learning framework for proximal policy optimization, obtain the capability evaluation results of the execution unit through alternating training of the policy network and the value network, construct a digital mapping model based on the capability evaluation results to encode the state of the execution unit into the latent space of the variational autoencoder, and output the predicted probability distribution of the future state of the execution unit;

[0067] The federated learning operation is a distributed machine learning method that aims to solve the problem of data privacy protection. The differential privacy desensitization is a data privacy protection technology that aims to blur data by adding noise so that changes in any single data point will not significantly affect the final analysis results. The state dependency matrix is ​​a matrix that represents the state transition law in a dynamic system, which describes the transition probability or transition function between different states in the system and is usually used in Markov decision processes. The digital mapping model is a model that maps discrete digital or symbolic data to continuous or high-dimensional space. Through mapping rules or algorithms, each element of the original data is converted into a new representation for subsequent calculations, analysis or prediction tasks.

[0068] In an optional embodiment,

[0069] Collect the load data and operation status data of the execution unit, establish a distributed data processing node to perform federated learning operations, perform differential privacy desensitization and standardization on the collected data to obtain a standardized data set, build a hybrid neural network of a bidirectional long short-term memory network and a graph structure attention network based on the standardized data set, and use the bidirectional long short-term memory network to extract timing execution features, including:

[0070] Deploy a hierarchical data collection system to collect load data and operating status data of the execution unit, the load data includes CPU usage, memory occupancy, network bandwidth utilization, disk read and write rate and task queue length, and the operating status data includes device temperature, power consumption, vibration frequency and noise;

[0071] Construct a distributed storage cluster to store the load data and operation status data of the execution group, deploy federated learning service nodes to build a hierarchical aggregation architecture, perform local data processing through the bottom nodes in the hierarchical aggregation architecture, perform regional data aggregation through the middle nodes in the hierarchical aggregation architecture, perform global model updates through the top nodes in the hierarchical aggregation architecture, perform differential privacy protection on the load data and operation status data of the execution group based on the Laplace mechanism, and set a privacy budget according to the sensitivity of the data;

[0072] Calculate the noise range according to the set privacy budget, limit the noise range of the load data and the operating status data within the ratio range of the corresponding data sensitivity and the privacy budget, sample Laplace noise from the noise range, add the sampled Laplace noise to the original data, verify whether the data after adding the noise meets the differential privacy protection requirements by combining the privacy budget, and perform piecewise normalization processing on the data that meets the differential privacy protection requirements to obtain a standardized data set;

[0073] A two-layer hybrid neural network model including a time series feature extraction layer and a topological relationship modeling layer is constructed, and the standardized data set is added to the two-layer hybrid neural network model. In the time series feature extraction layer, the standardized data set is segmented by a sliding window to generate a time series data sequence, and forward and reverse processing is performed through a first bidirectional long short-term memory network to obtain a first feature output. The first feature output is subjected to deep feature extraction through a second long short-term memory network to obtain a bidirectional context feature and a time series pooling operation is performed to obtain a time series execution feature.

[0074] The Laplace mechanism is a commonly used differential privacy implementation method, which protects data privacy by adding noise from the Laplace distribution to the query results. The Laplace noise refers to random noise generated according to the Laplace distribution. The Laplace distribution has a sharper peak and a heavier tail, and can effectively add noise to ensure the protection effect of differential privacy.

[0075] A hierarchical data collection system is deployed to collect the load data and operating status data of the execution unit. The load data includes the CPU usage rate, memory occupancy rate, network bandwidth utilization rate, disk read and write rate, and task queue length. The operating status data includes device temperature, power consumption, vibration frequency, and noise. The collection system adopts a hierarchical architecture. Multiple data collection nodes are set up at the bottom layer, which are distributed on each execution unit to collect raw data; data aggregation nodes are set up in the middle layer to perform preliminary processing and aggregation of the data collected at the bottom layer; and central data processing nodes are set up at the top layer to uniformly store and analyze the aggregated data.

[0076] For example, for a certain execution unit, the collected load data may be: CPU usage 80%, memory occupancy 75%, network bandwidth utilization 60%, disk read / write rate 100MB / s, task queue length 20. The running status data may be: device temperature 45°C, power consumption 500W, vibration frequency 60Hz, noise 65dB. These data will be reported to the aggregation node in real time by the collection node.

[0077] Build a distributed storage cluster to store the collected data, and deploy federated learning service nodes to build a hierarchical aggregation architecture. The distributed storage cluster uses technologies such as HDFS to distribute and store data on multiple nodes to improve storage capacity and access efficiency. The federated learning service nodes adopt a hierarchical architecture, with the bottom-level nodes responsible for local data processing, the middle-level nodes responsible for regional data aggregation, and the top-level nodes responsible for global model updates.

[0078] Based on the Laplace mechanism, differential privacy protection is performed on the data. The privacy budget ε is set according to the sensitivity of the data. For example, ε=0.1 is set for CPU usage and ε=0.5 is set for device temperature. The noise range is calculated and the noise is limited to the ratio of data sensitivity to the privacy budget. Laplace noise is sampled from the noise range and added to the original data. The combined privacy budget is used to verify whether the data after adding noise meets the differential privacy protection requirements.

[0079] For example, adding noise to a CPU usage rate of 80% may change it to 78.5%, and adding noise to a device temperature of 45°C may change it to 44.2°C. This protects the privacy of the original data while retaining the statistical characteristics of the data.

[0080] The data that meets the requirements of differential privacy protection is subjected to piecewise normalization to obtain a standardized data set. Piecewise normalization divides the data range into multiple intervals, and normalizes the data in each interval separately to avoid interference between data of different dimensions.

[0081] For example, the CPU usage can be divided into three intervals: 0-50%, 50%-80%, and 80%-100%, and normalized separately. In the processed standardized data set, the values ​​of each indicator are mapped to between 0 and 1, which is convenient for subsequent modeling and analysis.

[0082] A two-layer hybrid neural network model consisting of a time series feature extraction layer and a topological relationship modeling layer is constructed, and the standardized data set is input into the model for processing. In the time series feature extraction layer, the data is first segmented by a sliding window to generate a time series data sequence. The window size can be set to 24 hours with a step size of 1 hour, so that a continuous 24-hour data sequence can be generated.

[0083] The data sequence is input into the first bidirectional LSTM network for forward and reverse processing. The bidirectional LSTM network contains two LSTM layers, forward and reverse, which can capture the forward and backward dependencies of the sequence at the same time. The forward LSTM network processes backward from the beginning of the sequence, and the reverse LSTM network processes forward from the end of the sequence. Finally, the outputs of the two directions are spliced ​​to obtain the first feature output.

[0084] The first feature output is input into the second long short-term memory network for deep feature extraction. The second long short-term memory network adopts a multi-layer stacking structure to further extract deep temporal features. Finally, the temporal pooling operation is performed on the output of the long short-term memory network to obtain a temporal execution feature vector of a fixed dimension. Temporal pooling can use methods such as maximum pooling or average pooling to map sequence features of different lengths to the same dimension.

[0085] For example, for a 24-hour data sequence of a certain execution unit, a 48-dimensional feature vector is obtained after being processed by a bidirectional long short-term memory network, and then a 128-dimensional time series execution feature vector is finally obtained after a three-layer long short-term memory network and time series maximum pooling. This feature vector contains the time series change characteristics of the execution unit load and operating status, which can be used for subsequent prediction and analysis tasks.

[0086] In this embodiment, through hierarchical data collection and distributed storage, efficient collection and management of large-scale execution unit data are achieved, providing a reliable data foundation for subsequent analysis. Federated learning and differential privacy technologies are used to achieve collaborative modeling of multi-party data while protecting data privacy, overcoming the limitations of traditional centralized learning. The constructed two-layer hybrid neural network model can effectively extract timing execution features, capture the timing dependencies and contextual information of the data, and provide high-quality feature representation for subsequent predictive analysis.

[0087] In an optional embodiment,

[0088] The topological association between the execution units is calculated through the graph structure attention network to obtain a state dependency matrix, the temporal execution features and the state dependency matrix are input into the deep reinforcement learning framework of proximal policy optimization, the ability evaluation result of the execution unit is obtained by alternately training the policy network and the value network, and a digital mapping model is constructed according to the ability evaluation result to encode the state of the execution unit into the latent space of the variational autoencoder, and the predicted probability distribution of the future state of the output execution unit includes:

[0089] Construct a graph structure attention network in the pre-constructed topological relationship modeling layer, map the execution units into graph structure nodes with attributes, map the physical connection relationships between the execution units into graph structure edges, initialize the weights of the graph structure edges based on the physical distance, construct a multi-head attention mechanism to calculate the attention distribution between nodes, perform normalization on the attention scores to obtain the attention weights, and construct a state dependency matrix based on the attention weights;

[0090] Input the timing execution feature and the state dependency matrix into a proximal policy optimization framework, construct a policy network and a value network in the proximal policy optimization framework, wherein both the policy network and the value network adopt a multi-layer fully connected structure, construct an experience replay buffer to store training samples, sample training batches from the experience replay buffer, and calculate the policy gradient loss of the policy network according to the sampled training batches, wherein the policy gradient loss includes an action probability ratio term, an entropy regularization term, and a value loss term;

[0091] updating the parameters of the policy network according to the policy gradient loss, limiting the difference between the updated policy distribution and the original policy distribution within the confidence domain, calculating the value loss of the value network based on the temporal difference error, wherein the value loss represents the deviation between the predicted value and the actual return, updating the parameters of the value network according to the value loss, using the updated value estimate to calculate the advantage function, and performing policy network optimization and value network optimization alternately until the policy distribution converges and the value estimate is stable;

[0092] Based on the trained proximal policy optimization framework, the execution unit capability is evaluated, and a digital mapping model of a variational autoencoder structure is constructed, wherein the variational autoencoder structure includes an encoder and a decoder. The encoder is used to map the state vector of the execution unit to a low-dimensional latent space, and reparameterized sampling is performed on the latent variables to generate multiple samples, and the predicted probability distribution of the future state of the execution unit is estimated and output.

[0093] The proximal policy optimization framework is a reinforcement learning algorithm that aims to optimize the training process of the policy network. The temporal difference error is an error metric in reinforcement learning, which is usually used to evaluate the gap between the current strategy and future rewards. The value loss is a loss function used in reinforcement learning to measure the difference between the predicted value and the actual reward, which represents the error between the model's prediction of the value function of a certain state and the actual reward obtained.

[0094] Construct a graph structure attention network in the pre-built topological relationship modeling layer. Map each execution unit to a node in the graph structure, and the node attributes include the state feature vector of the execution unit. Map the physical connection relationship between the execution units to the edge in the graph structure, and initialize the edge weight to the inverse of the physical distance between the two execution units. For example, if the physical distance between execution units A and B is 100 meters, the corresponding edge weight is initialized to 0.01.

[0095] A multi-head attention mechanism is constructed to calculate the attention distribution between nodes. For each attention head, the attribute feature vector of the node is mapped to the query vector, key vector and value vector through linear transformation. The dot product of the query vector and the key vector is calculated to obtain the original attention score. The original attention score is scaled by dividing it by the square root of the key vector dimension and normalized by the softmax function to obtain the final attention weight. Finally, the attention weight is multiplied by the value vector and summed to obtain the output of the attention head. The outputs of multiple attention heads are concatenated and linearly transformed again to obtain the final node representation. The state dependency matrix is ​​constructed based on the calculated attention weights, and each element in the matrix represents the dependency strength between two execution units.

[0096] The timing execution features and state dependency matrix are input into the proximal policy optimization framework for training. In this framework, a policy network and a value network are constructed, and both networks adopt a multi-layer fully connected structure. The policy network is used to generate action probability distribution, and the value network is used to estimate the state value. An experience replay buffer is constructed to store training samples, including information such as state, action, reward, and next state. A training batch is randomly sampled from the buffer, for example, containing 32 samples.

[0097] For the sampled training batch, the policy gradient loss of the policy network is first calculated. This loss consists of three parts: the action probability ratio term, the entropy regularization term, and the value loss term. The action probability ratio term measures the ratio of the action probabilities under the new and old policies, the entropy regularization term is used to encourage policy exploration, and the value loss term is used to guide the policy to update in the direction of high value. The parameters of the policy network are updated according to the calculated policy gradient loss. During the update process, the KL divergence between the updated policy distribution and the original policy distribution is limited to not exceed a preset threshold, such as 0.01, to ensure the stability of the policy update.

[0098] Calculate the value loss of the value network based on the temporal difference error. This loss characterizes the deviation between the predicted value and the actual return. Update the parameters of the value network based on the calculated value loss. Use the updated value estimate to calculate the advantage function as the baseline for the policy gradient. Alternate between policy network optimization and value network optimization until the policy distribution converges and the value estimate is stable. For example, you can set the maximum number of iterations to 1000, and terminate training early if the policy KL divergence is less than 0.001 and the value loss change is less than 0.01 for 10 consecutive iterations.

[0099] The ability of the execution unit is evaluated based on the trained proximal policy optimization framework. The state characteristics of the execution unit are input into the trained policy network to obtain the optimal action probability distribution under this state. The higher the action probability, the stronger the ability of the execution unit. For example, for a certain execution unit, the action probability distribution obtained after inputting its current state characteristics is [0.1, 0.2, 0.7], then the third action with the highest probability is selected as the ability evaluation result.

[0100] Construct a digital mapping model of variational autoencoder structure. The model consists of two parts: encoder and decoder. The encoder consists of a multi-layer fully connected network, which maps the state vector of the actuator group to a low-dimensional latent space. For example, a 100-dimensional state vector is mapped to a 10-dimensional latent variable. Reparameterized sampling is performed on the latent variable to generate multiple sample points. The decoder also consists of a multi-layer fully connected network, which reconstructs the sample points in the latent space into vectors in the original state space. The variational autoencoder is trained by minimizing the reconstruction error and KL divergence. After training, the current state vector is input, mapped to the latent space through the encoder and sampled multiple times, and then reconstructed by the decoder to obtain multiple future state samples. Based on these samples, the predicted probability distribution of the future state of the actuator group is estimated and output.

[0101] In this embodiment, the topological relationship between the execution units is effectively modeled through the graph structure attention network, which overcomes the limitation of the traditional method of ignoring the mutual influence between the units and improves the accuracy of the characterization of the state dependency relationship. The deep reinforcement learning framework with proximal strategy optimization is used to evaluate the capacity of the execution units, avoiding the excessive reliance of traditional methods on historical data, and can better adapt to the dynamically changing environment, thereby improving the reliability and robustness of the evaluation results. The digital mapping model based on the variational autoencoder realizes the probabilistic prediction of the future state of the execution unit. Compared with the deterministic prediction method, it can more comprehensively characterize the uncertainty and provide richer information support for subsequent decision-making.

[0102] S2. Perform task allocation control based on the predicted probability distribution, calculate the initial allocation strategy in combination with the improved soft actor critic algorithm and add it to the pre-set two-way adversarial generative network, generate multiple candidate allocation schemes and determine the quality of the scheme in combination with the discriminator, select the candidate allocation scheme with the highest quality score as the initial task allocation scheme, perform spectral clustering operation on the initial task allocation scheme and construct an adaptive similarity matrix in combination with the load characteristic network topology of the execution unit, determine the number of clusters in combination with the dynamic density peak algorithm and divide the execution unit into different load level groups, perform hierarchical spatiotemporal graph attention network analysis on each load level group, perform task decomposition in combination with the hierarchical deep deterministic policy gradient algorithm and optimize the decomposition scheme according to the adaptive crossover mutation operator to obtain the final decomposition result;

[0103] The improved soft actor critic algorithm is a reinforcement learning algorithm and belongs to the policy gradient method. By introducing the entropy term as part of the reward, the strategy can explore more action options, thereby improving learning efficiency and maintaining better exploratory performance. The bidirectional generative adversarial network is a variant of the generative adversarial network (GAN) with two generators and two discriminators. Data is generated by interacting with each other, and the discriminator optimizes the network by judging the generated data and the real data. The spectral clustering operation is a clustering method based on graph theory, which mainly groups data by analyzing the similarity matrix (or adjacency matrix) of data points. The load characteristic network topology is a topological structure used to represent the load characteristics in the power network or communication network. By analyzing the characteristics of nodes (such as power stations, communication stations, etc.) and edges (such as transmission lines, links, etc.) in the network, the load characteristic network topology is a topological structure used to represent the load characteristics in the power network or communication network. The load-feature network topology can reflect the overall load distribution of the network and help perform tasks such as resource optimization and load balancing. The dynamic density peak algorithm is a clustering algorithm that aims to identify cluster centers by analyzing the local density of data. By calculating the density of each data point and the distance to the nearest point, the center point of the cluster is dynamically determined and grouped. The hierarchical spatiotemporal graph attention network is a deep learning model that combines graph neural networks and spatiotemporal data processing. By learning spatiotemporal features and weighting the attention mechanism at different levels of the graph, it can better capture the complex correlations in spatiotemporal data. The hierarchical deep deterministic policy gradient algorithm is a reinforcement learning algorithm that is suitable for problems in continuous action spaces. It combines a hierarchical structure and a deep learning model, and optimizes policies at multiple levels so that the model can make decisions efficiently in complex tasks.

[0104] In an optional embodiment,

[0105] Based on the predicted probability distribution, task allocation control is performed, and the initial allocation strategy is calculated in combination with the improved soft actor critic algorithm and added to the preset two-way adversarial generative network, multiple candidate allocation schemes are generated and the quality of the scheme is determined in combination with the discriminator, and the candidate allocation scheme with the highest quality score is selected as the initial task allocation scheme, and the spectral clustering operation is performed on the initial task allocation scheme and the adaptive similarity matrix is ​​constructed in combination with the load characteristics of the execution unit. The network topology includes:

[0106] An improved soft actor-critic model is constructed, in which the actor network adopts a multi-layer fully connected structure and is regularized by a Dropout layer, and the critic network adopts a dual-branch structure to process state information and action information respectively;

[0107] Based on the predicted probability distribution, the predicted distribution of processor usage, the predicted distribution of memory occupancy, and the predicted distribution of network bandwidth corresponding to the execution group are determined and added as input state information to the improved soft actor-critic model, and the initial allocation strategy is obtained by the probability distribution of the task allocation action output by the actor network. The critic network combines the state action to evaluate the computational value, constructs an experience replay buffer to store training samples and regularly updates sample data, calculates the temporal difference target value and policy gradient by sampling training batches, and updates network parameters based on the adaptive moment estimation optimizer;

[0108] The initial allocation strategy is added to a pre-set bidirectional adversarial generative network including a generator and a discriminator. The generator adopts a multi-layer transposed convolution structure and introduces a residual connection mechanism. Each layer of transposed convolution is connected to a batch normalization layer and a nonlinear activation function. The discriminator adopts a multi-scale structure including a global discriminator and a local discriminator. The global discriminator evaluates the overall allocation plan, and the local discriminator evaluates the detailed task allocation. The generator generates multiple candidate allocation plans based on randomly sampled latent variables. The candidate allocation plans include task mapping relationships, task priorities, and resource demand constraint information. The discriminator evaluates the quality score of each candidate allocation plan based on load balancing, network communication overhead, and task dependency.

[0109] The load characteristics, topology characteristics and historical performance characteristics of the execution groups are extracted, and based on the load characteristics, topology characteristics and historical performance characteristics, the similarity values ​​between the execution groups are calculated in combination with an adaptive Gaussian kernel function. A similarity matrix is ​​constructed based on the similarity values, and a spectral clustering operation is performed on the candidate scheme with the highest quality score and the similarity matrix to obtain an initial clustering result.

[0110] The sampling training batch is a batch sample strategy used when training deep learning models. A certain number of samples are randomly selected from the training data to form a training batch for network training. The temporal difference target value is a target value commonly used in reinforcement learning. It is calculated by estimating the difference between the value of the current state and the expected reward of the future state. The policy gradient is an optimization method in reinforcement learning. The strategy is directly optimized by calculating the gradient of the strategy to the expected cumulative reward. The adaptive Gaussian kernel function is a method for adaptively adjusting kernel function parameters, which is usually used in machine learning algorithms such as support vector machines and kernel ridge regression.

[0111] An improved soft actor-critic model is constructed. In this model, the actor network adopts a 4-layer fully connected structure, with 256, 128, 64, and 32 neurons in each layer, and a dropout layer is added after each layer with a dropout ratio of 0.3. The critic network adopts a dual-branch structure, with both the state processing branch and the action processing branch being 3-layer fully connected networks, with 128, 64, and 32 neurons, respectively. Finally, the outputs of the two branches are concatenated and passed through a fully connected layer to obtain the value evaluation result.

[0112] Determine the predicted resource usage distribution of the execution group based on the predicted probability distribution. Specifically, for each execution group, extract its processor usage, memory occupancy, and network bandwidth data for the past 24 hours, and use the long short-term memory network to predict the resource usage in the next hour to obtain the predicted probability distribution. For example, for a certain execution group, the predicted processor usage distribution has a mean of 60% and a standard deviation of 5%, a memory occupancy distribution has a mean of 40% and a standard deviation of 3%, and a network bandwidth distribution has a mean of 500Mbps and a standard deviation of 50Mbps. These predicted distributions are input into the improved soft actor critic model as state information.

[0113] The actor network outputs the probability distribution of task allocation actions to obtain the initial allocation strategy. For example, for 5 tasks and 3 execution units, the output probability distribution matrix is: [[0.6, 0.3, 0.1], [0.2, 0.5, 0.3], [0.4, 0.4, 0.2], [0.1, 0.7, 0.2], [0.3, 0.2, 0.5]]. The critic network combines the state-action pair calculation value evaluation, builds an experience replay buffer with a capacity of 10,000 to store training samples, and updates sample data every 500 steps. Each training randomly samples 64 samples from the buffer as a training batch, calculates the time difference target value and policy gradient, and uses the Adam optimizer with a learning rate of 0.001 to update the network parameters.

[0114] The obtained initial allocation strategy is added to the pre-set bidirectional adversarial generative network. The generator of the network adopts a 4-layer transposed convolution structure, with a convolution kernel size of 3x3, a stride of 2, and the number of channels is 256, 128, 64 and 32 respectively. Each layer of transposed convolution is connected to a batch normalization layer and a ReLU activation function, and a residual connection is introduced. The discriminator includes a global discriminator and a local discriminator. The global discriminator is a 4-layer convolutional network, and the local discriminator is a 3-layer convolutional network. The generator generates 10 candidate allocation schemes based on randomly sampled 100-dimensional normal distribution latent variables. Each scheme contains the mapping relationship between tasks and execution groups, task priority (integer 1-5) and resource requirement constraints (CPU, memory, bandwidth). The discriminator evaluates the quality score of each candidate scheme based on load balancing, network communication overhead and task dependency, with a score range of 0-1.

[0115] Extract the features of the execution groups and construct an adaptive similarity matrix. For each execution group, extract its load features such as CPU utilization, memory usage, network throughput, topological features such as node degree and clustering coefficient, and historical performance features such as average response time and throughput. Use the Gaussian kernel function to calculate the similarity value between the execution groups, and the bandwidth parameter of the kernel function is adaptively determined by the median method. Construct a similarity matrix based on the calculated similarity values. Perform spectral clustering operations on the candidate solutions with the highest quality scores and the similarity matrix, and set the number of clusters to 1 / 3 of the number of execution groups to obtain the initial clustering results.

[0116] In this embodiment, by combining the improved soft actor critic algorithm with the bidirectional generative adversarial network, it is possible to effectively utilize historical data and prediction information to generate high-quality task allocation solutions, thereby improving the accuracy and efficiency of task allocation. By using spectral clustering and adaptive similarity matrix, the similarity and correlation between the execution groups are fully considered, making task allocation more reasonable, which is conducive to load balancing and improving resource utilization. It has strong adaptability and generalization capabilities, can cope with complex and changeable task environments, and provides a new solution for task scheduling optimization of large-scale distributed systems.

[0117] In an optional embodiment,

[0118] Combined with the dynamic density peak algorithm, the number of clusters is determined and the execution units are divided into different load level groups. For each load level group, a hierarchical spatiotemporal graph attention network analysis is performed. Combined with the hierarchical deep deterministic policy gradient algorithm, the task is decomposed and the decomposition scheme is optimized according to the adaptive crossover mutation operator. The final decomposition results include:

[0119] Based on the pre-acquired initial clustering results and the dynamic density peak algorithm, the local density value and relative distance value of the execution group are calculated. The local density is determined by counting the neighborhood samples through the cutoff kernel function. The influence range of the density peak point is determined based on the density reachability analysis. The cluster center is iteratively updated to divide the execution group into different load level groups.

[0120] Construct a hierarchical spatiotemporal graph attention network to process the load level group, use causal convolution layers with different expansion rates in the time dimension to obtain multi-scale temporal features, and construct a dynamic graph structure modeling execution unit dependency in the space dimension, wherein the spatiotemporal graph attention network includes parallel time channels and space channels, captures temporal dependencies through a self-attention mechanism, aggregates node information through a graph attention layer, constructs a hierarchical policy network to perform task decomposition, splits tasks into subtask sets based on computational complexity and data dependencies through a high-level policy network, assigns subtask priorities through a middle-level policy network, determines the execution scheduling order through a low-level policy network, and uses a deep deterministic policy gradient algorithm to train each layer of policy networks to obtain an initial decomposition plan;

[0121] Based on the adaptive crossover mutation operator, an adaptive evolutionary operator is constructed to optimize the task decomposition scheme. The crossover probability and crossover position are dynamically adjusted according to the fitness value and population diversity. The mutation intensity is adaptively adjusted based on the population convergence degree. The non-dominated sorting method is used to select the optimization scheme that meets the load balancing and execution efficiency constraints to obtain the final decomposition result.

[0122] The cutoff kernel function is a kernel function that maps high-dimensional data to a low-dimensional space. A threshold (cutoff) is set to limit the scope of influence of the kernel function. The density reachability analysis is a technique used in cluster analysis to evaluate the density connectivity between data points. The adaptive crossover mutation operator is a mutation operation method in an evolutionary algorithm. It dynamically adjusts the crossover and mutation operation modes according to the fitness of the current individual. The population convergence degree is an indicator to measure the diversity and convergence state of a population during the evolution process. A population with a lower convergence degree indicates that its solution space has been explored more extensively and can avoid falling into a local optimal solution. A higher convergence degree means that the population has concentrated on a certain solution, which may cause the algorithm to converge too quickly.

[0123] Get the initial clustering results of the execution groups, including the load index data of each execution group. Based on these data, the dynamic density peak algorithm is used to determine the optimal number of clusters, and the execution groups are divided into different load level groups, and the local density value and relative distance value of each execution group are calculated. The local density value is determined by counting the number of neighborhood samples using the cutoff kernel function, and the relative distance value is calculated based on the distance to the high-density point. Then, based on the density reachability analysis, the influence range of the density peak point is determined, and the cluster center is iteratively updated until the clustering result is stable. For example, for a system containing 100 execution groups, it may eventually be divided into three load level groups: low load group, medium load group, and high load group.

[0124] A hierarchical spatiotemporal graph attention network is constructed for each load level group for analysis. The network contains parallel time channels and space channels. In the time channel, causal convolution layers with different dilation rates are used to obtain multi-scale temporal features. For example, for load data sampled once per hour, three layers of causal convolution with dilation rates of 1, 2, and 4 can be used to capture temporal patterns at 1 hour, 2 hours, and 4 hours scales, respectively. In the space channel, a dynamic graph structure is constructed to model the dependencies between executors. The edge weights of the graph can be determined based on the communication frequency or data dependency between executors.

[0125] The self-attention mechanism is used to capture temporal dependencies, and the graph attention layer is used to aggregate node information. The self-attention mechanism can adaptively assign different weights to different time steps, while the graph attention layer aggregates information based on the importance of neighboring nodes. This can effectively extract spatiotemporal features and provide a basis for subsequent task decomposition.

[0126] A hierarchical policy network is constructed to perform task decomposition. The network includes three policy networks: high-level, middle-level, and low-level. The high-level policy network splits the task into a set of subtasks based on computational complexity and data dependencies. For example, for a large data analysis task, it can be split into multiple parallel subtasks based on data distribution. The middle-level policy network assigns subtask priorities and can sort them based on the urgency of the task and resource requirements. The low-level policy network determines the specific execution scheduling order and performs dynamic scheduling considering the current load status of the execution unit.

[0127] The deep deterministic policy gradient algorithm is used to train the policy networks at each layer. This algorithm combines the advantages of Q learning and policy gradient and can effectively process continuous action space. During the training process, by simulating different task decomposition scenarios, the policy network learns the optimal decomposition strategy. After the training is completed, the initial task decomposition solution can be obtained.

[0128] The initial decomposition scheme is optimized based on the adaptive crossover mutation operator. The crossover probability and crossover position are dynamically adjusted according to the fitness value and diversity of the population. For example, when the population diversity is low, the crossover probability is increased to increase exploration. At the same time, the mutation intensity is adaptively adjusted based on the convergence degree of the population. In the early stage, a larger mutation step is used, and in the later stage, the mutation intensity is reduced to achieve fine adjustment. The non-dominated sorting method is used to select the optimization scheme that meets the load balancing and execution efficiency constraints. Load balancing can be measured by the load variance between the executing units, and the execution efficiency is evaluated based on the task completion time. Through multiple iterative optimizations, the task decomposition results that balance the load balancing and execution efficiency are obtained.

[0129] In this embodiment, the optimal number of clusters is adaptively determined through the dynamic density peak clustering algorithm, which avoids the subjectivity of manually setting the number of clusters and improves the accuracy and robustness of load level division. The clustering process based on density reachability analysis can effectively process irregularly distributed data and has stronger adaptability. The layered spatiotemporal graph attention network is used to analyze the load level group, which makes full use of the temporal characteristics of the load data and the spatial correlation between the execution units. The multi-scale causal convolution extracts patterns of different time granularities, and the dynamic graph structure characterizes the dynamic dependency relationship of the execution units, providing a comprehensive feature representation for task decomposition. The layered strategy network is combined with the deep deterministic policy gradient algorithm to realize the adaptive decomposition of tasks. The high, medium and low three-layer strategy network optimizes task allocation from the macro, meso and micro perspectives respectively, improving the flexibility of decomposition. The adaptive crossover mutation operator further optimizes the decomposition scheme and achieves a good balance between load balancing and execution efficiency.

[0130] S3. Based on the final decomposition result, a spatiotemporal heterogeneous graph neural network is constructed and the temporal dependency and spatial constraints of the tasks are integrated to obtain a multidimensional dependency model. The grouping status of the execution group is added to the multidimensional dependency model, and the suboptimal solution is pruned through Monte Carlo search to obtain a task scheduling plan. The Thompson sampling multi-armed bandit algorithm is executed on the task scheduling plan to calculate the importance weights of different scheduling goals and adjust the direction. The adjusted scheduling plan is added to a pre-set hierarchical reinforcement learning framework to plan the global scheduling strategy and implement task allocation by executing scheduling instructions. The task execution status is monitored in real time. If there is a performance anomaly, the samples with performance anomalies are learned through a priority experience playback mechanism. The learning results are added to the graph structure neural network and a knowledge graph is constructed to store the execution experience and optimize the global scheduling strategy. Repeat the optimization to obtain the global optimal scheduling strategy.

[0131] The spatiotemporal heterogeneous graph neural network is a graph neural network model that combines spatiotemporal information and is designed to process heterogeneous data with spatiotemporal structures. The Monte Carlo search is a method based on random sampling and is used to solve the optimal strategy or optimal solution in the decision-making process. The Thompson sampling multi-armed bandit algorithm is a reinforcement learning strategy that is applied to the multi-armed bandit problem. The purpose is to gradually find the best choice among multiple choices. The priority experience replay mechanism is an experience replay technology in reinforcement learning, which aims to improve learning efficiency by giving priority to certain experienced state-action pairs.

[0132] In an optional embodiment,

[0133] Based on the final decomposition result, a spatiotemporal heterogeneous graph neural network is constructed and the temporal dependency and spatial constraint of the task are integrated to obtain a multidimensional dependency model. The grouping state of the execution group is added to the multidimensional dependency. The suboptimal solution is pruned by Monte Carlo search to obtain a task scheduling plan. The Thompson sampling multi-armed bandit algorithm is executed on the task scheduling plan to calculate the importance weights of different scheduling targets and adjust the direction. The adjusted scheduling plan is added to a pre-set hierarchical reinforcement learning framework to plan a global scheduling strategy and implement task allocation by executing scheduling instructions. The task execution status is monitored in real time. If there is a performance anomaly, the samples with performance anomalies are learned through a priority experience playback mechanism. The learning results are added to a graph structure neural network and a knowledge graph is constructed to store execution experience and optimize the global scheduling strategy. Repeated optimization to obtain the global optimal scheduling strategy includes:

[0134] Based on the final decomposition result, a spatiotemporal heterogeneous graph neural network is constructed, and task nodes are divided into computing-intensive nodes, storage-intensive nodes, and communication-intensive nodes, wherein the task nodes include CPU requirements, memory requirements, network bandwidth requirements, expected execution time, task priority, data dependencies, resource constraints, and deadline requirements, and the number of processor cores, memory capacity, network bandwidth, and current load of the execution unit node are constructed as node attributes;

[0135] Extract serial dependencies, parallel dependencies, and semi-dependencies between adjacent tasks to construct timing dependency edges. The serial dependency indicates that tasks must be executed in a strict order, the parallel dependency indicates that tasks can be executed simultaneously, and the semi-dependency indicates that only part of the data dependency needs to be satisfied. Different weights are assigned to each timing dependency edge based on the dependency, where the serial dependency has the highest weight and the parallel dependency has the lowest weight.

[0136] The cosine similarity between the resource requirement vector of the computing task and the resource capacity vector of the execution unit is used to construct a spatial constraint edge, and a high-weight edge, a medium-weight edge, and a low-weight edge are constructed according to the size of the cosine similarity, wherein the similarity corresponds to different levels of weights from high to low, and when the similarity is lower than a preset similarity threshold, no edge connection is constructed;

[0137] Collect the group status information of the execution units within a preset time period to construct dynamic attributes, wherein the group status information includes the CPU utilization, task queue length, and average response time of the computing load group, the memory occupancy, disk IO rate, and cache hit rate of the storage load group, and the network throughput, message loss rate, and end-to-end delay of the communication load group. Based on the Monte Carlo tree search tree method, multiple sampling sequences are expanded from the initial state, and the sampling depth is set according to the total number of tasks. The sequence score is calculated based on the completion time, resource utilization, and load balancing. The sequence with a standardized score result that is higher than a preset score threshold is used as the task scheduling plan;

[0138] A multi-armed bandit model is constructed, and a multi-armed bandit algorithm of Thompson sampling is executed on the task scheduling scheme. The multi-armed bandit algorithm sets minimizing completion time, maximizing resource utilization, and optimizing load balancing as scheduling arms, and calculates the importance weights of different scheduling objectives through Thompson sampling, and adjusts the direction of the task scheduling scheme based on the importance weights;

[0139] Construct a hierarchical reinforcement learning framework including a global policy network and a local execution network, wherein the global policy network and the local execution network are both composed of multiple layers of fully connected layers, input the adjusted task scheduling scheme into the hierarchical reinforcement learning framework, the global policy network receives the system state vector and outputs a global scheduling policy, the local execution network converts the scheduling decision into a scheduling instruction, collects performance data on task execution progress, resource usage, and system response time at preset time intervals, divides anomalies into multiple levels according to the degree to which performance indicators deviate from expectations, and uses abnormal execution sequences as key training samples;

[0140] A three-layer knowledge graph is constructed to store execution experience, wherein the knowledge graph includes a task feature layer, an execution environment layer, and a scheduling strategy layer. The information between layers is connected by associative edges and the edge weights are set based on the association strength. The global scheduling strategy is optimized based on the knowledge graph, and the network parameters are updated through the policy gradient method to make the network output biased towards the scheduling mode that has been historically verified to be effective, until the scheduling performance tends to stabilize and the global optimal scheduling strategy is obtained.

[0141] The sampling depth refers to the exploration depth from the current state to the target state in each iteration in Monte Carlo search or other sampling-based algorithms, and the scheduling arm is a term in the multi-armed bandit problem, which refers to selecting an arm (or action) to pull or execute at a given time or decision stage.

[0142] Task nodes are divided into three categories: compute-intensive, storage-intensive, and communication-intensive. For each task node, attribute information such as CPU demand, memory demand, network bandwidth demand, expected execution time, task priority, data dependency, resource constraints, and deadline requirements are included. For example, a compute-intensive task node may have the following attributes: CPU requirement 4 cores, memory requirement 8GB, bandwidth requirement 100Mbps, expected execution time 30 minutes, high priority, dependent on the completion of tasks A and B, GPU resources required, and deadline 2 hours later. For the execution group node, the number of processor cores, memory capacity, network bandwidth, and current load are constructed as node attributes. For example, the attributes of an execution group node are: 8-core CPU, 32GB memory, 1Gbps bandwidth, and current CPU utilization 60%.

[0143] Extract the dependencies between adjacent tasks and construct sequential dependency edges. Serial dependencies mean that tasks must be executed in strict sequence, and are assigned the highest weight, such as 0.9. Parallel dependencies mean that tasks can be executed simultaneously, and are assigned the lowest weight, such as 0.1. Semi-dependencies mean that only part of the data dependency needs to be met, and are assigned a medium weight, such as 0.5. For example, if task A and task B have serial dependencies, a directed edge with a weight of 0.9 is constructed between A and B.

[0144] Calculate the cosine similarity between the task resource requirement vector and the execution unit resource capacity vector, and construct the spatial constraint edge. Assume that a task requirement vector is [4, 8, 100], and a execution unit capacity vector is [8, 16, 1000], and the calculated similarity is 0.98. Construct edges with different weights according to the similarity, such as 0.9-1.0 for high weight 0.8, 0.7-0.9 for medium weight 0.5, and 0.5-0.7 for low weight 0.2. When the similarity is lower than the preset threshold of 0.5, no edge connection is constructed.

[0145] The dynamic attributes are constructed by collecting the group status information of the execution units within the preset time period. The computing load group includes CPU utilization, task queue length, and average response time. The storage load group includes memory occupancy, disk IO rate, cache hit rate, and the communication load group includes network throughput, message loss rate, and end-to-end delay. For example, the average status of a computing load group within 5 minutes is: CPU utilization 75%, task queue length 10, and average response time 2 seconds.

[0146] Based on the Monte Carlo tree search method, multiple sampling sequences are expanded from the initial state. The sampling depth is set to the total number of tasks, such as 100 tasks, and the sampling depth is 100. The scores of the three indicators of completion time, resource utilization, and load balancing are calculated for each sequence, and then standardized. For example, the original score of a sequence is [90, 85, 80], and the standardized score is [0.95, 0.89, 0.84]. The sequences with scores higher than the preset threshold of 0.8 are selected as candidate task scheduling solutions.

[0147] Construct a multi-armed bandit model and apply Thompson sampling to the task scheduling scheme. Set minimizing completion time, maximizing resource utilization, and optimizing load balancing as three scheduling arms. Calculate the importance weights of the three objectives through Thompson sampling, such as [0.5, 0.3, 0.2]. Adjust the task scheduling scheme based on the weights, and favor the scheme with shorter completion time.

[0148] A hierarchical reinforcement learning framework consisting of a global policy network and a local execution network is constructed. Both networks consist of multiple fully connected layers, such as 100 nodes in the input layer, 50 nodes in the hidden layer, and 10 nodes in the output layer. The global policy network receives the system state vector and outputs the global scheduling strategy, and the local execution network converts the decision into specific scheduling instructions. Performance data such as task execution progress, resource usage, and system response time are collected every 5 minutes. According to the degree of deviation of the indicators from expectations, the anomalies are divided into three levels: mild, moderate, and severe, and the execution sequences with severe anomalies are used as key training samples.

[0149] Construct a three-layer knowledge graph to store execution experience. The task feature layer contains information such as task type and resource requirements, the execution environment layer contains information such as hardware configuration and load status, and the scheduling strategy layer contains information such as scheduling algorithm and parameter configuration. The layers are connected by association edges, and the edge weights represent the strength of the association. For example, the association edge weight between task A and GPU resources is 0.8, indicating a strong correlation. Optimize the global scheduling strategy based on the knowledge graph, update the network parameters through the policy gradient method, and make the network output biased towards the scheduling mode that has been historically verified to be effective. Continue to optimize until the scheduling performance tends to be stable, and obtain the global optimal scheduling strategy.

[0150] In this embodiment, by constructing a spatiotemporal heterogeneous graph neural network and integrating task timing dependencies and spatial constraints, the multi-dimensional dependencies in complex scheduling scenarios are comprehensively characterized, the modeling capabilities of task characteristics and execution environments are improved, and a richer and more accurate information basis is provided for subsequent scheduling decisions. A combination of Monte Carlo tree search and multi-armed bandit algorithm is adopted to quickly locate high-quality scheduling solutions in a vast solution space, and multi-objective balanced optimization is achieved through dynamic weight adjustment, which greatly improves the quality and robustness of the scheduling solutions. A hierarchical reinforcement learning framework and knowledge graph are introduced to achieve coordinated optimization of global strategies and local execution, and the scheduling strategy is continuously optimized through experience accumulation and knowledge transfer, which significantly improves the learning efficiency and adaptability of the scheduling system, and provides a more intelligent and efficient solution for task scheduling in complex dynamic environments.

[0151] In an optional embodiment,

[0152] The multi-armed bandit algorithm of Thompson sampling is executed on the task scheduling scheme, wherein the multi-armed bandit algorithm sets minimizing completion time, maximizing resource utilization, and optimizing load balancing as scheduling arms, and calculates the importance weights of different scheduling objectives through Thompson sampling. Adjusting the direction of the task scheduling scheme based on the importance weights includes:

[0153] Construct a multi-armed bandit model, set minimizing completion time, maximizing resource utilization, and optimizing load balancing as the scheduling arm, create a Beta distribution as a prior distribution for the scheduling arm, and initialize the number of success parameters, number of failure parameters, cumulative reward value, and number of selections of the scheduling arm;

[0154] Performing Thompson sampling on each scheduling arm, extracting sample values ​​from the Beta distribution of the scheduling arm, the sample values ​​representing the expected rate of return of the scheduling arm, normalizing the sample values ​​to obtain importance weights, the importance weights being obtained by dividing the sample values ​​of each scheduling arm by the sum of the sample values ​​of all scheduling arms;

[0155] Calculate a comprehensive score of the scheduling scheme based on the importance weight, the comprehensive score includes a completion time score, a resource utilization score, and a load balance score, the completion time score is obtained by multiplying the importance weight of the first scheduling arm by the inverse of the completion time, the resource utilization score is obtained by multiplying the importance weight of the second scheduling arm by the resource utilization, and the load balance score is obtained by multiplying the importance weight of the third scheduling arm by the inverse of the load variance, and the comprehensive score is obtained by adding the completion time score, the resource utilization score, and the load balance score;

[0156] The task scheduling scheme is adjusted according to the comprehensive score, including reordering task priorities, adjusting resource allocation ratios, and updating task execution order.

[0157] The Beta distribution is a continuous distribution commonly used to represent probability distribution, especially in Bayesian statistics. It has two shape parameters and can describe the probability distribution of event success and failure. It is particularly suitable for describing the prior information of the binomial distribution. In algorithms such as Thompson sampling, the Beta distribution is often used to represent the reward probability distribution of each choice.

[0158] For the scheduling arm that minimizes the completion time, its success number parameter α can be initialized to 1, the failure number parameter β can be initialized to 1, the cumulative reward value can be initialized to 0, and the selection number can be initialized to 0. Similarly, the scheduling arms that maximize resource utilization and optimize load balancing are initialized in the same way.

[0159] Perform Thompson sampling on each scheduling arm. Taking the scheduling arm that minimizes the completion time as an example, randomly extract a sample value, such as 0.7, from its corresponding Beta (1, 1) distribution. Perform the same sampling on the scheduling arms that maximize resource utilization and optimize load balancing, assuming that sample values ​​0.5 and 0.3 are obtained respectively. Then normalize these sample values ​​to obtain the importance weight of each scheduling arm. Specifically, divide 0.7, 0.5, and 0.3 by their sum 1.5, respectively, and obtain the normalized importance weights of 0.47, 0.33, and 0.2, respectively.

[0160] The comprehensive score of the scheduling scheme is calculated based on the importance weight. Assume that the completion time of a scheduling scheme is 100 seconds, the resource utilization is 80%, and the load variance is 0.2. Then its completion time score is 0.47*(1 / 100)=0.0047, the resource utilization score is 0.33*0.8=0.264, and the load balance score is 0.2*(1 / 0.2)=1. The comprehensive score is the sum of these three items, which is 1.2687.

[0161] Adjust the task scheduling plan based on the comprehensive score. For example, you can reorder the tasks from high to low based on the comprehensive score to increase the priority of tasks with high scores. For resource allocation, you can adjust the resource allocation ratio of each task based on the resource utilization score ratio. In addition, you can adjust the execution order of tasks based on the load balance score to make the load more balanced.

[0162] After the adjustment, recalculate the various indicators of the scheduling plan. If the adjusted completion time is shortened, increase the number of successes parameter α of the scheduling arm that minimizes the completion time; otherwise, increase its number of failures parameter β. Similarly, update the parameters of the other two scheduling arms. The cumulative reward value increases the reward for this time, and the number of selections increases by 1.

[0163] The above process is repeated continuously to optimize the task scheduling scheme through multiple iterations. As the number of iterations increases, the Beta distribution of each scheduling arm will gradually converge, thus finding the optimal scheduling strategy.

[0164] In this embodiment, a good balance is achieved between exploration and utilization. It can fully utilize known information and explore potential better solutions, avoid falling into local optimality, adaptively adjust the importance of different scheduling objectives, find a good balance between multiple scheduling objectives, and improve the overall performance of the scheduling scheme. It has strong flexibility and scalability, and can easily add or adjust scheduling objectives. It is suitable for various complex task scheduling scenarios.

[0165] In an optional implementation, the present invention also includes an optional embodiment:

[0166] The following modules work together to achieve efficient scheduling and management of tasks:

[0167] Task management module: responsible for receiving, parsing and scheduling tasks, automatically matching the appropriate execution unit according to the task type and dependency. This module is one of the core parts of the system, and ensures the overall operating efficiency of the system through efficient task parsing and scheduling mechanisms;

[0168] Execution group management module: responsible for the creation, destruction and management of execution groups, and supports dynamic adjustment of the computing power of execution groups according to computing needs. Developers use this module to configure appropriate execution groups to ensure that tasks are executed in the appropriate environment and reduce compatibility issues;

[0169] Resource monitoring module: monitors the operating status and resource usage of the execution unit in real time. This module helps ensure that computing power matches task requirements by monitoring system resources;

[0170] Node switching module: When an execution node fails or has insufficient resources, it automatically switches nodes to ensure continuous execution of tasks. This module can quickly identify node failures and take countermeasures, so that task processing can be smoothly transferred to other available nodes.

[0171] Methods for dynamically allocating and managing execution groups:

[0172] Task type analysis: After receiving a task request, the system will analyze the task type, dependent libraries, Python version, etc. to determine the required operating environment for the task. Through detailed analysis of the task, the system can better match the appropriate execution environment and improve execution efficiency.

[0173] Execution unit matching: According to the task analysis results, the appropriate execution unit is matched from the existing execution units to execute the task; the execution unit needs to be manually pre-allocated by the developer according to the task requirements. Developers can configure the appropriate execution unit through the management tool so that the system can select the optimal execution environment according to the task requirements.

[0174] Dynamic resource adjustment allows developers to manually add executors to the target executor group when they find that resources are insufficient, so as to achieve dynamic management. This method can not only cope with fluctuations in task load, but also improve the overall resource utilization of the system and reduce operating costs.

[0175] Automatic retry and node switching strategy:

[0176] Automatic retry mechanism: When a task fails, the system will automatically trigger the retry mechanism and try to re-execute the task to improve the success rate of the task. Through automatic retry, the system can maintain the continuous execution of tasks in the event of a short-term failure, reducing human intervention.

[0177] Node switching strategy: When an execution node is unavailable or has insufficient resources, the system will automatically migrate tasks to other available nodes to ensure that tasks are executed without interruption and completed correctly. The system selects new nodes based on the node's load, resources, and historical performance to reduce performance impact and ensure efficient execution of tasks. At the same time, the node switching strategy also considers the priority of tasks to ensure that critical tasks are processed first.

[0178] Figure 2 FIG. 1 is a schematic diagram of the structure of the execution unit management system of the adaptive task scheduling according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0179] The first unit is used to collect the load data and operation status data of the execution unit, establish a distributed data processing node to perform federated learning operations, perform differential privacy desensitization and standardization on the collected data to obtain a standardized data set, build a hybrid neural network of a bidirectional long short-term memory network and a graph structure attention network based on the standardized data set, use the bidirectional long short-term memory network to extract timing execution features, calculate the topological associations between the execution units through the graph structure attention network to obtain a state dependency matrix, input the timing execution features and the state dependency matrix into a deep reinforcement learning framework for proximal policy optimization, obtain the capability evaluation result of the execution unit through alternating training of the policy network and the value network, construct a digital mapping model based on the capability evaluation result to encode the state of the execution unit into the latent space of the variational autoencoder, and output the predicted probability distribution of the future state of the execution unit;

[0180] The second unit is used to perform task allocation control based on the predicted probability distribution, calculate the initial allocation strategy in combination with the improved soft actor critic algorithm and add it to a pre-set two-way adversarial generative network, generate multiple candidate allocation schemes and determine the quality of the scheme in combination with the discriminator, select the candidate allocation scheme with the highest quality score as the initial task allocation scheme, perform spectral clustering operation on the initial task allocation scheme and construct an adaptive similarity matrix in combination with the load characteristic network topology of the execution unit, determine the number of clusters in combination with the dynamic density peak algorithm and divide the execution unit into different load level groups, perform hierarchical spatiotemporal graph attention network analysis on each load level group, perform task decomposition in combination with the hierarchical deep deterministic policy gradient algorithm and optimize the decomposition scheme according to the adaptive crossover mutation operator to obtain the final decomposition result;

[0181] The third unit is used to construct a spatiotemporal heterogeneous graph neural network based on the final decomposition result and integrate the temporal dependency and spatial constraint of the task to obtain a multidimensional dependency model, add the grouping state of the execution group to the multidimensional dependency model, prune the suboptimal solution through Monte Carlo search to obtain a task scheduling plan, execute the Thompson sampling multi-armed bandit algorithm on the task scheduling plan to calculate the importance weights of different scheduling goals and adjust the direction, add the adjusted scheduling plan to a pre-set hierarchical reinforcement learning framework to plan the global scheduling strategy and realize task allocation by executing scheduling instructions, monitor the task execution status in real time, if there is a performance anomaly, learn the samples with performance anomaly through the priority experience playback mechanism, add the learning results to the graph structure neural network and construct a knowledge graph to store the execution experience and optimize the global scheduling strategy, and repeat the optimization to obtain the global optimal scheduling strategy.

[0182] According to a third aspect of the embodiments of the present invention,

[0183] An electronic device is provided, comprising:

[0184] processor;

[0185] a memory for storing processor-executable instructions;

[0186] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0187] A fourth aspect of the embodiments of the present invention is:

[0188] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0189] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0190] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An execution unit management method for adaptive task scheduling, characterized in that: include: Collect the load data and operation status data of the execution unit, establish a distributed data processing node to perform federated learning operations, perform differential privacy desensitization and standardization on the collected data to obtain a standardized data set, build a hybrid neural network of a bidirectional long short-term memory network and a graph structure attention network based on the standardized data set, use the bidirectional long short-term memory network to extract timing execution features, calculate the topological associations between the execution units through the graph structure attention network to obtain a state dependency matrix, input the timing execution features and the state dependency matrix into a deep reinforcement learning system for proximal strategy optimization, obtain the capability evaluation result of the execution unit through alternating training of the policy network and the value network, construct a digital mapping model based on the capability evaluation result to encode the state of the execution unit into the latent space of the variational autoencoder, and output the predicted probability distribution of the future state of the execution unit; Based on the predicted probability distribution, task allocation control is performed, and the initial allocation strategy is calculated in combination with the improved soft actor critic algorithm and added to the preset two-way adversarial generative network, multiple candidate allocation schemes are generated and the quality of the scheme is determined in combination with the discriminator, and the candidate allocation scheme with the highest quality score is selected as the initial task allocation scheme, and the spectral clustering operation is performed on the initial task allocation scheme and an adaptive similarity matrix is ​​constructed in combination with the load characteristic network topology of the execution unit, and the number of clusters is determined in combination with the dynamic density peak algorithm and the execution unit is divided into different load level groups, and a hierarchical spatiotemporal graph attention network analysis is performed for each load level group, and the task is decomposed in combination with the hierarchical deep deterministic policy gradient algorithm and the decomposition scheme is optimized according to the adaptive crossover mutation operator to obtain the final decomposition result; Based on the final decomposition result, a spatiotemporal heterogeneous graph neural network is constructed and the temporal dependency and spatial constraints of the tasks are integrated to obtain a multidimensional dependency model. The grouping status of the execution group is added to the multidimensional dependency model, and the suboptimal solution is pruned through Monte Carlo search to obtain a task scheduling plan. The Thompson sampling multi-armed bandit algorithm is executed on the task scheduling plan to calculate the importance weights of different scheduling goals and adjust the direction. The adjusted scheduling plan is added to a pre-set hierarchical reinforcement learning system to plan the global scheduling strategy and realize task allocation by executing scheduling instructions. The task execution status is monitored in real time. If there is a performance anomaly, the samples with performance anomalies are learned through the priority experience playback mechanism. The learning results are added to the graph structure neural network and a knowledge graph is constructed to store the execution experience and optimize the global scheduling strategy. Repeated optimization obtains the global optimal scheduling strategy.

2. The method according to claim 1, characterized in that Collect the load data and operation status data of the execution unit, establish a distributed data processing node to perform federated learning operations, perform differential privacy desensitization and standardization on the collected data to obtain a standardized data set, build a hybrid neural network of a bidirectional long short-term memory network and a graph structure attention network based on the standardized data set, and use the bidirectional long short-term memory network to extract timing execution features, including: Deploy a hierarchical data collection system to collect load data and operating status data of the execution unit, the load data includes CPU usage, memory occupancy, network bandwidth utilization, disk read and write rate and task queue length, and the operating status data includes device temperature, power consumption, vibration frequency and noise; Construct a distributed storage cluster to store the load data and operation status data of the execution group, deploy federated learning service nodes to build a hierarchical aggregation architecture, perform local data processing through the bottom nodes in the hierarchical aggregation architecture, perform regional data aggregation through the middle nodes in the hierarchical aggregation architecture, perform global model updates through the top nodes in the hierarchical aggregation architecture, perform differential privacy protection on the load data and operation status data of the execution group based on the Laplace mechanism, and set a privacy budget according to the sensitivity of the data; Calculate the noise range according to the set privacy budget, limit the noise range of the load data and the operating status data within the ratio range of the corresponding data sensitivity and the privacy budget, sample Laplace noise from the noise range, add the sampled Laplace noise to the original data, verify whether the data after adding the noise meets the differential privacy protection requirements by combining the privacy budget, and perform piecewise normalization processing on the data that meets the differential privacy protection requirements to obtain a standardized data set; A two-layer hybrid neural network model including a time series feature extraction layer and a topological relationship modeling layer is constructed, and the standardized data set is added to the two-layer hybrid neural network model. In the time series feature extraction layer, the standardized data set is segmented by a sliding window to generate a time series data sequence, and forward and reverse processing is performed through a first bidirectional long short-term memory network to obtain a first feature output. The first feature output is subjected to deep feature extraction through a second long short-term memory network to obtain a bidirectional context feature and a time series pooling operation is performed to obtain a time series execution feature.

3. The method according to claim 1, characterized in that The topological association between the execution units is calculated through the graph structure attention network to obtain a state dependency matrix, the temporal execution features and the state dependency matrix are input into the proximal policy optimized deep reinforcement learning system, the ability evaluation result of the execution unit is obtained by alternately training the policy network and the value network, and a digital mapping model is constructed according to the ability evaluation result to encode the state of the execution unit into the latent space of the variational autoencoder, and the predicted probability distribution of the future state of the output execution unit includes: Construct a graph structure attention network in the pre-constructed topological relationship modeling layer, map the execution units into graph structure nodes with attributes, map the physical connection relationships between the execution units into graph structure edges, initialize the weights of the graph structure edges based on the physical distance, construct a multi-head attention mechanism to calculate the attention distribution between nodes, perform normalization on the attention scores to obtain the attention weights, and construct a state dependency matrix based on the attention weights; Input the timing execution feature and the state dependency matrix into a proximal policy optimization system, construct a policy network and a value network in the proximal policy optimization system, wherein both the policy network and the value network adopt a multi-layer fully connected structure, construct an experience replay buffer to store training samples, sample training batches from the experience replay buffer, and calculate the policy gradient loss of the policy network according to the sampled training batches, wherein the policy gradient loss includes an action probability ratio term, an entropy regularization term, and a value loss term; updating the parameters of the policy network according to the policy gradient loss, limiting the difference between the updated policy distribution and the original policy distribution within the confidence domain, calculating the value loss of the value network based on the temporal difference error, wherein the value loss represents the deviation between the predicted value and the actual return, updating the parameters of the value network according to the value loss, using the updated value estimate to calculate the advantage function, and performing policy network optimization and value network optimization alternately until the policy distribution converges and the value estimate is stable; The capability of the execution unit is evaluated based on the trained proximal policy optimization system, the policy gradient loss is determined based on the timing execution characteristics and the state dependency matrix in combination with the policy network in the proximal policy optimization system, the value loss is determined through the value network, and the capability evaluation result is calculated according to the policy gradient loss and the value loss. A digital mapping model of a variational autoencoder structure is constructed based on the capability evaluation result, wherein the variational autoencoder structure includes an encoder and a decoder, and the encoder is used to map the state vector of the execution unit to a low-dimensional latent space, and reparameterized sampling is performed on the latent variables to generate multiple samples, and the predicted probability distribution of the future state of the execution unit is estimated and output.

4. The method according to claim 1, characterized in that Based on the predicted probability distribution, task allocation control is performed, and the initial allocation strategy is calculated in combination with the improved soft actor critic algorithm and added to the preset two-way adversarial generative network, multiple candidate allocation schemes are generated and the quality of the scheme is determined in combination with the discriminator, and the candidate allocation scheme with the highest quality score is selected as the initial task allocation scheme, and the spectral clustering operation is performed on the initial task allocation scheme and the adaptive similarity matrix is ​​constructed in combination with the load characteristics of the execution unit. The network topology includes: An improved soft actor-critic model is constructed, in which the actor network adopts a multi-layer fully connected structure and is regularized by a Dropout layer, and the critic network adopts a dual-branch structure to process state information and action information respectively; Based on the predicted probability distribution, the predicted distribution of processor usage, the predicted distribution of memory occupancy, and the predicted distribution of network bandwidth corresponding to the execution group are determined and added as input state information to the improved soft actor-critic model, and the initial allocation strategy is obtained by the probability distribution of the task allocation action output by the actor network. The critic network combines the state action to evaluate the computational value, constructs an experience replay buffer to store training samples and regularly updates sample data, calculates the temporal difference target value and policy gradient by sampling training batches, and updates network parameters based on the adaptive moment estimation optimizer; The initial allocation strategy is added to a pre-set bidirectional adversarial generative network including a generator and a discriminator. The generator adopts a multi-layer transposed convolution structure and introduces a residual connection mechanism. Each layer of transposed convolution is connected to a batch normalization layer and a nonlinear activation function. The discriminator adopts a multi-scale structure including a global discriminator and a local discriminator. The global discriminator evaluates the overall allocation plan, and the local discriminator evaluates the detailed task allocation. The generator generates multiple candidate allocation plans based on randomly sampled latent variables. The candidate allocation plans include task mapping relationships, task priorities, and resource demand constraint information. The discriminator evaluates the quality score of each candidate allocation plan based on load balancing, network communication overhead, and task dependency. The load characteristics, topology characteristics and historical performance characteristics of the execution groups are extracted, and based on the load characteristics, topology characteristics and historical performance characteristics, the similarity values ​​between the execution groups are calculated in combination with an adaptive Gaussian kernel function. A similarity matrix is ​​constructed based on the similarity values, and a spectral clustering operation is performed on the candidate scheme with the highest quality score and the similarity matrix to obtain an initial clustering result.

5. The method according to claim 1, characterized in that Combined with the dynamic density peak algorithm, the number of clusters is determined and the execution units are divided into different load level groups. For each load level group, a hierarchical spatiotemporal graph attention network analysis is performed. Combined with the hierarchical deep deterministic policy gradient algorithm, the task is decomposed and the decomposition scheme is optimized according to the adaptive crossover mutation operator. The final decomposition results include: Based on the pre-acquired initial clustering results and the dynamic density peak algorithm, the local density value and relative distance value of the execution group are calculated. The local density is determined by counting the neighborhood samples through the cutoff kernel function. The influence range of the density peak point is determined based on the density reachability analysis. The cluster center is iteratively updated to divide the execution group into different load level groups. Construct a hierarchical spatiotemporal graph attention network to process the load level group, use causal convolution layers with different expansion rates in the time dimension to obtain multi-scale temporal features, and construct a dynamic graph structure modeling execution unit dependency in the space dimension, wherein the spatiotemporal graph attention network includes parallel time channels and space channels, captures temporal dependencies through a self-attention mechanism, aggregates node information through a graph attention layer, constructs a hierarchical policy network to perform task decomposition, splits tasks into subtask sets based on computational complexity and data dependencies through a high-level policy network, assigns subtask priorities through a middle-level policy network, determines the execution scheduling order through a low-level policy network, and uses a deep deterministic policy gradient algorithm to train each layer of policy networks to obtain an initial decomposition plan; Based on the adaptive crossover mutation operator, an adaptive evolutionary operator is constructed to optimize the task decomposition scheme. The crossover probability and crossover position are dynamically adjusted according to the fitness value and population diversity. The mutation intensity is adaptively adjusted based on the population convergence degree. The non-dominated sorting method is used to select the optimization scheme that meets the load balancing and execution efficiency constraints to obtain the final decomposition result.

6. The method according to claim 1, characterized in that Based on the final decomposition result, a spatiotemporal heterogeneous graph neural network is constructed and the temporal dependency and spatial constraint of the task are integrated to obtain a multi-dimensional dependency model. The grouping state of the execution group is added to the multi-dimensional dependency. The suboptimal solution is pruned by Monte Carlo search to obtain a task scheduling plan. The Thompson sampling multi-armed bandit algorithm is executed on the task scheduling plan to calculate the importance weights of different scheduling targets and adjust the direction. The adjusted scheduling plan is added to a pre-set hierarchical reinforcement learning system to plan a global scheduling strategy and implement task allocation by executing scheduling instructions. The task execution status is monitored in real time. If there is a performance anomaly, the samples with performance anomalies are learned through a priority experience playback mechanism. The learning results are added to a graph structure neural network and a knowledge graph is constructed to store execution experience and optimize the global scheduling strategy. Repeated optimization to obtain the global optimal scheduling strategy includes: Based on the final decomposition result, a spatiotemporal heterogeneous graph neural network is constructed, and task nodes are divided into computing-intensive nodes, storage-intensive nodes, and communication-intensive nodes, wherein the task nodes include CPU requirements, memory requirements, network bandwidth requirements, expected execution time, task priority, data dependencies, resource constraints, and deadline requirements, and the number of processor cores, memory capacity, network bandwidth, and current load of the execution unit node are constructed as node attributes; Extract serial dependencies, parallel dependencies, and semi-dependencies between adjacent tasks to construct timing dependency edges. The serial dependency indicates that tasks must be executed in a strict order, the parallel dependency indicates that tasks can be executed simultaneously, and the semi-dependency indicates that only part of the data dependency needs to be satisfied. Different weights are assigned to each timing dependency edge based on the dependency, where the serial dependency has the highest weight and the parallel dependency has the lowest weight. The cosine similarity between the resource requirement vector of the computing task and the resource capacity vector of the execution unit is used to construct a spatial constraint edge, and a high-weight edge, a medium-weight edge, and a low-weight edge are constructed according to the size of the cosine similarity, wherein the similarity corresponds to different levels of weights from high to low, and when the similarity is lower than a preset similarity threshold, no edge connection is constructed; Collect the group status information of the execution units within a preset time period to construct dynamic attributes, wherein the group status information includes the CPU utilization, task queue length, and average response time of the computing load group, the memory occupancy, disk IO rate, and cache hit rate of the storage load group, and the network throughput, message loss rate, and end-to-end delay of the communication load group. Based on the Monte Carlo tree search tree method, multiple sampling sequences are expanded from the initial state, and the sampling depth is set according to the total number of tasks. The sequence score is calculated based on the completion time, resource utilization, and load balancing. The sequence with a standardized score result that is higher than a preset score threshold is used as the task scheduling plan; A multi-armed bandit model is constructed, and a multi-armed bandit algorithm of Thompson sampling is executed on the task scheduling scheme. The multi-armed bandit algorithm sets minimizing completion time, maximizing resource utilization, and optimizing load balancing as scheduling arms, and calculates the importance weights of different scheduling objectives through Thompson sampling, and adjusts the direction of the task scheduling scheme based on the importance weights; Construct a hierarchical reinforcement learning system including a global policy network and a local execution network, wherein the global policy network and the local execution network are both composed of multiple layers of fully connected layers, input the adjusted task scheduling plan into the hierarchical reinforcement learning system, the global policy network receives the system state vector and outputs a global scheduling policy, the local execution network converts the scheduling decision into a scheduling instruction, collects performance data on task execution progress, resource usage, and system response time at preset time intervals, divides anomalies into multiple levels according to the degree to which performance indicators deviate from expectations, and uses abnormal execution sequences as key training samples; A three-layer knowledge graph is constructed to store execution experience, wherein the knowledge graph includes a task feature layer, an execution environment layer, and a scheduling strategy layer. The information between layers is connected by associative edges and the edge weights are set based on the association strength. The global scheduling strategy is optimized based on the knowledge graph, and the network parameters are updated through the policy gradient method to make the network output biased towards the scheduling mode that has been historically verified to be effective, until the scheduling performance tends to stabilize and the global optimal scheduling strategy is obtained.

7. The method according to claim 6, characterized in that The multi-armed bandit algorithm of Thompson sampling is executed on the task scheduling scheme, wherein the multi-armed bandit algorithm sets minimizing completion time, maximizing resource utilization, and optimizing load balancing as scheduling arms, and calculates the importance weights of different scheduling objectives through Thompson sampling. Adjusting the direction of the task scheduling scheme based on the importance weights includes: Construct a multi-armed bandit model, set minimizing completion time, maximizing resource utilization, and optimizing load balancing as the scheduling arm, create a Beta distribution as a prior distribution for the scheduling arm, and initialize the number of success parameters, number of failure parameters, cumulative reward value, and number of selections of the scheduling arm; Performing Thompson sampling on each scheduling arm, extracting sample values ​​from the Beta distribution of the scheduling arm, the sample values ​​representing the expected rate of return of the scheduling arm, normalizing the sample values ​​to obtain importance weights, the importance weights being obtained by dividing the sample values ​​of each scheduling arm by the sum of the sample values ​​of all scheduling arms; Calculate a comprehensive score of the scheduling scheme based on the importance weight, the comprehensive score includes a completion time score, a resource utilization score, and a load balance score, the completion time score is obtained by multiplying the importance weight of the first scheduling arm by the inverse of the completion time, the resource utilization score is obtained by multiplying the importance weight of the second scheduling arm by the resource utilization, and the load balance score is obtained by multiplying the importance weight of the third scheduling arm by the inverse of the load variance, and the comprehensive score is obtained by adding the completion time score, the resource utilization score, and the load balance score; The task scheduling scheme is adjusted according to the comprehensive score, including reordering task priorities, adjusting resource allocation ratios, and updating task execution order.

8. An execution unit management system for adaptive task scheduling, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to collect the load data and operation status data of the execution unit, establish a distributed data processing node to perform federated learning operations, perform differential privacy desensitization and standardization on the collected data to obtain a standardized data set, build a hybrid neural network of a bidirectional long short-term memory network and a graph structure attention network based on the standardized data set, use the bidirectional long short-term memory network to extract timing execution features, calculate the topological associations between the execution units through the graph structure attention network to obtain a state dependency matrix, input the timing execution features and the state dependency matrix into a deep reinforcement learning system for proximal strategy optimization, obtain a capability evaluation result of the execution unit through alternating training of the policy network and the value network, construct a digital mapping model based on the capability evaluation result to encode the state of the execution unit into the latent space of the variational autoencoder, and output the predicted probability distribution of the future state of the execution unit; The second unit is used to perform task allocation control based on the predicted probability distribution, calculate the initial allocation strategy in combination with the improved soft actor critic algorithm and add it to a pre-set two-way adversarial generative network, generate multiple candidate allocation schemes and determine the quality of the scheme in combination with the discriminator, select the candidate allocation scheme with the highest quality score as the initial task allocation scheme, perform spectral clustering operation on the initial task allocation scheme and construct an adaptive similarity matrix in combination with the load characteristic network topology of the execution unit, determine the number of clusters in combination with the dynamic density peak algorithm and divide the execution unit into different load level groups, perform hierarchical spatiotemporal graph attention network analysis on each load level group, perform task decomposition in combination with the hierarchical deep deterministic policy gradient algorithm and optimize the decomposition scheme according to the adaptive crossover mutation operator to obtain the final decomposition result; The third unit is used to construct a spatiotemporal heterogeneous graph neural network based on the final decomposition result and integrate the temporal dependency and spatial constraint of the task to obtain a multidimensional dependency model, add the grouping state of the execution group to the multidimensional dependency model, prune the suboptimal solution through Monte Carlo search to obtain a task scheduling plan, execute the Thompson sampling multi-armed bandit algorithm on the task scheduling plan to calculate the importance weights of different scheduling goals and adjust the direction, add the adjusted scheduling plan to a pre-set hierarchical reinforcement learning system to plan the global scheduling strategy and realize task allocation by executing scheduling instructions, monitor the task execution status in real time, if there is a performance anomaly, learn the samples with performance anomaly through the priority experience playback mechanism, add the learning results to the graph structure neural network and construct a knowledge graph to store the execution experience and optimize the global scheduling strategy, and repeat the optimization to obtain the global optimal scheduling strategy.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Artificial intelligence selection and configuration

    CN115413346A

  • Edge calculation method based on AI

    CN119046010A