Scheduling method and system for intelligently and dynamically optimizing computing power resources
By obtaining multi-dimensional resource data of computing nodes, clustering and constructing graph neural network partition matching model, combined with multi-objective optimization algorithm, the multi-dimensional and heterogeneity problems of resource scheduling in distributed computing systems are solved, and computing efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202510670606.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-15
AI Technical Summary
The existing computing power resource scheduling systems fail to fully consider the multi-dimensional and heterogeneity of distributed computing systems, resulting in low computing task efficiency and insufficient resource utilization.
By obtaining the multi-dimensional computing resource data of the computing nodes, clustering and obtaining abstract descriptions, building a partition matching model based on graph neural network, combining task requirements and multi-objective optimization algorithms, dynamically adjusting node allocation to achieve optimal partitioning and optimal execution of tasks.
It significantly improves the computing efficiency of computing tasks and the resource utilization rate of computing systems, ensures that computing efficiency is always at a high level, and provides efficient computing task execution guarantees.
Smart Images

Figure CN120492170A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computing resource management and scheduling, and more specifically to a scheduling method and system for intelligently and dynamically optimizing computing resources. Background Art
[0002] With the rapid development of new-generation information technologies such as artificial intelligence, cloud computing, and distributed computing, the transmission, storage, and calculation of massive data have placed higher demands on network and computing capabilities.
[0003] Current computing systems have formed a three-tiered architecture of cloud, edge, and end, and are continuously evolving towards cloud-edge-end collaboration. Cloud data centers are typically deployed centrally. While computing resources are abundant, they are limited by the physical distance and network latency between the cloud and end devices, making them incapable of meeting the demands of low-latency, real-time computing tasks. Edge servers, deployed at the edge of the network, typically provide a fixed amount of computing power at a fixed operating frequency. While they meet low-latency requirements, computing power and storage resources are limited, particularly heterogeneous computing resources. Local end devices, as the end nodes of the network, offer strong data security and privacy protection capabilities. However, due to their specific configuration and the need to perform other tasks, their computing power is low and highly volatile, and their heterogeneous computing capabilities are even weaker. To better complete computing tasks, computing resource scheduling technologies are widely used. However, existing computing resource scheduling systems typically focus on a single resource dimension, failing to consider the impact of multiple resource dimensions on computing task efficiency and cluster resource utilization. This inadequately integrates and schedules heterogeneous computing resources from multiple sources, and the accuracy of resource scheduling within the three-tiered architecture of distributed computing systems needs to be improved.
[0004] Therefore, how to consider the multi-dimensionality and heterogeneity of resources in distributed computing systems, allocate and schedule computing tasks, and thus improve the computing efficiency of computing tasks and the resource utilization of computing systems is an issue that technical personnel in this field urgently need to solve. Summary of the Invention
[0005] In view of this, the present invention provides a scheduling method and system for intelligently and dynamically optimizing computing resources to improve the computing efficiency of computing tasks and the resource utilization of the computing system.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] The present invention discloses a scheduling method for intelligent dynamic optimization of computing resources, the specific steps of which are as follows:
[0008] Step 1: Obtain the computing resource data of each computing node in the computing system and cluster them to obtain an abstract description of the computing node;
[0009] Step 2: Collect historical data of different computing tasks of different users and construct a historical task dataset; establish a task prediction model and train it using the historical task dataset; use the trained task prediction model to predict the task requirements of each task;
[0010] Step 3: Based on the abstract description of each computing node and the network connection relationship, a partition matching model based on a graph neural network is constructed, and the optimal partition is matched for each task to be assigned according to the task requirements;
[0011] Step 4: Obtain the current computing load of each computing node and use a multi-objective optimization algorithm to determine the optimal node for each task execution;
[0012] Step 5: Monitor the execution status of each task, determine the tasks that need to be reallocated to the node, and return to step 4.
[0013] Furthermore, the computing resource data of each computing node includes: the CPU single-core computing power of each node, the number of CPU cores, the total memory, the total GPU computing power, the total video memory, the total storage and the total network bandwidth; the values of various types of computing resource data are arranged in order to determine the computing resource vector corresponding to each node.
[0014] Furthermore, the clustering obtains an abstract description of the computing nodes, specifically:
[0015] Step 1.1: Based on the computing resource vector, calculate the Euclidean distance between each node and determine the average distance between each node and all other nodes;
[0016] Step 1.2: Taking the average distance of each node as the radius, count the number of nodes within the radius of each node and sort them, and select the top K nodes with the largest number as the initial cluster centers;
[0017] Step 1.3: Assign each node to the cluster closest to the initial cluster center based on the Euclidean distance between the nodes; determine the mean vector of the cluster based on the mean of each dimension of the computing resource vector of all nodes in each cluster, and use it as the cluster center of each cluster;
[0018] Step 1.4: Calculate the average distance between all nodes in each cluster and the cluster center;
[0019] Step 1.5: Calculate the Euclidean distance between each cluster center and determine the nearest cluster of each cluster;
[0020] Step 1.6: Calculate the equilibrium value between the two clusters based on the Euclidean distance between the cluster center of the cluster and its nearest cluster center, and the average distance between the cluster nodes of the two clusters, and then calculate the average equilibrium value of the current clustering result;
[0021] Step 1.7: Determine whether the average equilibrium value is less than the lower threshold value. If so, K=K-1 and return to step 1.2; otherwise, proceed to step 1.8;
[0022] Step 1.8: Determine whether the average balance value is greater than the upper threshold. If so, K=K+1 and return to step 1.2; otherwise, the computing resource vector of each cluster center is used as an abstract description of all nodes in each cluster.
[0023] Furthermore, the historical data of the computing task includes: user type, task type, data volume, task duration and CPU single-core computing power, as well as the number of CPU cores occupied, GPU computing power occupied, memory occupied, video memory occupied, storage occupied, and network bandwidth occupied at each moment when the task was executed;
[0024] The construction of the historical task dataset includes: uniquely encoding the user type, task type, data volume, task duration and CPU single-core computing power of each computing task as the input features of the sample; according to the set sampling event interval, the number of CPU core occupancy, GPU computing power occupancy, memory occupancy, video memory occupancy, storage occupancy, and network bandwidth occupancy at each moment are converted into corresponding 6 two-dimensional waveform data as the output labels of the sample.
[0025] Furthermore, the task prediction model is a prediction model based on a generative adversarial network, including 6 generative adversarial network sub-models, each of which includes a generator composed of a multi-layer perceptron and a discriminator composed of a multi-layer perceptron;
[0026] The input of the generator of each generative adversarial network sub-model is the input feature, and the output is respectively the predicted two-dimensional waveform data corresponding to the number of CPU core occupancy, GPU computing power occupancy, memory occupancy, video memory occupancy, storage occupancy, and network bandwidth occupancy;
[0027] The input of the discriminator of each generative adversarial network sub-model is the predicted two-dimensional waveform data generated by the respective generator, and the corresponding real two-dimensional waveform data in the output label;
[0028] Each of the generative adversarial network sub-models updates the parameters of the generator and the discriminator in an alternating training manner until the generator and the discriminator reach a balance.
[0029] Furthermore, the task requirements of each task are predicted using the trained task prediction model, including:
[0030] The user type, task type, data volume, task duration, and CPU single-core computing power data of the new task to be run are one-hot encoded and input into the trained generators, and each generator outputs the predicted CPU core occupancy, GPU computing power occupancy, memory occupancy, video memory occupancy, storage occupancy, and network bandwidth occupancy at each moment; the average and maximum values of each occupancy are calculated to obtain the task requirements, including: minimum task latency, CPU single-core computing power, average and maximum CPU core occupancy, average and maximum GPU computing power occupancy, average and maximum memory occupancy, average and maximum video memory occupancy, average and maximum storage occupancy, and average and maximum network bandwidth occupancy.
[0031] Furthermore, the construction of a partition matching model based on a graph neural network includes: taking each of the computing nodes as a node of the graph neural network, adding an edge between any two nodes if there is a network connection between the two nodes, and taking the network delay between the nodes as the length of the edge, taking the abstract description of the computing nodes as the inherent feature of each node, and taking the average computing resource occupancy rate of each of the computing nodes as a dynamic feature, to obtain the partition matching model.
[0032] Furthermore, matching the best partition for each task according to the task requirements includes:
[0033] Step 3.1: Aggregate the inherent features of each node in the graph neural network through graph convolution to obtain the aggregated feature vector of each node;
[0034] Step 3.2: Encode the task requirements of a task to obtain a requirement feature vector;
[0035] Step 3.3: Input the aggregated feature vector and the required feature vector into a trained multi-layer perceptron, output the matching score of each node, and sort to determine a set of nodes with the highest scores;
[0036] Step 3.4: Determine the distance between each node in the set and the task node, determine whether it is less than the minimum distance determined by the task minimum delay in the task requirement, and retain the nodes with a distance less than the minimum distance as the optimal partition of a task;
[0037] Step 3.5: Repeat steps 3.2 to 3.4 until the optimal partitions for all tasks to be assigned are determined.
[0038] Furthermore, the step 4 specifically includes:
[0039] Step 4.1: Sort the currently assigned tasks by number, and randomly select a task execution node from the optimal partition of each task to be assigned, and organize the task execution node numbers into a chromosome in order;
[0040] Step 4.2: Repeat step 4.1 until the number of chromosomes is equal to the set value to obtain the initial population;
[0041] Step 4.3: Obtain the current computing load of each computing node. For the task allocation plan corresponding to a chromosome, simulate the average resource occupancy rate of each node after the task allocation, obtain the dynamic characteristics of the partition matching model, and determine the balance of each node through graph convolution. The sum of the results is used to obtain the fitness value of the current chromosome.
[0042] Step 4.4: Repeat step 4.3 to determine the fitness value of each chromosome;
[0043] Step 4.5: Determine whether the maximum number of iterations has been reached. If so, terminate the iteration and determine the optimal node for each task execution based on the task allocation scheme corresponding to the chromosome with the highest fitness value; otherwise, proceed to step 4.6.
[0044] Step 4.6: Select chromosomes with fitness values greater than the set threshold from the current population as parents, perform chromosome crossover to generate new chromosomes, and add them to the population; remove chromosomes with fitness values less than the set threshold to form a new population; return to step 4.3.
[0045] The present invention also discloses a scheduling system for intelligent dynamic optimization of computing resources, comprising:
[0046] Node description module: obtains the computing resource data of each computing node in the computing system and clusters it to obtain an abstract description of the computing node;
[0047] Task demand prediction module: collects historical data of different computing tasks of different users and constructs a historical task dataset; establishes a task prediction model and trains it using the historical task dataset; uses the trained task prediction model to predict the task demand of each task;
[0048] Partition matching module: Based on the abstract description of each computing node and the network connection relationship, a partition matching model based on a graph neural network is constructed, and the optimal partition is matched for each task according to the task requirements;
[0049] Task allocation module: obtains the computing power load of each computing node and uses a multi-objective optimization algorithm to determine the optimal node for executing each task;
[0050] Monitoring module: monitors the execution status of each task, determines the tasks that need to be reallocated to the node, and uses the task allocation module to reallocate the tasks.
[0051] Through the above technical solutions, it can be seen that compared with the prior art, the present invention discloses a scheduling method and system for intelligent dynamic optimization of computing resources, which fully considers the multi-dimensionality and heterogeneity of resources in distributed computing systems. By obtaining multi-dimensional computing resource data of each computing node and clustering to obtain an abstract description, it can accurately and abstractly characterize the unified characteristics of the computing nodes; using historical task data to build a data set training task prediction model, accurately predict task requirements, and provide a basis for reasonable task allocation; based on the graph neural network, a partition matching model is constructed, combining the abstract description of the computing node, the network connection relationship and the task requirements to match the best partition for the task; and then using a multi-objective optimization algorithm, the node computing power load is comprehensively considered to determine the optimal execution node, which greatly improves the rationality of the allocation of computing tasks. During the task execution process, the node allocation is monitored in real time and dynamically adjusted to ensure that the computing efficiency is always at a high level. The present invention significantly improves the computing efficiency of computing tasks, improves the resource utilization of the computing system, and provides a strong guarantee for the efficient execution of various computing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0053] Figure 1 Schematic diagram of the overall process of an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0055] The embodiment of the present invention discloses a scheduling method for intelligent dynamic optimization of computing resources, such as Figure 1 The specific steps are as follows:
[0056] Step 1: Obtain the computing resource data of each computing node in the computing system and cluster them to obtain an abstract description of the computing node;
[0057] Step 2: Collect historical data of different users' different computing tasks and build a historical task dataset; establish a task prediction model and train it using the historical task dataset; use the trained task prediction model to predict the task requirements of each task;
[0058] Step 3: Based on the abstract description of each computing node and the network connection relationship, a partition matching model based on the graph neural network is constructed, and the optimal partition is matched for each task to be assigned according to the task requirements;
[0059] Step 4: Obtain the current computing load of each computing node and use a multi-objective optimization algorithm to determine the optimal node for each task execution;
[0060] Step 5: Monitor the execution status of each task, determine the tasks that need to be reallocated to the node, and return to step 4.
[0061] In a specific embodiment, the computing resource data of each computing node includes: the CPU single-core computing power of each node, the number of CPU cores, the total memory, the total GPU computing power, the total video memory, the total storage and the total network bandwidth; the values of various types of computing resource data are arranged in order to determine the computing resource vector corresponding to each node.
[0062] In a specific embodiment, clustering obtains an abstract description of the computing nodes, specifically:
[0063] Step 1.1: Based on the computational resource vector, calculate the Euclidean distance between each node and determine the average distance between each node and all other nodes. The formula is: where v i 、v j Represent the computing resource vectors of the i-th and j-th nodes respectively, The average distance between node i and all other nodes, d(v i ,v j ) represents the Euclidean distance between nodes i and j, and V is the total number of nodes;
[0064] Step 1.2: Take the average distance of each node as the radius, count the number of nodes within the radius of each node and sort them, and select the top K nodes with the largest number as the initial cluster centers. The formula is: in, Indicates the number of nodes within the radius of node i; compare[] represents a comparison function, if the internal variable is greater than or equal to 0, the function value is 0, and if it is less than 0, the function value is 1;
[0065] Step 1.3: Based on the Euclidean distance between nodes, assign each node to the cluster closest to the initial cluster center. Based on the mean of each dimension of the computing resource vector of all nodes in each cluster, determine the mean vector of the cluster and use it as the cluster center of each cluster.
[0066] Step 1.4: Calculate the average distance from all nodes in each cluster to the cluster center; the formula is: in, N k and c k They represent the average distance between cluster nodes, the total number of nodes, and the computing resource vector of the cluster center of the k-th cluster respectively; represents the computing resource vector of the i-th node in the k-th cluster;
[0067] Step 1.5: Calculate the Euclidean distance between each cluster center and determine the nearest cluster of each cluster;
[0068] Step 1.6: Calculate the equilibrium value between the two clusters based on the Euclidean distance between the cluster center and the cluster center of the nearest cluster, as well as the average distance between the cluster nodes of the two clusters, and then calculate the average equilibrium value of the current clustering result; the formula is: Among them, B represents the average equilibrium value, c k,near and They represent the computing resource vector of the cluster center of the nearest cluster of the k-th cluster and the average distance of the cluster nodes respectively;
[0069] Step 1.7: Determine whether the average equilibrium value is less than the lower threshold. If so, K = K-1 and return to step 1.2; otherwise, proceed to step 1.8.
[0070] Step 1.8: Determine whether the average balance value is greater than the upper threshold. If so, K = K + 1 and return to step 1.2; otherwise, the computing resource vector of each cluster center is used as an abstract description of all nodes in each cluster.
[0071] By setting the upper and lower thresholds, the number of cluster categories is adjusted. In the actual operation of the computing power system, the upper and lower thresholds are adjusted according to the node scale of the computing power system to keep the number of cluster categories within a reasonable range. This can accurately and abstractly characterize the unified characteristics of the computing nodes without causing an excessive number of cluster categories.
[0072] In a specific embodiment, the historical data of computing tasks includes: user type (individual user, enterprise user, scientific research institution, etc.), task type (scientific computing, image rendering, deep learning training, big data mining, real-time data processing tasks, etc.), data volume (the amount of hard disk usage during execution is determined based on the task execution log), task duration and CPU single-core computing power (determined based on public CPU general computing evaluation data), as well as the number of CPU cores occupied, GPU computing power occupied, memory occupied, video memory occupied, storage occupied, and network bandwidth occupied at each moment during task execution (determined based on the task execution log);
[0073] Constructing a historical task dataset includes: uniquely encoding the user type, task type, data volume, task duration, and CPU single-core computing power of each computing task as the input features of the sample; according to the set sampling event interval, the number of CPU core occupancy, GPU computing power occupancy, memory occupancy, video memory occupancy, storage occupancy, and network bandwidth occupancy at each moment are converted into corresponding 6 two-dimensional waveform data as the output labels of the sample.
[0074] In a specific embodiment, the task prediction model is a prediction model based on a generative adversarial network, including six generative adversarial network sub-models, each of which includes a generator composed of a multi-layer perceptron and a discriminator composed of a multi-layer perceptron;
[0075] The input of the generator of each generative adversarial network sub-model is the input feature, and the output is the predicted two-dimensional waveform data corresponding to the number of CPU core occupancy, GPU computing power occupancy, memory occupancy, video memory occupancy, storage occupancy, and network bandwidth occupancy;
[0076] The input of the discriminator of each generative adversarial network sub-model is the predicted two-dimensional waveform data generated by its own generator and the corresponding real two-dimensional waveform data in the output label;
[0077] Each generative adversarial network sub-model uses an alternating training method to update the parameters of the generator and discriminator until the generator and discriminator reach a balance.
[0078] Specifically, in each training iteration, the generator is first fixed to train the discriminator. A batch of real data samples are randomly sampled from the training set. At the same time, the generator is asked to generate a batch of predicted data samples based on the corresponding user type, task type, data volume, task duration, and CPU single-core computing power. The discriminator distinguishes between real data samples and generated data samples, and updates the discriminator parameters by minimizing the binary cross-entropy loss function. The Adam optimizer is used as the optimizer. Then, the discriminator is fixed and the generator is trained. The generator parameters are updated by minimizing the adversarial loss function, also using the Adam optimizer. Through alternating training, the generator and discriminator reach a balance, that is, the discriminator has difficulty distinguishing between real data and generated data.
[0079] In a specific embodiment, using a trained task prediction model to predict the task requirements of each task includes:
[0080] The user type, task type, data volume, task duration, and CPU single-core computing power data of the new task to be run are one-hot encoded and input into the trained generators. Each generator outputs the predicted CPU core occupancy, GPU computing power occupancy, memory occupancy, video memory occupancy, storage occupancy, and network bandwidth occupancy at each moment; the average and maximum values of each occupancy are calculated to obtain the task requirements, including: minimum task latency, CPU single-core computing power, average and maximum CPU core occupancy, average and maximum GPU computing power occupancy, average and maximum memory occupancy, average and maximum video memory occupancy, average and maximum storage occupancy, and average and maximum network bandwidth occupancy.
[0081] Specifically, after the training of each generative adversarial network sub-model is completed, the user type, task type, data volume, task duration, and CPU single-core computing power data of a new task to be run are correspondingly encoded and input into the trained generator. The generator outputs the predicted number of CPU core occupancy, GPU computing power occupancy, memory occupancy, video memory occupancy, storage occupancy, and network bandwidth occupancy at each moment; and the average and maximum task occupancy are calculated. The task minimum delay determined by the task type of the new task to be run is combined with the task type (different types of tasks have different requirements for network delay and are set with a minimum lower limit. For example, real-time image processing tasks require extremely low network delay, while commercial big data computing and push services allow a certain delay, while scientific computing tasks are not sensitive to network delay and the minimum delay is set according to the characteristics of different tasks), and the CPU single-core computing power to obtain the task requirements, including: CPU single-core computing power, average and maximum CPU core occupancy, average and maximum GPU computing power occupancy, average and maximum memory occupancy, average and maximum video memory occupancy, average and maximum storage occupancy, and average and maximum network bandwidth occupancy.
[0082] In a specific embodiment, a partition matching model based on a graph neural network is constructed, including: treating each computing node as a node of the graph neural network, adding an edge between any two nodes if there is a network connection between the two nodes, using the network delay between the nodes as the length of the edge, using an abstract description of the computing node as the inherent feature of each node, and using the average computing resource utilization of each computing node as a dynamic feature to obtain a partition matching model. The average computing resource utilization is a weighted average of the average CPU core utilization, average GPU computing power utilization, average memory utilization, average video memory utilization, average storage utilization, and average network bandwidth utilization of each node.
[0083] In a specific embodiment, matching the best partition for each task according to task requirements includes:
[0084] Step 3.1: Aggregate the inherent features of each node in the graph neural network through graph convolution to obtain the aggregated feature vector of each node. The formula is:
[0085]
[0086] Among them, h i ′ represents the feature vector after node i is aggregated; h i represents the feature vector determined by node i based on inherent features and dynamic features; h j U represents the feature vector determined by node j based on inherent features and dynamic features; ij represents the set of nodes that have direct edges with node i; σ represents the nonlinear activation function; ε ij represents the aggregation weight, which is the inverse of the edge length between nodes i and j;
[0087] Step 3.2: Encode the task requirements of a task to obtain the requirement feature vector;
[0088] Step 3.3: Input the aggregated feature vector and the required feature vector into the trained multi-layer perceptron, output the matching score of each node, and sort to determine the set of nodes with the highest scores.
[0089] Specifically, the multilayer perceptron is a feedforward neural network consisting of an input layer, multiple hidden layers and an output layer. The layers are connected by weights, and there are no cross-layer connections or same-layer connections between neurons in each layer. The input layer is responsible for receiving external input data and passing it to the hidden layer; the neurons in the hidden layer perform weighted summation on the input data and perform nonlinear transformation in combination with the activation function (Sigmoid function or ReLU function), thereby learning the nonlinear relationship in the data and enhancing the expressive power of the model; the output layer performs calculations based on the output of the hidden layer to generate the final prediction result. In the present invention, the multilayer perceptron is used for a partition matching model based on a graph neural network. When matching the best partition for a task, the aggregated feature vectors of the nodes in the graph neural network and the requirement feature vectors obtained by encoding the task requirements are input into the multilayer perceptron, and these features are deeply processed and analyzed by the multilayer perceptron to output the matching score of each node.
[0090] When training a multi-layer perceptron, we first randomly match aggregate feature vectors with demand feature vectors. We then use each aggregate feature vector and its corresponding demand feature vector as a sample feature. Experts then vote on the matching score for each sample, which serves as the sample label. Through multiple random matching simulations, we generate multiple training samples. The training samples are then used to train and validate the multi-layer perceptron, resulting in a trained multi-layer perceptron.
[0091] Step 3.4: Determine the distance between each node in the set and the task node, and determine whether it is less than the minimum distance determined by the minimum task delay in the task requirement. Nodes with a distance less than the minimum distance are retained as the best partition for a task.
[0092] Step 3.5: Repeat steps 3.2 to 3.4 until the optimal partitions for all tasks to be assigned are determined.
[0093] In step 3, graph convolution aggregates the inherent features of each node to obtain an aggregated feature vector, and encodes the task requirements to obtain a requirement feature vector. Both are then input into a multi-layer perceptron to output a matching score. This process fully exploits the potential connections between nodes and tasks, enabling the model to accurately match node partitions with high GPU computing power and large memory, reducing task waiting time and data transmission overhead, significantly improving training speed. Compared with traditional random assignment or simple node allocation based on a single metric, this significantly shortens task execution time and improves overall computational efficiency.
[0094] In a specific embodiment, step 4 specifically includes:
[0095] Step 4.1: Sort the currently assigned tasks by number, and randomly select a task execution node from the best partition of each task to be assigned. The task execution node numbers are sequentially organized into a chromosome.
[0096] Step 4.2: Repeat step 4.1 until the number of chromosomes is equal to the set value to obtain the initial population;
[0097] Step 4.3: Obtain the current computing load of each computing node. For the task allocation plan corresponding to a chromosome, simulate the average resource occupancy of each node after the task allocation, obtain the dynamic characteristics of the partition matching model, and determine the balance of each node through graph convolution. The sum is obtained to obtain the fitness value of the current chromosome. The formula is:
[0098]
[0099] F represents the fitness value; V represents the total number of nodes; g i represents the average resource occupancy rate of node i, ω is the balancing coefficient; g j represents the average resource occupancy of node j, j∈U ij ; N ij Indicates the total number of nodes that have direct edges with node i;
[0100] Step 4.4: Repeat step 4.3 to determine the fitness value of each chromosome;
[0101] Step 4.5: Determine whether the maximum number of iterations has been reached. If so, terminate the iteration and determine the optimal node for each task execution based on the task allocation scheme corresponding to the chromosome with the highest fitness value; otherwise, proceed to step 4.6.
[0102] Step 4.6: Select chromosomes with fitness values greater than the set threshold from the current population as parents, perform chromosome crossover to generate new chromosomes, and add them to the population; remove chromosomes with fitness values less than the set threshold to form a new population; return to step 4.3.
[0103] Specifically, the multi-objective optimization algorithm utilizes a genetic algorithm. Each potential optimal solution is encoded as a "chromosome." New chromosomes are generated through chromosome inheritance, crossover, and mutation. The chromosomes in the population are iteratively updated, and the optimal solution in the population is determined based on its fitness value. This optimal node for executing each task is then determined based on the optimal solution. The number of iterations can be dynamically adjusted based on the number of tasks to be assigned. When the number of tasks to be assigned is large, the number of iterations is reduced to improve the real-time nature of task assignment. When the number of tasks to be assigned is small, a higher number of iterations can be used to improve the rationality of task assignment. Similarly, the crossover and mutation probabilities can also be dynamically adjusted based on the number of tasks to be assigned. When the number of tasks to be assigned is large, the number of genes making up the chromosome is large. Increasing the crossover and mutation probabilities increases the probability of gene variation, thereby increasing the chances of finding a more appropriate task. When the number of tasks to be assigned is small, the crossover and mutation probabilities are reduced. Through multiple iterations, a more reasonable task assignment is stably found. In step 4, after determining the optimal task matching partition, the multi-objective optimization algorithm is then used to determine the optimal execution node, taking into account node computing load. By first partitioning and then performing more refined node load optimization, we can take into account the adaptability between different tasks and different computing nodes, while also achieving better load balancing of computing nodes, greatly improving the rationality of the distribution of computing tasks.
[0104] The embodiment of the present invention further discloses a scheduling system for intelligently and dynamically optimizing computing resources, including:
[0105] Node description module: obtains the computing resource data of each computing node in the computing system and clusters it to obtain an abstract description of the computing node;
[0106] Task demand prediction module: collects historical data of different computing tasks of different users and constructs a historical task dataset; establishes a task prediction model and trains it using the historical task dataset; uses the trained task prediction model to predict the task demand of each task;
[0107] Partition matching module: Based on the abstract description of each computing node and the network connection relationship, a partition matching model based on graph neural network is constructed, and the optimal partition is matched for each task according to the task requirements;
[0108] Task allocation module: obtains the computing power load of each computing node and uses a multi-objective optimization algorithm to determine the optimal node for executing each task;
[0109] Monitoring module: monitors the execution of each task, determines the tasks that need to be reallocated to the node, and uses the task allocation module to reallocate tasks.
[0110] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0111] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for intelligently and dynamically optimizing the scheduling of computing resources, characterized in that: The specific steps are as follows: Step 1: Obtain the computing resource data of each computing node in the computing system and cluster them to obtain an abstract description of the computing node; Step 2: Collect historical data of different computing tasks of different users and construct a historical task dataset; establish a task prediction model and train it using the historical task dataset; use the trained task prediction model to predict the task requirements of each task; Step 3: Based on the abstract description of each computing node and the network connection relationship, a partition matching model based on a graph neural network is constructed, and the optimal partition is matched for each task to be assigned according to the task requirements; Step 4: Obtain the current computing load of each computing node and use a multi-objective optimization algorithm to determine the optimal node for each task execution; Step 5: Monitor the execution status of each task, determine the tasks that need to be reallocated to the node, and return to step 4.
2. The method for scheduling computing resources for intelligent dynamic optimization according to claim 1, characterized in that: The computing resource data of each computing node includes: the CPU single-core computing power, the number of CPU cores, the total memory, the total GPU computing power, the total video memory, the total storage and the total network bandwidth of each node; the values of various types of computing resource data are arranged in order to determine the computing resource vector corresponding to each node.
3. The method for scheduling intelligent dynamic optimization computing resources according to claim 2, characterized in that: The clustering obtains an abstract description of the computing nodes, specifically: Step 1.1: Based on the computing resource vector, calculate the Euclidean distance between each node and determine the average distance between each node and all other nodes; Step 1.2: Taking the average distance of each node as the radius, count the number of nodes within the radius of each node and sort them, and select the top K nodes with the largest number as the initial cluster centers; Step 1.3: Assign each node to the cluster closest to the initial cluster center based on the Euclidean distance between the nodes; determine the mean vector of the cluster based on the mean of each dimension of the computing resource vector of all nodes in each cluster, and use it as the cluster center of each cluster; Step 1.4: Calculate the average distance between all nodes in each cluster and the cluster center; Step 1.5: Calculate the Euclidean distance between each cluster center and determine the nearest cluster of each cluster; Step 1.6: Calculate the equilibrium value between the two clusters based on the Euclidean distance between the cluster center of the cluster and its nearest cluster center, and the average distance between the cluster nodes of the two clusters, and then calculate the average equilibrium value of the current clustering result; Step 1.7: Determine whether the average equilibrium value is less than the lower threshold value. If so, K=K-1 and return to step 1.2; otherwise, proceed to step 1.8; Step 1.8: Determine whether the average balance value is greater than the upper threshold. If so, K=K+1 and return to step 1.2; otherwise, the computing resource vector of each cluster center is used as an abstract description of all nodes in each cluster.
4. The method for scheduling computing resources for intelligent dynamic optimization according to claim 1, characterized in that: The historical data of the computing task includes: user type, task type, data volume, task duration and CPU single-core computing power, as well as the number of CPU cores occupied, GPU computing power occupied, memory occupied, video memory occupied, storage occupied, and network bandwidth occupied at each moment when the task was executed; The construction of the historical task dataset includes: uniquely encoding the user type, task type, data volume, task duration and CPU single-core computing power of each computing task as the input features of the sample; according to the set sampling event interval, the number of CPU core occupancy, GPU computing power occupancy, memory occupancy, video memory occupancy, storage occupancy, and network bandwidth occupancy at each moment are converted into corresponding 6 two-dimensional waveform data as the output labels of the sample.
5. The method for scheduling computing resources for intelligent dynamic optimization according to claim 4, characterized in that: The task prediction model is a prediction model based on a generative adversarial network, which includes 6 generative adversarial network sub-models, each of which includes a generator composed of a multi-layer perceptron and a discriminator composed of a multi-layer perceptron; The input of the generator of each generative adversarial network sub-model is the input feature, and the output is respectively the predicted two-dimensional waveform data corresponding to the number of CPU core occupancy, GPU computing power occupancy, memory occupancy, video memory occupancy, storage occupancy, and network bandwidth occupancy; The input of the discriminator of each generative adversarial network sub-model is the predicted two-dimensional waveform data generated by the respective generator, and the corresponding real two-dimensional waveform data in the output label; Each of the generative adversarial network sub-models updates the parameters of the generator and the discriminator in an alternating training manner until the generator and the discriminator reach a balance.
6. The method for scheduling computing resources for intelligent dynamic optimization according to claim 5, characterized in that: The task requirements of each task are predicted using the trained task prediction model, including: The user type, task type, data volume, task duration, and CPU single-core computing power data of the new task to be run are one-hot encoded and input into the trained generators, and each generator outputs the predicted CPU core occupancy, GPU computing power occupancy, memory occupancy, video memory occupancy, storage occupancy, and network bandwidth occupancy at each moment; the average and maximum values of each occupancy are calculated to obtain the task requirements, including: minimum task latency, CPU single-core computing power, average and maximum CPU core occupancy, average and maximum GPU computing power occupancy, average and maximum memory occupancy, average and maximum video memory occupancy, average and maximum storage occupancy, and average and maximum network bandwidth occupancy.
7. The method for scheduling computing resources for intelligent dynamic optimization according to claim 1, characterized in that: The construction of a partition matching model based on a graph neural network includes: taking each of the computing nodes as a node of the graph neural network, adding an edge between any two nodes if there is a network connection between the two nodes, and using the network delay between the nodes as the length of the edge, taking the abstract description of the computing nodes as the inherent feature of each node, and taking the average computing resource occupancy rate of each of the computing nodes as a dynamic feature to obtain the partition matching model.
8. The method for scheduling computing resources for intelligent dynamic optimization according to claim 7, characterized in that: Matching the best partition for each task according to the task requirements includes: Step 3.1: Aggregate the inherent features of each node in the graph neural network through graph convolution to obtain the aggregated feature vector of each node; Step 3.2: Encode the task requirements of a task to obtain a requirement feature vector; Step 3.3: Input the aggregated feature vector and the required feature vector into a trained multi-layer perceptron, output the matching score of each node, and sort to determine a set of nodes with the highest scores; Step 3.4: Determine the distance between each node in the set and the task node, determine whether it is less than the minimum distance determined by the task minimum delay in the task requirement, and retain the nodes with a distance less than the minimum distance as the optimal partition of a task; Step 3.5: Repeat steps 3.2 to 3.4 until the optimal partitions for all tasks to be assigned are determined.
9. The method for scheduling computing resources for intelligent dynamic optimization according to claim 7, characterized in that: The step 4 specifically includes: Step 4.1: Sort the currently assigned tasks by number, and randomly select a task execution node from the optimal partition of each task to be assigned, and organize the task execution node numbers into a chromosome in order; Step 4.2: Repeat step 4.1 until the number of chromosomes is equal to the set value to obtain the initial population; Step 4.3: Obtain the current computing load of each computing node. For the task allocation plan corresponding to a chromosome, simulate the average resource occupancy rate of each node after the task allocation, obtain the dynamic characteristics of the partition matching model, and determine the balance of each node through graph convolution. The sum of the results is used to obtain the fitness value of the current chromosome. Step 4.4: Repeat step 4.3 to determine the fitness value of each chromosome; Step 4.5: Determine whether the maximum number of iterations has been reached. If so, terminate the iteration and determine the optimal node for each task execution based on the task allocation scheme corresponding to the chromosome with the highest fitness value; otherwise, proceed to step 4.
6. Step 4.6: Select chromosomes with fitness values greater than the set threshold from the current population as parents, perform chromosome crossover to generate new chromosomes, and add them to the population; remove chromosomes with fitness values less than the set threshold to form a new population; return to step 4.
3.
10. A scheduling system for intelligent dynamic optimization of computing resources, applying the scheduling method for intelligent dynamic optimization of computing resources according to any one of claims 1 to 9, characterized in that: include: Node description module: obtains the computing resource data of each computing node in the computing system and clusters it to obtain an abstract description of the computing node; Task demand prediction module: collects historical data of different computing tasks of different users and constructs a historical task dataset; establishes a task prediction model and trains it using the historical task dataset; uses the trained task prediction model to predict the task demand of each task; Partition matching module: Based on the abstract description of each computing node and the network connection relationship, a partition matching model based on a graph neural network is constructed, and the optimal partition is matched for each task according to the task requirements; Task allocation module: obtains the computing power load of each computing node and uses a multi-objective optimization algorithm to determine the optimal node for executing each task; Monitoring module: monitors the execution status of each task, determines the tasks that need to be reallocated to the node, and uses the task allocation module to reallocate the tasks.
Citation Information
Cited By
Heterogeneous workflow task scheduling method, equipment and medium
CN121585664A