Artificial Intelligence-Based Dynamic Combination Optimization Scheduling Method for Computing Power Tasks and Resources
The AI-driven hybrid framework optimizes task-resource allocation in heterogeneous data centers by modeling server and task heterogeneity, improving efficiency and reducing energy costs through adaptive scheduling.
Patent Information
- Application Number
- CN202411095281.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-12
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-08-12
AI Technical Summary
The existing data center computing power task scheduling methods fail to truly reflect the actual operating conditions of large-scale data centers and cannot be refined management, resulting in waste of resources and high operating costs.
Adopting the dynamic combination optimization scheduling method of heterogeneous tasks-heterogeneous resources based on artificial intelligence, through refined modeling, building a traditional Markov decision-making model, and efficient scheduling is carried out under the framework of directed acyclic graph neural network-pointer network-soft actor critic reinforcement learning algorithms, and establishing the underlying logic that adapts to the large-scale dynamic computing power task scheduling of data centers.
It realizes efficient matching of computing power tasks and resources in heterogeneous data centers, reduces energy consumption and operation costs, and improves system performance and user experience, which is scalable and generalized.
Smart Images

Figure CN118964027B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and specifically relates to a dynamic combination optimization scheduling method for computing power tasks - resources based on artificial intelligence. Background Art
[0002] With the rapid development of artificial intelligence and large language models, as a computing power infrastructure, the energy consumption of data centers has been continuously increasing, and the electricity cost remains high. Data centers can play a unique energy - using flexibility by scheduling computing power tasks, mainly having the following four aspects of advantages:
[0003] (1) Optimize resource utilization: Data centers usually have a large number of heterogeneous servers and resources, and the utilization rate of these resources is directly related to the economic benefits of data centers. Through effective computing power task scheduling algorithms, the efficient use of resources can be achieved, minimizing resource waste to the greatest extent and reducing operating costs;
[0004] (2) Optimize performance: Data centers usually carry a large number of different types of user requests and computing power tasks. These computing power tasks often have different performance requirements and latency tolerances. By optimizing the computing power task scheduling strategy, resources can be reasonably allocated to ensure the performance indicators of key computing power tasks and improve the overall performance of the system and user satisfaction;
[0005] (3) Reduce energy consumption: The energy consumption of data centers is an important part of their operating costs. Through reasonable computing power task scheduling algorithms, active computing power tasks can be centrally scheduled to some servers, enabling other servers to enter a low - power or shutdown state, thereby reducing energy consumption and environmental impact.
[0006] When optimizing the computing power task scheduling of data centers, the impacts of multiple factors need to be comprehensively considered to formulate appropriate scheduling strategies to achieve the optimization and efficient operation of the system: (1) Computing power task attributes and requirements: The attributes of computing power tasks include computing requirements (CPU), memory requirements (I / O), network bandwidth requirements (NET), disk requirements (DISK), etc. In addition, the latency tolerance of computing power tasks is also an important factor to be considered; (2) Resource heterogeneity: Different servers may have different hardware configurations and performance characteristics; (3) User experience: In addition to energy - related goals such as energy consumption costs, another goal of computing power task scheduling is to optimize the performance and response time of the system to improve user experience and the overall efficiency of the system; (4) Generalization: It can dynamically adjust the allocation of computing power tasks and re - balance when the scale of computing power tasks changes to ensure the efficient operation of the system.
[0007] Existing methods for computing power task scheduling in data centers are limited by model scale and computing power. They usually simplify problems by restricting the solution set scale of computing power tasks or resources, or making assumptions such as homogenizing computing power tasks or resources. However, this simplification fails to truly reflect the actual operating conditions of large-scale data centers and cannot perform targeted fine-grained management. Therefore, there are significant limitations in industrial applications. Summary of the Invention
[0008] In view of the above deficiencies of the prior art, this application provides a dynamic combination optimization scheduling method for computing power tasks - resources based on artificial intelligence.
[0009] In the first aspect, this application proposes a dynamic combination optimization scheduling method for heterogeneous tasks - heterogeneous resources based on artificial intelligence, including the following steps:
[0010] Perform fine-grained modeling on the heterogeneous data center environment to obtain a fine-grained model of the heterogeneous data center environment;
[0011] Construct the fine-grained model of the heterogeneous data center environment into a traditional Markov decision model;
[0012] Optimize the traditional Markov decision model to form a heterogeneous task - heterogeneous resource dynamic combination optimization two-layer Markov decision model adapted to the underlying logic of large-scale dynamic computing power task scheduling in the data center, and perform efficient scheduling under the framework of the directed acyclic graph neural network - pointer network - soft actor-critic reinforcement learning algorithm.
[0013] In some embodiments, the performing fine-grained modeling on the heterogeneous data center environment to obtain a fine-grained model of the heterogeneous data center environment includes:
[0014] Perform fine-grained modeling on the resource occupancy and energy consumption characteristics of different types of server resources configured in the heterogeneous data center according to the energy consumption characteristics corresponding to servers with different hardware configurations:
[0015] Establish a server set:
[0016]
[0017] where represents the th server of the th energy consumption type;
[0018] The high-dimensional heterogeneous resource parameters corresponding to the server are:
[0019]
[0020] where represents the server type mark, , , and respectively represent the current CPU, memory, disk, and network occupancy parameters of the server. , , , and all represent the upper limits of server resource parameters;
[0021] According to the random arrival time, execution duration, heterogeneous resource requests, internal dependency structure, and latency tolerance of different types of computing power tasks, the characteristics of heterogeneous computing power tasks in the data center are modeled as a directed acyclic graph structure. Among them, the arrival trajectory of the computing power task is denoted as , and a computing power task includes: job information and the dependency relationship information between jobs. The dependency relationship information between jobs is represented as the adjacent edge set and the adjacency matrix and , providing basic information for the directed acyclic graph neural network operation, denoted as:
[0022]
[0023]
[0024] Among them, the high-dimensional heterogeneous features of the th computing power task include: cpu occupancy rate , mem occupancy rate , disk occupancy rate , net occupancy rate , execution duration: and latency tolerance as well as .
[0025] In some embodiments, constructing the refined model of the heterogeneous data center environment into a traditional Markov decision model includes a state space construction step:
[0026] The state space construction step includes:
[0027] Using DAGNN to transform the high-dimensional heterogeneous computing power task features into low-dimensional homogeneous computing power task embedding vectors, obtaining three embedding vectors, namely job vector, computing power task vector, and global vector;
[0028] Among them, according to the partial order defined by the computing power task, the jobs are processed, and the attention mechanism is used to aggregate the features, represented by the operator. For the of the layer , output the message vector , which represents the same layer calculated by all sub - jobs directly connected to the job of the job vector : the weighted combination of :
[0029]
[0030] Using the associative operator aggregate the job feature vector of the previous layer and the message vector of this job to generate an updated job vector :
[0031]
[0032] where , and are the input, past state, update state / output of the GRU respectively; it is stipulated that the initial state is 0;
[0033] After processing by layers, use the read - out operator to calculate the root job to generate a computing power task vector :
[0034]
[0035] Generate the global vector of all computing power tasks:
[0036]
[0037]
[0038] where and represent the non - linear transformation of the vector input;
[0039] The state space defined according to the traditional Markov decision process is as follows:
[0040]
[0041] consists of the time in the training cycle, the computing power task trajectory the set of servers and the electricity price ;
[0042] Through the three obtained embedding vectors, the high-dimensional heterogeneous computing power task state vector in the traditional Markov decision model is converted into a low-dimensional homogeneous vector, and the defined state space is compressed to:
[0043]
[0044] .
[0045] In some embodiments, constructing the refined model of the heterogeneous data center environment into a traditional Markov decision model includes an action space construction step:
[0046] The action space construction step includes:
[0047] Based on the embedding vectors compressed by DAGNN, construct an action space adapted to the homogeneous input of the pointer network:
[0048]
[0049]
[0050] Among them, represents the sequence of selectable actions of the action space adapted to the homogeneous input of the pointer network at time t, represents each job 's embedding vector, represents the embedding vector of the computing power task to which this job belongs 's embedding vector, represents the global embedding vector, represents the server information, represents that no job is selected for execution;
[0051] In some embodiments, constructing the refined model of the heterogeneous data center environment into a traditional Markov decision model includes a state transition construction step:
[0052] The state transition construction step includes:
[0053] Given state , the corresponding action of the reinforcement learning agent at time step is:
[0054]
[0055] Among them, represents a sequence composed of multiple job-server pairs formed, represents that the current job is not executed. If the action sequence selected at a certain time step is all composed of If it consists of, it means that no action is selected in this step;
[0056] When the action is determined, interact with the environment to obtain the next state and reward , this interaction process can be represented as a mapping function:
[0057]
[0058]
[0059] Among them, represents the resource occupancy corresponding to the corresponding job.
[0060] In some embodiments, constructing the heterogeneous data center environment refinement model into a traditional Markov decision model includes a reward feedback construction step:
[0061] The reward feedback construction step includes: designing a reward function:
[0062]
[0063] Among them, represents the running cycle of the current time step, and the value of this value is uncertain for each time step; represents the system performance weight coefficient; , and respectively represent the energy consumption characteristic functions of different servers, represents the CPU utilization rate, represents the memory access count, represents the disk read / write rate, represents the network read / write rate.
[0064] In some embodiments, optimizing the traditional Markov decision model to form a heterogeneous task-heterogeneous resource dynamic combination optimization two-layer Markov decision model adapted to the underlying logic of large-scale dynamic computing power task scheduling in the data center, and performing efficient scheduling according to the optimized model in the framework of a directed acyclic graph neural network-pointer network-soft actor-critic reinforcement learning algorithm, including:
[0065] Improve the traditional Markov decision model into a heterogeneous task-heterogeneous resource dynamic combination optimization two-layer Markov decision model to make it adapt to the underlying logic of large-scale dynamic computing power task scheduling in the data center;
[0066] Embed the pointer algorithm into the soft actor-critic algorithm framework to solve the temporal dynamic arrangement problem caused by the dynamic combination optimization of heterogeneous tasks and heterogeneous resources. Use soft policy iteration to maximize the objective, and alternately perform policy evaluation and policy improvement within the maximum entropy framework;
[0067] Finally, form an end-to-end training framework for the directed acyclic graph neural network-pointer network-soft actor-critic reinforcement learning algorithm to achieve real-time online decision-making.
[0068] In some embodiments, improving the traditional Markov decision model into a two-layer Markov decision model for the dynamic combination optimization of heterogeneous tasks and heterogeneous resources to adapt to the underlying logic of large-scale dynamic computing power task scheduling in the data center includes:
[0069] Reconstruct the traditional Markov decision model with N-dimensional actions into a combinatorial optimization Markov decision model containing N one-dimensional action sequences, and define the action value function of the reconstructed combinatorial optimization Markov decision model:
[0070]
[0071] Among them, define ; for different time steps is changing, and the action value values of the two models are equal in some states:
[0072]
[0073] Among them, represents the action value function of the reconstructed dynamic combinatorial optimization Markov decision model, represents the action value function of the traditional Markov decision model.
[0074] In some embodiments, embedding the pointer algorithm into the soft actor-critic algorithm framework to solve the temporal dynamic arrangement problem caused by the dynamic combination optimization of heterogeneous tasks and heterogeneous resources, using soft policy iteration to maximize the objective, and alternately performing policy evaluation and policy improvement within the maximum entropy framework, including pointer network soft policy evaluation, pointer network soft policy improvement, and pointer network soft policy iteration;
[0075] Embed the pointer algorithm into the soft actor-critic algorithm framework. First, use the pointer network encoder to encode to obtain a feature vector, and then use the decoder to gradually construct a solution in an autoregressive manner in combination with the attention calculation method to obtain the conditional probability :
[0076]
[0077] Among them, given the training pair , Represents the sequence of strategies to be trained, which is read into the encoder as input in sequence, and finally encodes to obtain a vector V storing the information of the input sequence. Meanwhile, the encoder obtains the hidden state of each matching pair during the calculation process ;
[0078] The decoder decodes the vector V. The decoder reads in V and outputs the first-layer hidden state , and uses the attention mechanism to calculate the probabilities of each matching pair according to and the hidden states of each matching pair obtained by the encoder . Select the matching pair with the highest probability as the matching pair for the first step. The decoder reads in the hidden layer output of the previous step and the feature vector of the matching pair, and outputs the current hidden state , and calculates the probabilities of each matching pair according to and the calculations of each matching pair. If the matching pair selected in a certain step does not meet the resource upper limit constraint, or when an empty action appears, the pointer network stops outputting.
[0079] Optimize the pointer network parameters using the discrete soft actor-critic algorithm to find a strategy that maximizes the discounted return over time:
[0080]
[0081] where, represents the policy; represents the optimal policy; represents the policy induced trajectory distribution; determines the relative importance of the entropy term with respect to the reward, represented as the temperature parameter; represents the policy at state , and let be the dimension of the input sequence at a certain time step, represents the action sequence, , and at the end of the time step, the reward obtained by the agent based on the pointer network decision;
[0082] The soft policy evaluation of the pointer network includes: Let be the dimension of a certain action sequence, represents the action sequence , and the Bellman operator applied to is convergent:
[0083]
[0084] where:
[0085]
[0086]
[0087] When it converges;
[0088] The improvement of the pointer network soft policy includes: defining the new goal of the policy as:
[0089]
[0090] Define as the optimizer for maximizing the formula. For each there is and ;
[0091] The iteration of the pointer network soft policy includes: starting from any policy or applying the pointer network soft policy evaluation and improvement, the policy sequence converges to there is , .
[0092] In a second aspect, the present application proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0093] Advantages of the present invention:
[0094] 1. For the heterogeneous data center environment with servers having different hardware configurations and energy consumption characteristics, considering the different resource request characteristics and latency tolerance characteristics of different types of computing power tasks, a matching optimization model for heterogeneous computing power tasks and heterogeneous resources is established;
[0095] 2. Realize the heterogeneous task-heterogeneous resource dynamic combination optimization Markov decision-making modeling that adapts to the underlying logic of large-scale computing power task scheduling in heterogeneous data centers;
[0096] 3. Construct an end-to-end algorithm framework of directed acyclic graph neural network-pointer network-soft actor-critic (DAGNN-PointerNetwork-SACD), based on which the problem of representing massive high-dimensional heterogeneous computing power tasks is solved, and the online efficient scheduling of heterogeneous task-heterogeneous resource dynamic combination optimization is realized;
[0097] 4. It has scalability and generalization. The model that converges in the small-scale computing power task scenario can be extended and applied to different large-scale computing power task scenarios. Description of the Drawings
[0098] Figure 1This is the overall flowchart of the present invention.
[0099] Figure 2 This is a schematic diagram of the traditional Markov decision model and the combined optimization Markov decision model.
[0100] Figure 3 This is a schematic diagram of the electricity consumption cost and the situation of overdue tasks in the data center.
[0101] Figure 4 This is a schematic diagram of the power consumption of the servers in the data center. Detailed implementation manners
[0102] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein; on the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be completely conveyed to those skilled in the art.
[0103] The key points of the present invention include:
[0104] (1) Conduct refined modeling for the heterogeneous data center environment facing servers with different hardware configurations and energy consumption characteristics, consider the different resource request characteristics and latency tolerance characteristics of different types of computing power tasks, and establish a matching optimization model for heterogeneous computing power tasks and heterogeneous resources;
[0105] (2) Establish a heterogeneous task-heterogeneous resource dynamic combined optimization Markov model that adapts to the underlying logic of large-scale computing power task scheduling in heterogeneous data centers, and reconstruct the traditional MDP model with N-dimensional actions into a reconstructed combined optimization MDP containing N 1-dimensional action sequences;
[0106] (3) This paper constructs a connection framework of DAGNN (Directed Acyclic Graph Neural Network)-PointerNetwork (Pointer Network). Use DAGNN to solve the problem of representing massive high-dimensional heterogeneous computing power tasks, compress the high-dimensional computing power task features into low-dimensional homogeneous embedding vectors, and form a low-dimensional homogeneous sequence vector that can be input into PointerNetwork together with the heterogeneous server features;
[0107] (4) Propose an efficient algorithm of Pointer Network-Soft Actor-Critic (Ptr-SACD) that adapts to the heterogeneous task-heterogeneous resource matching dynamic combined optimization MDP model, decompose the combined optimization problem in the new MDP as a sequence-to-sequence solution process, and use Pointer Network to solve the temporal dynamic arrangement problem therein: the input of Pointer Network is all the possibilities of combinations at the current time step, and the output is the sorting of the optimal combination determined according to the current network policy;
[0108] (5) Finally, a reinforcement learning algorithm framework for end-to-end solution of DAGNN-Ptr-SACD generalization is formed, which has scalability and generalization. The model converged in the small-scale computing power task scenario can be extended and applied to different large-scale computing power task scenarios.
[0109] Specifically, in the actual computing power task scheduling scenario of heterogeneous data centers, the order of magnitude of computing power tasks is in the tens of millions. Taking n computing power tasks and m servers as an example, the scale of the computing power task scheduling problem is , which is an NP-hard combinatorial optimization problem. From the perspectives of tasks and resources respectively, there are more complex non-linear constraints:
[0110] For tasks, the execution of a task will be decomposed into jobs with different dependency relationships. In the industrial and academic circles, it is usually modeled as a directed acyclic topology (DAG). A reasonable DAG computing power task execution order is closely related to system performance. With the rapid development of the Internet, data centers need to face different types of users (such as government agencies, enterprises, private users, etc.), so they also need to process a variety of DAG computing power tasks (such as form queries, machine learning model training, online shopping, etc.). The arrival of these computing power tasks is also random. Therefore, the DAG computing power tasks in actual working conditions have a high-dimensional heterogeneous topology with different numbers of jobs and different job dependency relationships, and have different requirements for heterogeneous resources such as CPU, MEM, DISK, and NET. Online tasks and offline tasks also have different latency tolerances.
[0111] For resources, in heterogeneous data centers, servers with different hardware configurations have different energy consumption characteristics, and the resource types that affect their energy consumption are also different. According to statistics, the energy consumption of CPU resources only accounts for about 30%-60% of the total energy of the server. Therefore, only considering the single resource of CPU will not only cause other resources to violate the resource upper limit constraint and cause the failure of computing power task execution, but also ignore the heterogeneous energy consumption characteristics and cause cost increase.
[0112] To sum up, during the scheduling process, the scheduler needs to make online real-time matching decisions for heterogeneous computing power tasks and heterogeneous resources in the scenarios of refined heterogeneous resource operation, random arrival of computing power tasks, dynamic change of computing power task scale, and real-time clearing of electricity prices, and is expected to have high generalization ability in different fresh states. The scheduling model and algorithm need to take into account the energy consumption cost and the system performance limitations inherent in the data center.
[0113] Therefore, on the one hand, the present application proposes an artificial intelligence-based dynamic combination optimization scheduling method for computing power tasks-resources, including the following steps:
[0114] S100: Refined modeling of the heterogeneous data center environment to obtain a refined model of the heterogeneous data center environment;
[0115] In some embodiments, the fine-grained modeling of the heterogeneous data center environment to obtain a fine-grained model of the heterogeneous data center environment includes:
[0116] Fine-grained modeling of heterogeneous servers in the data center:
[0117] Fine-grained modeling of the resource occupancy and energy consumption characteristics of different types of server resources configured in the heterogeneous data center according to the energy consumption characteristics corresponding to servers with different hardware configurations;
[0118] Establish a server set:
[0119]
[0120] where represents the th server of the th energy consumption type;
[0121] The high-dimensional heterogeneous resource parameters corresponding to the server are:
[0122]
[0123] wherein, represents the server type tag, , , and respectively represent the current CPU, memory, disk, and network occupancy parameters of the server, , , , and all represent the upper limits of the server resource parameters;
[0124] Modeling of massive heterogeneous DAG computing power tasks in the data center:
[0125] According to the random arrival time, execution duration, heterogeneous resource requests, internal dependency structure, and latency tolerance of different types of computing power tasks, the characteristics of heterogeneous computing power tasks in the data center are modeled as a directed acyclic graph structure. Among them, the arrival trajectory of the computing power task is recorded as , and a said computing power task includes: job information , as well as the dependency relationship information between jobs. The dependency relationship information between jobs is represented as the adjacent edge set and the adjacency matrix , providing basic information for the directed acyclic graph neural network operation, denoted as:
[0126]
[0127]
[0128] Among them, the th computing power task 's th job has high-dimensional heterogeneous features including: CPU occupancy , mem occupancy , disk occupancy , net occupancy , execution duration: and latency tolerance ;
[0129] Specifically, online computing power tasks, according to their characteristics , that is, the latency tolerance is 0, and different offline computing power tasks have different latency tolerances, ranging from more than ten minutes to several hours.
[0130] S200: Construct the refined model of the heterogeneous data center environment into a traditional Markov decision model;
[0131] In some embodiments, the constructing the refined model of the heterogeneous data center environment into a traditional Markov decision model includes a state space construction step:
[0132] The state space construction step includes:
[0133] Use DAGNN to transform the high-dimensional heterogeneous computing power task features into low-dimensional homogeneous computing power task embedding vectors, obtaining three embedding vectors, namely the job vector, the computing power task vector, and the global vector;
[0134] Among them, process the jobs according to the partial order defined by the computing power tasks, and use the attention mechanism to aggregate the features, represented by the operator. For the th layer of , output the message vector , which represents the weighted combination of the job vectors of the same layer computed by for all sub-jobs directly connected to the job :
[0135]
[0136] Specifically, at the 0th layer, , that is, the features of the job itself; the weight coefficient follows the value-key (query-key) design in the traditional attention mechanism, where the job vector of the previous layer As a query:
[0137]
[0138] Among them, and are model parameters. In this embodiment, it is in the form of addition instead of the usual dot product form, so that fewer parameters are involved.
[0139] Using the associative operator to aggregate the job feature vector of the previous layer and the message vector of this job to generate an updated job vector :
[0140]
[0141] Wherein , and are the input, past state, and updated state / output of the GRU respectively; it is stipulated that the initial state is 0. This design is different from most message passing mechanisms that use simple summation or concatenation for aggregation. In this embodiment, the computing power task features are used as the input and the messages are used as the hidden states for periodic updates.
[0142] After layers of processing, the root job calculated using the readout operator generates a computing power task vector :
[0143]
[0144] Generate the global vector of all computing power tasks:
[0145]
[0146]
[0147] Among them, and represent the non-linear transformation of the vector input;
[0148] The state space defined according to the traditional Markov decision process is as follows:
[0149]
[0150] Through the three obtained embedding vectors, the high-dimensional heterogeneous computing power task vector is converted into a low-dimensional The vector compresses the defined state space into:
[0151]
[0152]
[0153] Among them, represents the time in the training cycle; represents the day-ahead market electricity price at the current moment.
[0154] The main innovation of the state space: Utilize the advantage of DAGNN in processing the partial order structure to compress the massive heterogeneous computing power task features into low-dimensional homogeneous computing power task embedding vectors, so as to fully represent the heterogeneous feature information of DAG computing power tasks and the dependencies between their internal operations, and compress the state space to obtain three embedding vectors, namely the partial order aggregation operation embedding vector (abbreviated as the operation vector), the computing power task pooling embedding vector (abbreviated as the computing power task vector), and the non-linear change global embedding vector (abbreviated as the global vector). After that, it can be combined with heterogeneous resource features to form a low-dimensional homogeneous decoupling solution suitable for the pointer network.
[0155] In some embodiments, constructing the refined model of the heterogeneous data center environment into a traditional Markov decision model includes the action space construction step:
[0156] Among them, the combinatorial optimization problem of matching heterogeneous computing power tasks with heterogeneous servers, that is, selecting the optimal job-server pairs from all possible sequences of job-server pairs to form an optimal sequence, is an NP-hard problem with an exponentially increasing scale. And due to the uncertainty of the arrival of computing power tasks and the heterogeneity of computing power tasks, the input of the policy network, that is, the length of the sequence of job-server matching pairs that can be selected at each time step and the length of the optimal sequence of decisions, are variable. Therefore, the decision of the action can be regarded as a special sequence-to-sequence problem.
[0157] Using the pointer network, the encoder is used to encode the input sequence of the combinatorial optimization problem to obtain a feature vector, and then the decoder combines the attention calculation method to gradually construct a solution in an autoregressive manner. Autoregressive means selecting a task-server pair each time and selecting the next action based on the actions that have been selected until a complete solution is constructed.
[0158] The action space construction step includes:
[0159] Based on the embedding vectors compressed by DAGNN, construct an action space that adapts to the homogeneous input of the pointer network:
[0160]
[0161]
[0162] Among them, represents the sequence of optional actions at time t in the action space that adapts to the isomorphic input of the pointer network, represents each job 's embedding vector, represents the job belonging to the computing power task 's embedding vector, represents the global embedding vector, represents the server information, represents that no job is selected for execution;
[0163] Main innovation points of the action space: Due to the uncertainty of the arrival of computing power tasks and the heterogeneity of computing power tasks, the size of the action space / feasible region that needs to be decided at each time step is different, that is, the length of the sequence of job-server matching pairs that can be selected at each time step and the length of the optimal sequence to be decided are variable. As a result, the matching decision problem of heterogeneous computing power tasks and heterogeneous servers is a temporal dynamic permutation problem. The present invention models this NP-hard problem with an exponentially increasing scale.
[0164] In some embodiments, constructing the refined model of the heterogeneous data center environment into a traditional Markov decision model includes a state transition construction step:
[0165] The state transition construction step includes:
[0166] Given state , the corresponding action of the reinforcement learning agent at time step is:
[0167]
[0168] Among them, represents a sequence composed of multiple job-server pairs , represents that the current job is not executed. If the action sequence selected at a certain time step is all composed of , it means that no action is selected in this step;
[0169] When the action is determined, interacts with the environment to obtain the next state and the reward . This interaction process can be represented as a mapping function:
[0170]
[0171]
[0172] Among them, represents the resource occupation corresponding to the corresponding operation, which is restricted and characterized by operation rules and corresponding physical rules, so as to ensure the feasibility of the actions selected in each time step.
[0173] Innovation points of state transition: From to the conversion process has the following randomness: 1. New tasks arrive; 2. An operation is completed.
[0174] In some embodiments, constructing the refined model of the heterogeneous data center environment into a traditional Markov decision model includes a reward feedback construction step:
[0175] At the end of the time step, the agent obtains its reward. The goal of the agent is to optimize the scheduling plan for each task - operation, minimize the energy consumption cost of the data center, and appropriately sacrifice system performance to give full play to time flexibility when high electricity prices occur.
[0176] Therefore, the reward function can be designed in the following aspects:
[0177] 1. The negative cost of server running energy consumption, which has a certain functional relationship with resource occupation;
[0178] 2. Penalty for violating the tolerance constraint of computing power task latency;
[0179] The reward feedback construction step includes: designing a reward function:
[0180]
[0181] Among them, represents the running period of the current time step, and the value of this in each time step is uncertain; represents the system performance weight coefficient; , and respectively represent the energy consumption characteristic functions of different servers, represents the CPU utilization rate, represents the memory access count, represents the disk read - write rate, represents the network read - write rate.
[0182] Innovation points of the reward: It considers both the energy consumption costs of three heterogeneous servers and the performance of the data center.
[0183] S300: Optimize the traditional Markov decision model to form a heterogeneous task-heterogeneous resource dynamic combination optimization double-layer Markov decision model adapted to the underlying logic of large-scale dynamic computing power task scheduling in the data center, and perform efficient scheduling based on the optimized model in the framework of the directed acyclic graph neural network-pointer network-soft actor-critic reinforcement learning algorithm.
[0184] In some embodiments, the step of optimizing the traditional Markov decision model to form a heterogeneous task-heterogeneous resource dynamic combination optimization double-layer Markov decision model adapted to the underlying logic of large-scale dynamic computing power task scheduling in the data center, and performing efficient scheduling based on the optimized model in the framework of the directed acyclic graph neural network-pointer network-soft actor-critic reinforcement learning algorithm includes:
[0185] Use the discrete soft actor-critic algorithm to optimize the pointer network parameters in the traditional Markov decision model to form a heterogeneous task-heterogeneous resource dynamic combination optimization double-layer Markov decision model adapted to the underlying logic of large-scale dynamic computing power task scheduling in the data center;
[0186] Use soft policy iteration to maximize the objective, and alternately perform policy evaluation and policy improvement within the maximum entropy framework;
[0187] Finally, form a directed acyclic graph neural network-pointer network-soft actor-critic reinforcement learning algorithm framework, and realize real-time online decision-making through the optimized model.
[0188] Among them, the established traditional Markov decision model cannot adapt to the dynamic combination optimization of heterogeneous tasks and heterogeneous resources under multiple boundary conditions. Therefore, the Markov decision model is further improved to form an end-to-end training framework, and a double-layer Markov decision model (reconstructed combination optimization Markov decision model) is established to model the problem of energy consumption cost management in large-scale heterogeneous data centers under the scenario of random arrival of massive heterogeneous computing power tasks. Based on this framework, a matching scheme for the dynamic combination optimization of heterogeneous tasks and heterogeneous resources under multiple boundary conditions is proposed. The Ptr algorithm is embedded in the SACD algorithm to solve the dynamic sequence-sequence decision problem caused by the dynamic combination optimization of heterogeneous tasks and heterogeneous resources. Different strategies are fully explored in the environment of electricity price fluctuations, random arrival of massive heterogeneous computing power tasks, and dynamic changes in the states of heterogeneous resources. To maximize the objective, soft policy iteration is used, and policy evaluation and policy improvement need to be alternately performed within the maximum entropy framework. Finally, a DGANN-Ptr-SACD end-to-end training framework is formed to realize real-time online decision-making.
[0189] In some embodiments, the step of improving the traditional Markov decision model into a heterogeneous task-heterogeneous resource dynamic combination optimization double-layer Markov decision model to adapt it to the underlying logic of large-scale dynamic computing power task scheduling in the data center includes:
[0190] Furthermore, due to the heterogeneous and random arrival characteristics of computing power tasks, the action decision of the traditional Markov decision model is a sequential dynamic arrangement problem, which cannot be solved by existing reinforcement learning algorithms. Therefore, both the model and the algorithm are improved. This step first improves the model.
[0191] Consider a certain state under which the action is considered. Regarding an N-dimensional action at a time step t as a sequence of N one-dimensional actions introduces a new MDP composed of states , where the subscript is aligned with the state of the traditional MDP, and the superscript represents the time offset of the reconstructed combinatorial optimization MDP relative to . Therefore, is a tuple that contains the state from the traditional MDP and the history of other states defined by the reconstructed combinatorial optimization MDP (the concatenation of previously selected actions );
[0192] The transition of this reconstructed combinatorial optimization MDP can be defined by two rules: when all one-dimensional actions are executed, i.e., 1 step is calculated in the N-dimensional environment in the last step, a new state is received, a reward is obtained, and the action is reset to ; while in all other previous transitions, the previously selected action is appended to and a reward of 0 is obtained;
[0193] As Figure 2 shown, the two Markov decision models are different forms in the same environment;
[0194] Reconstruct the traditional Markov decision model with N-dimensional actions into a combinatorial optimization Markov decision model that contains N sequences of one-dimensional actions, and define the action value function of the reconstructed combinatorial optimization Markov decision model:
[0195]
[0196] where is defined; for different time steps it varies, and the action value of the two models has equality in some states:
[0197]
[0198] where represents the action value function of the reconstructed dynamic combinatorial optimization Markov decision model. Represents the action value function of the traditional Markov decision model.
[0199] Among them, let be the agent's experience trajectory of length from the offline dataset. For a given time step , and the corresponding action in the trajectory, define the action as composed of multi-dimensional actions included in the one-step decision of the pointer network: Let represent the action dimension vector from the first dimension to the th dimension of the pointer network.
[0200] In some embodiments, embedding the pointer algorithm into the soft actor-critic algorithm framework to solve the temporal dynamic arrangement problem caused by the dynamic combination optimization of heterogeneous tasks and heterogeneous resources, using soft policy iteration to maximize the objective, and alternately performing policy evaluation and policy improvement within the maximum entropy framework, including pointer network soft policy evaluation, pointer network soft policy improvement, and pointer network soft policy iteration;
[0201] Embed the pointer algorithm into the soft actor-critic algorithm framework, use the pointer network encoder to encode to obtain a feature vector, and then use the decoder to gradually construct the solution in an autoregressive manner in combination with the attention calculation method:
[0202]
[0203] Among them, given the training pair , represents the policy sequence to be trained, represents, is read into the encoder as input in sequence, and finally encodes to obtain a vector V storing the information of the input sequence. At the same time, the encoder obtains the hidden layer state of each matching pair during the calculation process;
[0204] Decode the vector V through the decoder. The decoder reads in V and outputs the first-layer hidden layer state , uses the attention mechanism to calculate the probabilities of each matching pair according to and the hidden layer states of each matching pair obtained by the encoder, selects the matching pair with the highest probability as the matching pair for the first step. The decoder reads in the hidden layer output of the previous step and the feature vector of the matching pair, and outputs the current hidden layer state , calculates the probabilities of each matching pair according to and the calculation of each matching pair. If the matching pair selected in a certain step does not meet the resource upper limit constraint, or when an empty action When this happens, the pointer network stops outputting.
[0205] Among them, the specific formula for calculating the probability of each matching pair is:
[0206]
[0207] Among them, 、 and are learnable parameters of the output model; is the number of input vectors ; As a pointer pointing to the input element, the value (query) vector of the attention mechanism represents the matching degree between the rd output and the th input; softmax normalizes the vector (with length n) into an output distribution based on the input dictionary;
[0208] Use the discrete soft actor-critic algorithm to optimize the pointer network parameters and find a policy that maximizes the discounted return over time:
[0209]
[0210] Among them, represents the policy; represents the optimal policy; represents the trajectory distribution induced by the policy ; determines the relative importance of the entropy term relative to the reward, represented as the temperature parameter; represents the entropy of the policy at the state . Let be the dimension of the input sequence at a certain time step, represents the action sequence, , at the end of the time step, the reward obtained by the agent based on the pointer network decision;
[0211] The soft policy evaluation of the pointer network includes: Let be the dimension of a certain action sequence, represents the action sequence , applied to the Bellman operator is convergent:
[0212]
[0213] Among them:
[0214]
[0215]
[0216] When it converges;
[0217] The improvement of the pointer network soft policy includes: defining the new goal of the policy as:
[0218]
[0219] Define as the optimizer for maximizing the formula. For each , then there is and ;
[0220] The iteration of the pointer network soft policy includes: starting from any policy or applying the evaluation and improvement of the pointer network soft policy, the policy sequence converges to , then there is , .
[0221] In a second aspect, the present application proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0222] The present invention has the following three advantages:
[0223] (1) Facing the underlying logic of dynamic computing power task scheduling in large-scale heterogeneous data centers, it more accurately reflects the actual situation of the dynamic combined optimization and matching of heterogeneous computing power tasks and heterogeneous resources in the data center, ensuring both reduced power consumption costs and system performance and user experience, that is, the number of tasks violating the latency tolerance is small, as Figure 3 shown.
[0224] (2) Facing the scenarios of refined heterogeneous resource operation, random arrival of computing power tasks, dynamic change of computing power task scale, and real-time clearing of electricity prices, it can make online real-time decisions on the dynamic combined optimization and matching of large-scale heterogeneous computing power tasks and heterogeneous resources, as Figure 4 shown. The proposed method can make the power consumption of heterogeneous servers adapt to the dynamically changing electricity prices according to their own characteristics: due to the energy consumption characteristics of server 2, the no-load energy consumption is high, so it maintains full-load operation once started; while server 1 has average performance and less energy consumption, so it maintains basic operation; server 3 has low no-load energy consumption but high running energy consumption. Therefore, when the electricity price is high and the number of computing power tasks is small, only a small number of computing power tasks are assigned to server 3 for execution to maintain low power consumption. When the electricity price is low and the number of computing power tasks is large, more computing power tasks are assigned to server 3 for execution. Therefore, the power consumption of server 3 shows an obvious following trend according to the high or low electricity price and the amount of tasks.
[0225] (3) The RL algorithm framework of DAGNN-Ptr-SACD is formed, which has scalability and generalization, improves the efficiency and accuracy of problem-solving, and makes the invention have wider applicability and practicability in actual industrial applications.
[0226] The above are only the preferred embodiments of the present invention. It should be noted that for those skilled in the art, without departing from the technical solution of the present invention, several modified and improved technical solutions should also be regarded as falling within the scope protected by this claim book.
Claims
1. An artificial intelligence-based computing power task-resource dynamic combination optimization scheduling method, characterized in that: Including the following steps: Fine-grained modeling of the heterogeneous data center environment to obtain a fine-grained model of the heterogeneous data center environment, including: Fine-grained modeling of the resource occupancy and energy consumption characteristics of different types of server resources configured in the heterogeneous data center according to the energy consumption characteristics corresponding to servers with different hardware configurations: According to the random arrival time, execution duration, heterogeneous resource requests, internal dependency structure, and latency tolerance of different types of computing power tasks, the characteristics of heterogeneous computing power tasks in the data center are modeled as a directed acyclic graph structure. Among them, the arrival trajectory of the computing power task is denoted as . A computing power task includes: job information , and the dependency relationship information between jobs, which is represented as a set of adjacent edges and an adjacency matrix , providing basic information for the operation of the directed acyclic graph neural network, denoted as: Among them, the high-dimensional heterogeneous features of the th computing power task 's th job include: CPU occupancy , mem occupancy , disk occupancy , net occupancy , execution duration: and latency tolerance ; Constructing the fine-grained model of the heterogeneous data center environment into a traditional Markov decision model; Optimizing the traditional Markov decision model to form a heterogeneous task-heterogeneous resource dynamic combination optimization double-layer Markov decision model adapted to the underlying logic of large-scale dynamic computing power task scheduling in the data center. Scheduling is performed under the framework of the directed acyclic graph neural network-pointer network-soft actor-critic reinforcement learning algorithm, including: Improving the traditional Markov decision model into a heterogeneous task-heterogeneous resource dynamic combination optimization double-layer Markov decision model to adapt it to the underlying logic of large-scale dynamic computing power task scheduling in the data center. Reconstructing the traditional Markov decision model with N-dimensional actions into a combined optimization Markov decision model containing N one-dimensional action sequences, defining the action value function of the reconstructed combined optimization Markov decision model, and constructing a new action value function for this reconstructed combined optimization Markov decision model; Embedding the pointer algorithm into the soft actor-critic algorithm framework to solve the problem of temporal dynamic arrangement caused by heterogeneous task-heterogeneous resource dynamic combination optimization. Using the soft policy to iterate to obtain the maximum objective, and alternately performing policy evaluation and policy improvement within the maximum entropy framework, including pointer network soft policy evaluation, pointer network soft policy improvement, and pointer network soft policy iteration; Embed the pointer algorithm into the soft actor-critic algorithm framework. First, use the pointer network encoder to encode and obtain the feature vector, and then use the decoder combined with the attention calculation method to gradually construct the solution in an autoregressive manner to obtain the conditional probability : Among them, given training pairs , represents the sequence of policies to be trained and is read in sequentially as the input of the encoder. Finally, a vector V storing the information of the input sequence is encoded, and at the same time, the hidden state of each matching pair is obtained during the calculation of the encoder ; Decode the vector V through a decoder, where the decoder reads in V and outputs the first-layer hidden state , and use the attention mechanism to calculate the probabilities of each matching pair according to and the hidden states of each matching pair obtained from the encoder . Select the matching pair with the highest probability as the matching pair for the first step. The decoder reads in the hidden output of the previous step and the feature vector of the matching pair, and outputs the current hidden state . Calculate the probabilities of each matching pair according to and the calculations of each matching pair. When the selected matching pair does not satisfy the resource upper limit constraint, or when there is an empty action , the pointer network stops outputting; Using the discrete soft actor-critic algorithm to optimize the pointer network parameters to find a policy that maximizes the discounted return over time: Finally, forming an end-to-end training framework for the directed acyclic graph neural network-pointer network-soft actor-critic reinforcement learning algorithm to achieve real-time online decision-making.
2. The method according to claim 1, wherein: The step of constructing the fine-grained model of the heterogeneous data center environment into a traditional Markov decision model includes the state space construction step: The state space construction step includes: Using DAGNN to transform the high-dimensional heterogeneous computing power task features into low-dimensional homogeneous computing power task embedding vectors, obtaining three embedding vectors, namely the job vector, the computing power task vector, and the global vector; Among them, the jobs are processed according to the partial order defined by the computing power tasks, and the attention mechanism is used to aggregate the features of , which is represented by the operator. For the of the layer, the output message vector is obtained, which represents the weighted combination of the job vectors of all sub-jobs that are directly adjacent to the job in the same layer calculated by Using an associative operator aggregate the job feature vector of the previous layer and the message vector of the job to generate an updated job vector : Among them , and are the input, past state, and update state / output of the GRU respectively; it is stipulated that the initial state is 0; After layer processing, use the readout operator to perform max pooling calculation on the root job to generate a computing power task vector : Generate a global vector for all computing power tasks: Among them, and represent non-linear transformations of the vector input; The state space defined according to the traditional Markov decision process is as follows: composed of time in the training cycle , computing power task trajectory , server set and electricity price ; By using the three obtained embedding vectors, the high-dimensional heterogeneous computing power task state vector in the traditional Markov decision model is converted into a low-dimensional homogeneous vector, and the defined state space is compressed to: 。 3. The method according to claim 2, characterized in that: The step of constructing the fine-grained model of the heterogeneous data center environment into a traditional Markov decision model includes the action space construction step: The action space construction step includes: Based on the embedding vectors compressed by DAGNN, constructing an action space adapted to the homogeneous input of the pointer network: Among them, represents the selectable action sequence of the action space adapting to the isomorphic input of the pointer network at time t, represents each job 's embedding vector, represents this job belonging to the computing power task 's embedding vector, represents the global embedding vector, represents the server information, represents that no job is selected for execution.
4. The method according to claim 3, wherein: The step of constructing the fine-grained model of the heterogeneous data center environment into a traditional Markov decision model includes the state transition construction step: The state transition construction step includes: Given state , the corresponding action of the reinforcement learning agent at time step is as follows: Among them, represents a sequence composed of multiple job-server pairs ; indicates that the current job is not executed. If the action sequence selected at a certain time step is all composed of , it means that no action is selected in this step; When the action is determined, interact with the environment to obtain the next state and reward , this interaction process can be represented as a mapping function: Among them, represents the resource occupation corresponding to the corresponding job.
5. The method according to claim 4, characterized in that: The step of constructing the fine-grained model of the heterogeneous data center environment into a traditional Markov decision model includes the reward feedback construction step: The reward feedback construction step includes: designing a reward function: Among them, represents the running period of the current time step, and the value of this parameter for each time step is not fixed; represents the system performance weight coefficient; , and respectively represent the energy consumption characteristic functions of different servers, represents the CPU utilization rate, represents the memory access count, represents the disk read / write rate, represents the network read / write rate.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Cloud computing heterogeneous task scheduling and container management method based on artificial intelligence
CN117648174A
Multi-policy intelligent scheduling method and apparatus oriented to heterogeneous computing power
US20240111586A1