A Container Cluster Resource Scheduling Method and System Based on Deep Reinforcement Learning

By automatically scheduling container cluster resources based on deep reinforcement learning, the problem of manual adjustment in the existing technology is solved, and scheduling efficiency and load balancing are improved.

CN114443249BActive Publication Date: 2025-07-18SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210051579.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-17
Publication Date
2025-07-18
Estimated Expiration
2042-01-17

AI Technical Summary

Technical Problem

Existing container cluster scheduling methods require manual adjustment when the cluster structure or task state changes, resulting in inefficient resource scheduling.

Method used

An agent based on deep reinforcement learning is adopted to generate the action probability distribution of task scheduling by inputting the container cluster node resource usage status and task characteristic values to be scheduled, and update network parameters according to rewards to achieve automatic scheduling strategy adjustment.

Benefits of technology

Automatic scheduling when the cluster node or task state changes is realized, resource scheduling efficiency is improved, task average completion time is shortened, and node resource load balancing is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114443249B_ABST
    Figure CN114443249B_ABST
Patent Text Reader

Abstract

The present invention proposes a container cluster resource scheduling method and system based on deep reinforcement learning, including: S1: Establish a deep reinforcement learning agent. S2: When the allocable resources of a certain container cluster node satisfy the resource request of a certain to-be-scheduled task in the task queue, input the resource usage status of the container cluster node and the characteristic values of the to-be-scheduled task into the deep reinforcement learning agent to obtain the action probability distribution of task scheduling. S3: According to the action probability distribution of task scheduling, schedule the to-be-scheduled task to execute in the container cluster node, calculate the reward, and update the network parameters of the agent according to the reward. S4: Repeat S2 to S3 to train the deep reinforcement learning agent, enabling the agent to continuously learn and adjust. By establishing a deep reinforcement learning agent and continuously learning and training the agent, the agent can automatically generate corresponding scheduling strategies and schedule tasks to the corresponding container cluster nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cluster scheduling, and more specifically, to a container cluster resource scheduling method and system based on deep reinforcement learning. Background Art

[0002] A container is a form of operating system virtualization. In practice, a single container can be used to run all the contents of a small microservice or a large-scale application software process. Different container tasks in a container cluster have different natures. In order to better target the characteristics of container tasks to shorten the average task completion time and achieve load balancing among node resources, it is necessary to design a scheduling strategy method that takes into account the current node resource occupancy and also the impact of future node resource occupancy on resource contention.

[0003] There is an existing scheduling method for a container cluster, which includes obtaining information and load data of all terminals in the container cluster at a preset time interval, where the load data includes the CPU utilization rate, memory occupancy rate, and transmission rate occupied by the containers of the terminals; calculating the comprehensive load rate of the terminals according to the terminal identifiers and the load data; and scheduling the container cluster when the comprehensive load rate of the terminals exceeds 50%. However, the above scheduling method is a method of compiling rules by experts. When the cluster structure or task status changes, it is necessary to manually adjust the scheduling algorithm again, thus wasting a large amount of manpower and material resources and reducing the scheduling efficiency of cluster resources. Summary of the Invention

[0004] In order to overcome the defect of the prior art that when the cluster structure or task status changes, it is necessary to readjust the manual scheduling algorithm, the present invention provides a container cluster resource scheduling method and system based on deep reinforcement learning by modeling multi-dimensional features of container tasks.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows:

[0006] In a first aspect, the present invention proposes a container cluster resource scheduling method based on deep reinforcement learning, including the following steps:

[0007] S1: Establish a deep reinforcement learning agent;

[0008] S2: When the allocable resources of a certain container cluster node satisfy the resource request of a certain to-be-scheduled task in the task queue, input the resource usage status of the container cluster node and the characteristic values of the to-be-scheduled task into the deep reinforcement learning agent to obtain the action probability distribution of task scheduling;

[0009] S3: The container cluster scheduler schedules the tasks to be scheduled to the container cluster nodes according to the action probability distribution of the task scheduling, calculates the rewards, and updates the network parameters of the deep reinforcement learning agent according to the rewards;

[0010] S4: Repeat S2 to S3 to train the deep reinforcement learning agent so that the deep reinforcement learning agent can continuously learn and adjust.

[0011] In this technical solution, a deep reinforcement learning agent is established, and the resource usage status of the container cluster nodes and the characteristic values of the tasks to be scheduled are input into the deep reinforcement learning agent to obtain the action probability distribution of task scheduling. The container cluster scheduler schedules the tasks to be scheduled to the corresponding container cluster nodes according to the action probability distribution, and calculates the rewards. The network parameters of the deep reinforcement learning agent are updated according to the rewards to continuously learn and train, so that the deep reinforcement learning agent can automatically adjust according to changes in the container cluster nodes or task status information. Finally, the deep reinforcement learning agent can generate a corresponding scheduling strategy to instruct the container cluster scheduler to schedule the tasks to be scheduled to the corresponding container cluster nodes.

[0012] Preferably, in S3, the calculation formula of the reward is as follows:

[0013] r j,i (c) = γ p ·priority j (c)-γ i1 imblance1(c)-γ i2 ·imblance2(c)

[0014] Among them, priority j (c) indicates that the deep reinforcement learning agent schedules task T in the time state of scheduling the cth task. j The task priority reward obtained by scheduling to a suitable node for execution, imblance1(c) represents the imbalance degree of resource usage among the nodes in the container cluster, imblance2(c) represents the imbalance degree of resource usage among the nodes in the container cluster, and γ p is the weight coefficient of task priority, γ i1 is the weight coefficient of the penalty term for imbalanced resource usage within the container cluster node, γ i2 The weight coefficient for uneven resource usage among nodes in a container cluster.

[0015] Preferably, the task priority reward priority j The specific formula of (c) is as follows:

[0016]

[0017] In the formula, γ1 is the weight reward coefficient of the task waiting time, γ2 is the priority reward of the task resource request volume, and γ3 is the priority reward of the task delay sensitivity; It represents the task priority reward obtained by scheduling tasks from the perspective of the task resource request volume, and its formula is as follows:

[0018]

[0019] Among them, n is the number of nodes, r CPU is the weight reward coefficient of the CPU resource, r Mem is the weight reward coefficient of the memory resource, r GPU is the weight reward coefficient of the corresponding GPU resource; It represents the maximum available amount of CPU resources of node N i , It represents the maximum available amount of memory resources of node N i , It represents the maximum available amount of GPU resources of node N i ; It represents the CPU resource request volume of the task T to be scheduled j , It represents the memory resource request volume of the task T to be scheduled j , It represents the GPU resource request volume of the task T to be scheduled j ;

[0020] It represents the task priority reward obtained by scheduling tasks from the perspective of the task waiting time, and its formula is as follows:

[0021]

[0022] Among them, Q represents the task queue, w j (c) represents the waiting time of task T j , w k (c) represents the waiting time of task T k ;

[0023] It represents the task priority reward obtained by scheduling tasks from the perspective of the task delay sensitivity. The tasks to be scheduled include inference tasks and training tasks, and its formula is as follows:

[0024]

[0025] Preferably, when considering the dependency relationship between tasks to be scheduled, the task priority reward function priority j(c) is represented as:

[0026]

[0027] where priority z (c) represents the priority reward of task T that depends on task Tj z , and child(j) represents the subtask that depends on task T z .

[0028] Preferably, in S2, when the task priority reward of a certain task is greater than 0, it means that the allocable resources of a certain container cluster node meet the resource requests of the task to be scheduled.

[0029] Preferably, the expressions for the degree of imbalance imblance1(c) of resources within the container cluster node and the degree of imbalance imblance2(c) of resource usage between container cluster nodes are as follows:

[0030]

[0031]

[0032] where U i (c) represents the resource utilization rate of node N at the time state of the c-th scheduling by the deep reinforcement learning agent i , including the CPU resource utilization rate the memory resource utilization rate the GPU resource utilization rate and its expression is as follows:

[0033]

[0034]

[0035]

[0036]

[0037] where represents the available CPU resources of node N in the state of the c-th scheduling i , represents the available memory resources of node N in the state of the c-th scheduling i , represents the available GPU resources of node N in the state of the c-th scheduling i .

[0038] Preferably, in S1, the initial eigenvalue includes the resource request volume, waiting time, dependent task index, and task latency sensitivity of the task to be scheduled, as well as the maximum available resource volume, current available resource volume, and current resource utilization rate of the container cluster node.

[0039] Preferably, the deep reinforcement learning agent includes a policy network. The state information of the task to be scheduled and the container cluster node is input into the policy network, and the policy network outputs the action probability distribution of task scheduling.

[0040] Preferably, the deep reinforcement learning agent further includes a value network, and the value network is used to score the policy network;

[0041] The initial eigenvalue is input into the value network to obtain the score of the policy network. According to the reward and the score of the policy network, the network parameters of the policy network and the value network are updated.

[0042] In a second aspect, the present invention also proposes a container cluster resource scheduling system based on deep reinforcement learning, including:

[0043] A deep reinforcement learning agent, configured to obtain the action probability distribution of task scheduling according to the resource usage status of the container cluster node and the eigenvalue of the task to be scheduled when the available resources of a certain container cluster node satisfy the resource request of a certain task to be scheduled in the task queue;

[0044] A container cluster scheduler, configured to schedule the task to be scheduled to the container cluster node for execution according to the action probability distribution output by the deep reinforcement learning agent;

[0045] An optimization module, configured to calculate the reward, update the network parameters of the deep reinforcement learning agent according to the reward, and continuously optimize the deep reinforcement learning agent.

[0046] Compared with the prior art, the beneficial effect of the technical solution of the present invention is that: by establishing a deep reinforcement learning agent, the present invention inputs the resource usage status of the container cluster node and the eigenvalue of the task to be scheduled into the deep reinforcement learning agent to obtain the action probability distribution of task scheduling. The container cluster scheduler schedules the task to be scheduled to the corresponding container cluster node according to the action probability distribution, calculates the reward, and updates the network parameters of the deep reinforcement learning agent according to the reward to continuously learn and train, so that the deep reinforcement learning agent can automatically adjust when the resource usage status of the container cluster node or the task state information changes. Finally, the deep reinforcement learning agent can generate the corresponding scheduling strategy to instruct the container cluster scheduler to schedule the task to be scheduled to the corresponding container cluster node. Description of the Drawings

[0047] Figure 1It is a flowchart of a container cluster resource scheduling method based on deep reinforcement learning.

[0048] Figure 2 It is a schematic diagram of the container cluster resource scheduling method based on deep reinforcement learning in Embodiment 1.

[0049] Figure 3 It is a flowchart of the task scheduling algorithm in Embodiment 3.

[0050] Figure 4 It is an architecture diagram of a container cluster resource scheduling system based on deep reinforcement learning. Detailed implementation manners

[0051] The attached drawings are only for illustrative purposes and should not be construed as limitations on this patent;

[0052] The technical solutions of the present invention will be further described below in conjunction with the attached drawings and embodiments.

[0053] Embodiment 1

[0054] Please refer to Figure 1 - Figure 2 , this embodiment proposes a container cluster resource scheduling method based on deep reinforcement learning, including the following steps:

[0055] S1: Establish a deep reinforcement learning agent;

[0056] S2: When the allocable resources of a certain container cluster node satisfy the resource request of a certain to-be-scheduled task in the task queue, input the resource usage status of the container cluster node and the feature values of the to-be-scheduled task into the deep reinforcement learning agent to obtain the action probability distribution of task scheduling;

[0057] S3: The container cluster scheduler schedules the to-be-scheduled task to be executed in the container cluster node according to the action probability distribution of task scheduling, calculates the reward, and updates the network parameters of the deep reinforcement learning agent according to the reward;

[0058] S4: Repeat S2 to S3 to train the deep reinforcement learning agent, so that the deep reinforcement learning agent continuously learns and adjusts.

[0059] As Figure 2 shown, Figure 2 It is the system architecture diagram of the container cluster resource scheduling method based on deep reinforcement learning in this embodiment. Assume task T jArriving discretely in the form of containers in a Kubernetes cluster, they enter a task queue with a fixed length of m maintained by the container cluster. These tasks request certain system resources, including CPU, memory, and GPU. The scheduler selects a task from the task queue according to the resource requests of the tasks, the waiting time for scheduling, the dependency index, and the response latency sensitivity, and schedules it to a suitable node according to the optimal scheduling strategy to achieve the shortest average scheduling completion time and resource load balancing among nodes.

[0060] By establishing a deep reinforcement learning agent, the resource usage status of the container cluster nodes and the characteristic values of the tasks to be scheduled are input into the deep reinforcement learning agent to obtain the action probability distribution of task scheduling. The container cluster scheduler schedules the tasks to be scheduled to the corresponding cluster nodes according to the action probability distribution, calculates the reward, and updates the network parameters of the deep reinforcement learning agent according to the reward to continuously learn and train, so that the deep reinforcement learning agent can be adjusted, update the network parameters of the deep reinforcement learning agent to continuously learn and train, and finally the deep reinforcement learning agent can automatically generate the optimal scheduling strategy, instructing the container cluster scheduler to schedule the tasks to be scheduled to the corresponding cluster nodes.

[0061] Embodiment 2

[0062] This embodiment proposes a container cluster resource scheduling method based on deep reinforcement learning, including:

[0063] S1: Establish a deep reinforcement learning agent.

[0064] In this embodiment, the deep reinforcement learning agent interacts with the system environment and needs to take actions, that is, select a certain task from the task queue, instruct the container cluster scheduler to schedule the task to the corresponding container cluster node for execution, and maximize the long-term reward by observing the environment and internal state. More specifically, assuming that the initial maintenance state of the deep reinforcement learning agent is s0, when there is a node with available resources that can meet the resource requirements of a certain task to be scheduled in the task queue, the deep reinforcement learning agent and the scheduling environment will interact. When the c-th scheduling is performed, the deep reinforcement learning agent maintains the state s c , which describes the current information of the system. On this basis, the deep reinforcement learning agent instructs the container cluster scheduler to select the action a c , that is, select a certain task from the task queue and schedule it to the corresponding container cluster node for execution. Then, the deep reinforcement learning agent receives the reward r c generated by the environment after the action, and generates a new state s c+1 . Therefore, the sequence of a series of operations of the deep reinforcement learning agent is:

[0065] s0,a0,r0,s1,a1,r1,…,s c+1 。

[0066] S2: When the allocable resources of a certain container cluster node meet the resource requests of a certain to-be-scheduled task in the task queue, input the resource usage status of the container cluster node and the characteristic values of the to-be-scheduled task into the deep reinforcement learning agent to obtain the action probability distribution of task scheduling.

[0067] In this embodiment, the resource usage status of the container cluster node and the characteristic values of the to-be-scheduled task include the resource request volume, waiting time, dependent task index, and task delay sensitivity of the to-be-scheduled task, as well as the maximum available resources, current allocable resources, and current resource utilization rate of the container cluster node. In the Kubernetes container cluster N = {N1, N2, …, N n}, there are n nodes in total. The computing resources that each node can provide are different. Assume that the computing resources are CPU, memory, and GPU.

[0068] In this embodiment, the deep reinforcement learning agent includes a policy network. In the specific implementation process, based on the A2C algorithm, input the state information of the to-be-scheduled task and the container cluster node into the policy network π(a|s; θ1) of the deep reinforcement learning, and the policy network outputs the action probability distribution π(a|s) of task scheduling.

[0069] As a decision-making unit, the deep reinforcement learning agent is essentially a neural network that takes the state information of the to-be-scheduled task and the container cluster node as input and outputs the action probability distribution π(a|s), where a represents an action and s represents state characteristics. The action of the deep reinforcement learning agent is defined as the deep reinforcement learning agent selects a certain task from the task queue and schedules it to the corresponding node for execution. There are n nodes in the Kubernetes container cluster, and a task queue with a length of m is maintained. Then, the deep reinforcement learning agent has n·m possible actions, and the network output is the probability distribution of these n·m actions.

[0070] The deep reinforcement learning agent includes a state space and an action space. In order to let the deep reinforcement learning agent understand the current scheduling environment in detail, rich features are needed to set the state and use it as the input of the network. In this embodiment, the initial characteristic values set according to the resource usage status of each container cluster node in the current container cluster system and the to-be-scheduled tasks in the task queue constitute the state space of the deep reinforcement learning agent.

[0071] Table 1 State Space of the Deep Reinforcement Learning Agent

[0072] Input Size Maximum available amount of node resources n Current allocatable amount of node resources n Current utilization rate of node resources n Requested amount of each resource for the task 3m Task waiting time m Task dependency index m * m Task delay sensitivity m

[0073] As shown in Table 1, Table 1 represents the state space of the deep reinforcement learning agent, where Table 1 includes the maximum available resource RC of the node, the currently allocable amount of node resources RA(c), the current resource utilization rate U(c) of the node resources, the resource request amount RT of each task to be scheduled in the task queue, the waiting time W, the index I of the dependent task, and the task latency sensitivity L. All these features are represented in the same unit ratio.

[0074] The expression of the action space is a c =(j, i), indicating that the deep reinforcement learning agent selects task T from the task queue j and schedules task T j to node N i for execution.

[0075] S3: The container cluster scheduler schedules the tasks to be scheduled to the container cluster nodes for execution according to the action probability distribution of task scheduling, calculates the reward, and updates the network parameters of the deep reinforcement learning agent according to the reward.

[0076] To effectively train the deep reinforcement learning agent, appropriate rewards need to be set.

[0077] To avoid long-running tasks with low response sensitivity from occupying a large amount of resources and blocking the execution of tasks with less resource occupancy, thereby increasing the average job completion time of the system, a task queue is maintained, and the tasks in the task queue are dynamically sorted by priority. The higher the priority of the task selected by the scheduler, the higher the reward. Among them, the priority sorting mainly starts from the task resource request amount, task waiting time, task dependency, and task latency sensitivity. The purpose is to hope that tasks with higher task priority rewards will be scheduled first. The reward model of the deep reinforcement learning agent is constructed according to the task priority reward, the uneven degree of resource use within the container cluster node, and the uneven degree of resource use between container cluster nodes:

[0078] To avoid tasks with large resource occupancy from affecting the operation of tasks with small resource occupancy, the task priority reward obtained by scheduling tasks from the perspective of task resource request amount is calculated as follows:

[0079]

[0080] where n is the number of nodes, r CPU is the weight reward coefficient of CPU resources, r Mem is the weight reward coefficient of memory resources, r GPU is the weight reward coefficient of GPU corresponding resources; represents the maximum available CPU resource of node N i and Represents the maximum available memory resource of node N i , Represents the maximum available GPU resource of node N i ; Represents the CPU resource request of the task T to be scheduled j , Represents the memory resource request of the task T to be scheduled j , Represents the GPU resource request of the task T to be scheduled j .

[0081] To avoid new tasks constantly coming in, resulting in tasks being in a "starving" state for a long time. Therefore, the higher the priority reward for tasks with longer waiting times, and the task priority reward obtained by scheduling tasks from the perspective of task waiting time The calculation formula is as follows:

[0082]

[0083] Among them, Q represents the task queue, w j (c) represents the waiting time of task T j , w k (c) represents the waiting time of task T k ;

[0084] In the cluster, different types of tasks have different latency sensitivities. For example, in a machine learning task cluster, inference tasks often require faster responses compared to training tasks. Therefore, for these two task types, the task priority reward obtained by scheduling tasks from the perspective of task latency sensitivity The formula is as follows:

[0085]

[0086] In addition, considering the situation where the cluster resources are tight and there are no nodes with available resources to meet the resource requirements of the task to be scheduled, the priority reward of this scheduling task is directly set to negative infinity. Therefore, the task priority reward function priority j (c), the specific formula is as follows:

[0087]

[0088] In the formula, γ1 is the weight reward coefficient of task waiting time, γ2 is the priority reward of task resource request, and γ3 is the priority reward of task latency sensitivity. priority j (c) represents that the deep reinforcement learning agent selects task T in the time state of scheduling the c-th task jThe priority reward obtained by scheduling to a suitable node for execution.

[0089] When the task priority reward of a certain task is negative infinity, there is no container cluster node with available resources to meet the resource requirements of a pending scheduling task. At this time, the task is not scheduled until there is a container cluster node with available resources that can meet the resource requirements of the pending scheduling task, that is, when the priority of the task is greater than 0, the pending scheduling task can be scheduled to a container cluster node.

[0090] In an actual cluster environment, there are often dependencies between tasks. In this case, the most dependent task should be executed first to shorten the completion time of the tasks that depend on it. Therefore, considering the dependencies between tasks, the priority reward can be expressed as:

[0091]

[0092] where priority z (c) represents the priority reward of task T that depends on task Tj, and child(j) represents the subtasks that depend on task T z . z

[0093] In addition, when calculating the scheduling reward obtained by the deep reinforcement learning agent, the imbalance degree of resource usage imblance1(c) within the container cluster node and the imbalance degree of resource usage imblance2(c) between container cluster nodes also need to be considered, and their expressions are as follows:

[0094]

[0095]

[0096] where U i (c) represents the resource utilization rate of node N at the time state of the c-th scheduling by the deep reinforcement learning agent, including the CPU resource utilization rate i , the memory resource utilization rate , and the GPU resource utilization rate.

[0097]

[0098]

[0099]

[0100]

[0101] where Denotes node N in the state of the c-th scheduling i Available CPU resources Denotes node N in the state of the c-th scheduling i Available memory resources Denotes node N in the state of the c-th scheduling i Available GPU resources

[0102] Therefore, the reward function r j,i (c) of the deep reinforcement learning agent is expressed as follows

[0103] r j,i (c) = γ p ·priority j (c) - γ i1 ·imblance1(c) - γ i2 ·imblance2(c)

[0104] Where γ p Is the weight coefficient of the task priority, γ i1 Is the weight coefficient of the penalty term for uneven resource usage within the node, γ i2 Is the weight coefficient of the uneven resource usage between nodes

[0105] When the task priority reward of a certain task is greater than 0, it means that the allocable resources of a certain container cluster node meet the resource requests of the to-be-scheduled task

[0106] In this embodiment, the deep reinforcement learning agent further includes a value network ν(s; θ2), and the value network ν(s; θ2) is used to score the policy network π(a|s; θ1) to help the policy network π(a|s; θ1) improve. In the specific implementation process, the state information s c And s c+1 Are respectively input into the value network ν(s; θ2) to obtain the corresponding scores And According to the scores Score And the reward r c , the time difference method (TD method) is used to train the deep reinforcement learning agent to update the network parameters of the policy network and the value network, and the specific expression formula is as follows

[0107]

[0108]

[0109]

[0110]

[0111] Among them, γ is the future reward weight factor, is the TD target, δ c is the TD error, θ′1 represents the updated policy network parameters, and θ′2 represents the updated value network parameters.

[0112] S4: Repeat S2 to S3 to train the deep reinforcement learning agent, enabling the deep reinforcement learning agent to continuously learn and adjust.

[0113] The goal of optimizing the deep reinforcement learning agent is to maximize the reward function r j,i (c) of the total reward model, so that the deep reinforcement learning agent generates an optimal scheduling policy, and based on the optimal scheduling policy, schedules the tasks to be scheduled to the corresponding container cluster nodes to shorten the average completion time of task scheduling and achieve load balancing among node resources.

[0114] The algorithm pseudocode of the container cluster resource scheduling method based on deep reinforcement learning is as follows:

[0115] Input:

[0116] Resource request volume RP of the task to be scheduled

[0117] Waiting time W of the task to be scheduled

[0118] Index I of the tasks on which the task to be scheduled depends

[0119] Task latency sensitivity L of the task to be scheduled

[0120] Maximum available resource RC of the container cluster node

[0121] Current available resource RA(c) of the container cluster node

[0122] Current resource utilization U(c) of the container cluster node

[0123] Reward function weight ω

[0124] Future reward weight factor γ

[0125] Output:

[0126] Task T to be scheduled j , node N for task scheduling i

[0127] The first step: Calculate the priority reward of the task to be scheduled

[0128] Step 2: Determine whether there is a task priority reward greater than 0. If there is a task priority reward greater than 0, proceed to the next step; otherwise, wait until the condition is met.

[0129] Step 3: Set the characteristic values of each status information and input them into the deep reinforcement learning agent.

[0130] Form the status feature s with the status information such as RC, RA(c), U(c) of the container cluster nodes and RP, W, I, L of the tasks to be scheduled in the c-th scheduling status C Input the policy network π(a|s; θ1)

[0131] Step 4: Make a scheduling decision according to the action probability distribution π(a|s) output by the policy network, and schedule the specified task to be executed on the specified node.

[0132] Select the task T to be scheduled according to the action probability distribution output by the policy network j

[0133] Obtain the selected scheduling node N according to the action probability distribution output by the policy network i

[0134] On the node N i Schedule the task T j

[0135] Step 5: Calculate the reward

[0136] Calculate the reward r according to the formula j,i (c)

[0137] Step 6: Update the network parameters

[0138] Input the status information feature s c and s c+1 into the value network respectively to obtain: and

[0139] Calculate the TD target and the TD error δ c

[0140] Update the value network parameter θ2

[0141] Update the policy network parameter θ1.

[0142] Example 4

[0143] Please refer to Figure 4 , this example proposes a container cluster resource scheduling system based on deep reinforcement learning, including a deep reinforcement learning agent, a container cluster scheduler, and an optimization module.

[0144] In the specific implementation process, when the allocable resources of a certain container cluster node meet the resource requests of a certain to-be-scheduled task in the task queue, the resource usage status of the container cluster node and the characteristic values of the to-be-scheduled task are input into the deep reinforcement learning agent. The policy network of the deep reinforcement learning agent outputs the action probability distribution of task scheduling. The container cluster scheduler schedules the to-be-scheduled task into the corresponding container cluster node according to the action probability distribution. The optimization module calculates the reward and updates the network parameters of the deep reinforcement learning agent according to the reward, continuously optimizing the deep reinforcement learning agent, so that the deep reinforcement learning agent can automatically adjust when the container cluster node or task status information changes. Finally, the deep reinforcement learning agent can generate the corresponding scheduling policy, instructing the container cluster scheduler to schedule the to-be-scheduled task to the corresponding container cluster node.

[0145] The terms used to describe the positional relationship in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent.

[0146] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A container cluster resource scheduling method based on deep reinforcement learning, characterized in that, It includes the following steps: S1: Establish a deep reinforcement learning agent; S2: When the allocable resources of a certain container cluster node meet the resource requests of a certain to-be-scheduled task in the task queue, input the resource usage status of the container cluster node and the characteristic values of the to-be-scheduled task into the deep reinforcement learning agent to obtain the action probability distribution of task scheduling; S3: The container cluster scheduler schedules the to-be-scheduled task to execute in the container cluster node according to the action probability distribution of task scheduling, calculates the reward, and updates the network parameters of the deep reinforcement learning agent according to the reward; Among them, the calculation formula of the reward is as follows: Among them, represents the task priority reward obtained by scheduling task c at the time state of the th task by the deep reinforcement learning agent to an appropriate node for execution, represents the degree of imbalance in resource utilization within the container cluster node, represents the degree of imbalance in resource utilization among the container cluster nodes, is the weight coefficient of the penalty term for the imbalance in resource utilization within the container cluster node, is the weight coefficient for the imbalance in resource utilization among the container cluster nodes; Among them, the degree of imbalance in resource utilization within the container cluster nodes and the degree of imbalance in resource utilization among the container cluster nodes are expressed as follows: Among them, represents the time state of the c th scheduling of the deep reinforcement learning agent, and the N i resource utilization rate of node , including CPU resource utilization rate , memory resource utilization rate , and GPU resource utilization rate. Its expression is as follows: Among them, represents the available CPU resources of node c under the state of the N i th scheduling; represents the available memory resources of node c under the state of the N i th scheduling; represents the available GPU resources of node c under the state of the N i th scheduling; S4: Repeat S2 to S3 to train the deep reinforcement learning agent so that the deep reinforcement learning agent continuously learns and adjusts.

2. The method for resource scheduling of container clusters based on deep reinforcement learning according to claim 1, wherein The task priority reward The specific formula is as follows: Wherein, is the weighted reward coefficient of the task waiting time, is the priority reward of the task resource request volume, is the priority reward of the task delay sensitivity; represents the task priority reward obtained by scheduling tasks from the perspective of the task resource request volume, and its formula is as follows: where n is the number of nodes, is the weight reward coefficient of the corresponding resources of the GPU; represents the node N i the maximum available amount of CPU resources, represents the node N i the maximum available amount of memory resources, represents the node N i the maximum available amount of GPU resources; represents the CPU resource request amount of the task to be scheduled T j ; represents the memory request amount of the task to be scheduled T j ; represents the GPU resource request amount of the task to be scheduled T j ; Indicates the task priority reward obtained by scheduling tasks from the perspective of task waiting time, and its formula is as follows: = Among them, Q represents the task queue, represents the task T j waiting time, represents the task T k waiting time; It represents the task priority reward obtained by scheduling tasks from the perspective of task latency sensitivity. The tasks to be scheduled include inference tasks and training tasks. The formula of 。 3. The method for resource scheduling of container clusters based on deep reinforcement learning according to claim 2, wherein When considering the dependencies between tasks to be scheduled, the task priority reward function is expressed as: Among them, represents the task that depends on task Tj T z is the priority reward of represents the subtask that depends on task T j of 4. The method for container cluster resource scheduling based on deep reinforcement learning according to claim 3, wherein In S2, when the task priority reward of a certain task is greater than 0, it means that the allocable resources of a certain container cluster node meet the resource requests of the to-be-scheduled task.

5. The method for resource scheduling of container clusters based on deep reinforcement learning according to any one of claims 1-4, characterized in that, In S1, the characteristic values include the resource request volume, waiting time, dependent task index, and task latency sensitivity of the to-be-scheduled task, as well as the maximum available resources, current allocable resources, and current resource utilization rate of the container cluster node.

6. The method for container cluster resource scheduling based on deep reinforcement learning according to claim 5, wherein The deep reinforcement learning agent includes a policy network. Input the characteristic values into the policy network, and the policy network outputs the action probability distribution of task scheduling.

7. The method for container cluster resource scheduling of deep reinforcement learning according to claim 6, characterized in that The deep reinforcement learning agent also includes a value network, and the value network is used to score the policy network; Input the characteristic values into the value network to obtain the score of the policy network. Update the network parameters of the policy network and the value network according to the reward and the score of the policy network.

8. A container cluster resource scheduling system based on deep reinforcement learning, which is applied to the container cluster resource scheduling method based on deep reinforcement learning according to any one of claims 1 to 7, and is characterized in that, It includes: A deep reinforcement learning agent, which is used to obtain the action probability distribution of task scheduling according to the resource usage status of the container cluster node and the characteristic values of the to-be-scheduled task when the allocable resources of a certain container cluster node meet the resource requests of a certain to-be-scheduled task in the task queue; A container cluster scheduler, which is used to schedule the to-be-scheduled task to execute in the container cluster node according to the action probability distribution output by the deep reinforcement learning agent; An optimization module, which is used to calculate the reward, update the network parameters of the deep reinforcement learning agent according to the reward, and continuously optimize the deep reinforcement learning agent.

Citation Information

Patent Citations

  • Multi-agent reinforcement learning scheduling method and system, and electronic device

    CN109947567A

  • Cluster resource management and task scheduling method and system based on deep reinforcement learning

    CN111966484A