Container scheduling method and device, computer equipment and storage medium

By using node determination model in container scheduling and considering the real-time state of the node, the container binding failure caused by static information-based scheduling strategies is solved, and more efficient and flexible container scheduling is achieved.

CN119960964APending Publication Date: 2025-05-09CHINA TELECOM CLOUD TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411796659.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-09

Smart Images

  • Figure CN119960964A_ABST
    Figure CN119960964A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of container management, and discloses a container scheduling method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining request resource information of a to-be-scheduled container and real-time resource information of a first preset number of binding nodes; the real-time resource information and the request resource information are input into a node determination model, a target node is obtained, and the node determination model is used for determining the target node in the first preset number of binding nodes according to the real-time resource information and the request resource information; and binding the to-be-scheduled container with the target node. The problem that when the optimal node is determined according to the static node information, the optimal node may not be used when the container is scheduled, and consequently binding of the container fails is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of container management, and in particular to a container scheduling method, device, computer equipment and storage medium. Background Art

[0002] Kubernetes is the current mainstream container management technology, which can uniformly manage and monitor containers of cluster services. In the process of Kubernetes managing containers, it is necessary to schedule containers according to the container scheduling strategy, consider many factors such as resource requirements, hardware constraints, software constraints, etc., determine the optimal node, bind the container to the optimal node, and use the optimal node to run the container's workload. The quality of the container scheduling strategy directly affects the system energy consumption, cluster resource utilization, and user experience.

[0003] Kubernetes clusters are highly dynamic, and the status of nodes changes in real time. However, the current container scheduling strategy can be divided into two stages: pre-selection and optimization. The optimization stage determines the optimal node based on static node information, and it is difficult to consider the real-time status of the node. The determined optimal node may be available before calculation in the optimization stage, but unavailable before the actual scheduling of the container, resulting in failure to bind the container to the optimal node.

[0004] Therefore, the related art has the problem that the optimal node is determined based on static node information, and the optimal node may not be available when scheduling the container, resulting in the failure of container binding. Summary of the invention

[0005] In view of this, the present invention provides a container scheduling method, apparatus, computer device and storage medium to solve the problem that when determining the optimal node based on static node information, the optimal node may be unavailable when scheduling the container, resulting in container binding failure.

[0006] In a first aspect, the present invention provides a container scheduling method, comprising:

[0007] Obtaining requested resource information of the container to be scheduled and real-time resource information of a first preset number of bindable nodes;

[0008] Inputting the real-time resource information and the requested resource information into a node determination model to obtain a target node, wherein the node determination model is used to determine the target node from a first preset number of bindable nodes according to the real-time resource information and the requested resource information;

[0009] Bind the container to be scheduled to the target node.

[0010] The container scheduling method provided in this embodiment models the strategy of the node optimization stage to obtain a node determination model. The real-time resource information and the requested resource information are input into the node determination model. The node determination model determines the target node by considering the real-time status of the node, binds the container to be scheduled to the target node, and completes the container scheduling. A more flexible, efficient and adaptable node scheduling is achieved, which improves the robustness of the scheduling strategy, cluster utilization and user experience. It solves the problem that the optimal node may not be available when scheduling the container when the optimal node is determined based on static node information, resulting in the failure of container binding.

[0011] In some optional implementations, inputting the real-time resource information and the requested resource information into a node determination model to obtain a target node includes:

[0012] Initialize the number of training iterations;

[0013] Determine whether the number of training iterations is less than a first preset threshold;

[0014] When the number of training iterations is less than a first preset threshold, initializing the number of execution steps;

[0015] Extracting a second preset number of experience data groups from the experience replay pool as training samples, wherein the number of experience data groups in the experience replay pool is greater than or equal to the second preset number;

[0016] Using a second preset number of training samples to train the node determination model to obtain a trained node determination model;

[0017] Determine, based on the trained node determination model, whether there is a candidate node that meets a preset condition among the first preset number of bindable nodes;

[0018] If there is no candidate node, and the number of execution steps is less than the second preset threshold, the number of execution steps is increased by the first preset step length, and a second preset number of experience data groups are extracted from the experience replay pool as training samples to start executing subsequent steps;

[0019] If there is a candidate node, or the number of execution steps is greater than or equal to the second preset threshold, the number of training iterations is increased by the second preset step length, the trained node determination model is used as the node determination model, and subsequent steps are performed from the determination of whether the number of training iterations is less than the first preset threshold, until the number of training iterations is equal to the first preset threshold, then the process ends, and a third preset number of candidate nodes is obtained;

[0020] A target node is determined from a third preset number of candidate nodes.

[0021] In this embodiment, training samples are extracted from the experience replay pool to train the node determination model to improve the accuracy of the model in determining the target node. By setting the number of training iterations and the number of execution steps, the node determination model is trained for multiple rounds, and a third preset number of candidate nodes are obtained during the training process, and the target node is determined from the third preset number of candidate nodes, further improving the accuracy of the model.

[0022] In some optional implementations, before extracting a second preset number of experience data groups from the experience replay pool as training samples, the method further includes:

[0023] Get model configuration parameters and reward function;

[0024] Determine a first cluster state according to the real-time resource information;

[0025] Determine a candidate action according to the model configuration parameters and the preset strategy, wherein the candidate action is used to bind the bindable node to the intermediate node;

[0026] Execute the candidate action and obtain the reward value and the second cluster state according to the model configuration parameters, reward function, real-time resource information, and requested resource information;

[0027] The first cluster state, the candidate action, the reward value, and the second cluster state are used as an experience data group, and the experience data group is added to the experience replay pool;

[0028] Subsequent steps are executed starting from determining candidate actions according to the model configuration parameters and the preset strategy until the number of experience data groups in the experience replay pool is greater than or equal to a third preset threshold, wherein the third preset threshold is greater than or equal to the second preset number.

[0029] In this embodiment, the first cluster state, candidate actions, reward values, and second cluster state are used as experience data groups, and the experience data groups are added to the experience replay pool. Subsequently, training samples are extracted from the experience replay pool through a preset strategy, and past experience can be reused to improve the effective utilization of data and reduce the number of required interactions. In addition, using a preset strategy to break the correlation of experience can significantly improve training efficiency and stability.

[0030] In some optional implementations, determining the target node from a third preset number of candidate nodes includes:

[0031] Obtaining an evaluation index of the candidate node, or obtaining a frequency of each candidate node in a third preset number of candidate nodes;

[0032] The candidate node with the highest evaluation index or the candidate node with the highest frequency is selected as the target node.

[0033] In some optional implementations, after obtaining a third preset number of candidate nodes, the method further includes:

[0034] The trained node determination model is used as the target model;

[0035] When the requested resource information of the container to be bound and the real-time resource information of the available nodes are obtained, the requested resource information of the container to be bound and the real-time resource information of the available nodes are input into the target model to obtain the evaluation index of the available nodes;

[0036] Bind the container to be bound to the available node with the highest evaluation index.

[0037] In this embodiment, the trained node determination model is used as the target model and saved. The target model is then directly used to determine the available nodes corresponding to the container to be bound and bind them. There is no need to spend time training the model, thereby improving the efficiency of container scheduling.

[0038] In some optional implementations, obtaining a reward function includes:

[0039] A first index and a second index are obtained according to the requested resource information and the real-time resource information, wherein the first index is used to determine the load of the bindable node, and the second index is used to determine the matching degree between the bindable node and the container to be scheduled;

[0040] According to the first index, the second index and the preset weight, a reward function is obtained.

[0041] In this implementation, the actual node load and matching degree in the optimal scheduling phase are fully considered to determine the reward function of the optimal scheduling strategy, thereby improving the flexibility and pertinence of the scheduling strategy and avoiding the problem of unreliable estimation caused by setting sparse rewards in traditional reinforcement learning methods.

[0042] In some optional implementations, obtaining the first index and the second index according to the requested resource information and the real-time resource information includes:

[0043] According to the real-time resource information, determine the total amount of resources and the unallocated amount of resources that can be bound to the node;

[0044] According to the ratio of the total amount of resources to the unallocated amount of resources, a first index is obtained;

[0045] Determine the requested resource vector of the container to be scheduled according to the requested resource information;

[0046] Determine the remaining resource vector of the bindable node according to the real-time resource information;

[0047] The similarity between the requested resource vector and the remaining resource vector is determined, and a second index is obtained according to the similarity.

[0048] In this embodiment, the first index is determined according to the ratio of the total amount of resources of the bindable node to the unallocated amount of resources. The second index is determined according to the similarity between the request resource vector of the bindable node and the remaining resource vector of the node to be scheduled. The first index and the second index respectively characterize the load condition and matching degree of the node, thereby improving the flexibility and pertinence of the scheduling strategy.

[0049] In a second aspect, the present invention provides a container scheduling device, comprising:

[0050] An acquisition module, used to acquire requested resource information of the container to be scheduled and real-time resource information of a first preset number of bindable nodes;

[0051] An input module, used to input the real-time resource information and the requested resource information into a node determination model to obtain a target node, wherein the node determination model is used to determine the target node from a first preset number of bindable nodes according to the real-time resource information and the requested resource information;

[0052] The binding module is used to bind the container to be scheduled to the target node.

[0053] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the container scheduling method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0054] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the container scheduling method of the first aspect or any corresponding embodiment thereof.

[0055] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, wherein the computer instructions are used to enable a computer to execute the container scheduling method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the related technologies, the drawings required for use in the specific embodiments or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0057] Figure 1 is a flowchart of a container scheduling method according to an embodiment of the present invention;

[0058] Figure 2 is a flowchart of an improved Scheduler scheduling strategy combined with a deep reinforcement learning optimization strategy according to an embodiment of the present invention;

[0059] Figure 3 is a schematic diagram of an action space according to an embodiment of the present invention;

[0060] Figure 4 is a flowchart of a training node determination model according to an embodiment of the present invention;

[0061] Figure 5 is a structural block diagram of a container scheduling device according to an embodiment of the present invention;

[0062] Figure 6 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0064] Combined with the application scenarios that the execution of the container scheduling method depends on, the application scenarios are described here. Scheduler is the scheduler component in the Kubernetes cluster, which is responsible for scheduling the defined container (Pod) to the available cluster nodes. In the process of container scheduling, many factors such as the resource requirements of the container, the hardware constraints of the node, and the software constraints need to be considered to determine the best node to bind to the container to run the workload. The Scheduler scheduling strategy can be divided into two stages: pre-selection and optimization. Among them, the pre-selection stage mainly filters unavailable nodes and selects nodes that meet the Pod application resources by traversing all current nodes (Node). The optimization stage mainly scores the optional nodes selected in the pre-selection stage and selects the highest-scoring Node for the Pod to bind. There is a optimization strategy in the optimization stage. The optimization strategy is responsible for scoring the Node selected in the pre-selection stage according to factors such as node availability, resource utilization, and load balancing, and selecting the Node with the highest score to run the workload to ensure efficient utilization and load balancing of the cluster. However, the current optimization strategy is calculated based on static node information, which makes it difficult to consider the real-time status of the node. The Kubemetes cluster is highly dynamic, and it is possible that the cluster Node is available before optimization calculation, but unavailable before actual scheduling, resulting in Pod binding failure.

[0065] DRL (Deep Reinforcement Learning) is a machine learning method. The Deep Q Network (DQN) is a neural network based on DRL. It combines deep learning perception ability with reinforcement learning decision-making ability, uses the state information of the neural network learning environment, and continuously iterates through trial and error to obtain reward and punishment information to optimize the strategy and finally achieve the learning goal. This application models the Scheduler optimization strategy based on deep reinforcement learning. The improved optimization scheduling model can be described as a Markov decision process. The Markov decision process can be represented by a five-tuple (S, A, P sa , R, γ) represents, where S represents the set of states, A represents the set of actions, and P sa is the probability of taking action a from state s to reach state s_, which can be expressed as p(s_|s, a), R is the reward function, and performing action a in state s to obtain reward r can be expressed as r=R(s, a), γ is the decay factor, and its value is between 0 and 1. The closer it is to 1, the more importance is attached to future rewards. The goal of deep reinforcement learning optimization strategy is to learn the optimal node scheduling strategy through trial and error interaction with the environment, that is, for any state s, there is π≥π′, where π is the optimal scheduling strategy, π′ is other strategies, and the optimal state value function v under the optimal strategy π (s) is as shown in formula (1).

[0066] v π (s)=∑π(a|s)∑p(s_|s,a)[r+γv π (s_)] (1)

[0067] The powerful decision-making ability of deep reinforcement learning provides a modeling idea for the Scheduler optimization strategy modeling. The optimization strategy mainly models the environment, state, action, and reward function elements to achieve node scheduling in the optimization stage.

[0068] Based on the above content, an embodiment of the present invention provides a container scheduling method, which combines an improved Scheduler scheduling strategy with a deep reinforcement learning optimization strategy, models the Scheduler optimization strategy based on deep reinforcement learning, and obtains a node determination model. For the cluster Scheduler pre-selection stage, a list of Nodes that meet the operating conditions is obtained; input into the node determination model, the optimized Node is obtained, and the Pod is finally bound. The Scheduler optimization strategy is modeled based on deep reinforcement learning to solve the problem that it is difficult to consider the real-time status of the node in the existing optimization strategy. In order to achieve the technical effect of fully considering the real-time status of the node, ensuring that the selected node is optimal and can be successfully bound to the container.

[0069] According to an embodiment of the present invention, an embodiment of a container scheduling method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, for example: a computer, a server, etc., and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0070] In this embodiment, a container scheduling method is provided, which can be used in the above-mentioned computer system. Figure 1 is a flowchart of a container scheduling method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0071] Step S101: obtaining requested resource information of a container to be scheduled and real-time resource information of a first preset number of bindable nodes.

[0072] Specifically, the Scheduler scheduling strategy can be divided into two stages: pre-selection and optimization. The pre-selection stage mainly filters out unavailable nodes, and filters out nodes that meet the Pod application resources by traversing all current nodes (Node), and obtains a Node list that meets the running conditions of the container to be scheduled. The Node list contains a first preset number of bindable nodes, and the first preset number is, for example: 1, 2, 3... The specific value is determined according to actual needs.

[0073] In addition, the number of Pods to be scheduled in the preferred stage can be one or more, and a Pod list can be set to store the containers to be scheduled. For example, the Pod list to be scheduled in the preferred stage can be represented as a set P = (Pod1, Pod2, Pod3, ..., Pod i , ..., Pod m ), the status is the current resource status of the first Pod in the Pod container list. Get the requested resource information of the container to be scheduled. The requested resource information can be expressed as a resource tuple vector (req cpu ,req mem ,req disk ,req network ), where req cpu ,req mem ,req disk ,req network Respectively represent the requests of the container to be scheduled for CPU, memory, disk, and network.

[0074] The current environment in the optimization stage can be described as the Node set N=(N1, N2, N3, ..., N i , ..., N pre), the Node node set N is the Node list. Get the real-time resource information of the first preset number of bindable nodes. The real-time resource information is the remaining resource of each Node node, which can be expressed as a resource tuple vector (rest cpu , rest mem , rest disk , rest network ), rest cpu , rest mem , rest disk , rest network Respectively indicate the remaining status of the current node's CPU, memory, disk, and network.

[0075] The above process is as follows Figure 2 As shown, determine the list of containers to be bound and the current Node list of the cluster. The current Node list of the cluster includes nodes N1 to N n , determine a container to be scheduled in the list of containers to be bound, and determine the pre-selected Node list through the pre-selection stage. The pre-selected Node list includes nodes N1 to N pre .

[0076] Step S102: input the real-time resource information and the requested resource information into a node determination model to obtain a target node, wherein the node determination model is used to determine the target node from a first preset number of bindable nodes according to the real-time resource information and the requested resource information.

[0077] Specifically, the node determination model, for example: a deep reinforcement learning optimization model obtained by modeling the Scheduler optimization strategy based on deep reinforcement learning, a pre-trained deep reinforcement learning optimization model, etc.

[0078] Input the real-time resource information and the requested resource information into the node determination model to obtain the optimal Nod e That is, the target node. It should be noted that if the node determination model is a deep reinforcement learning optimization model, the agent of the node determination model takes actions based on the current available node set N. When a Pod request arrives, the optimization strategy selects the Node node to complete the binding based on the Pod request resource situation.

[0079] Then the node model can select the optimal action in the action space, the action space is as follows Figure 3 As shown, the action space can be described as an action list A = (a1, a2, a3, ..., a pre ), each action means binding the container to be scheduled to a node in the pre-selected Node list, for example, action a1 means binding the container to be scheduled to node N1. After determining the optimal action, the node to which the container to be scheduled needs to be bound, i.e., the target node, can be determined.

[0080] Step S103: Bind the container to be scheduled to the target node.

[0081] Specifically, the container to be scheduled is bound to the target node, and the target node is used to run the workload of the container to be scheduled.

[0082] The above process is as follows Figure 2 As shown, a container to be scheduled is determined in the list of containers to be bound, and the deep reinforcement learning optimization stage determines the preferred Node based on the container to be scheduled and the pre-selected Node list. The container to be scheduled is bound to the preferred Node.

[0083] The container scheduling method provided in this embodiment models the strategy of the node optimization stage to obtain a node determination model. The real-time resource information and the requested resource information are input into the node determination model. The node determination model determines the target node by considering the real-time status of the node, binds the container to be scheduled to the target node, and completes the container scheduling. A more flexible, efficient and adaptable node scheduling is achieved, which improves the robustness of the scheduling strategy, cluster utilization and user experience. It solves the problem that the optimal node may not be available when scheduling the container when the optimal node is determined based on static node information, resulting in the failure of container binding.

[0084] In this embodiment, another container scheduling method is provided, which can be used in the above-mentioned computer system. The process of the method includes the following steps:

[0085] Step S401: Obtain requested resource information of a container to be scheduled and real-time resource information of a first preset number of bindable nodes.

[0086] For details, please see Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.

[0087] Step S402: input the real-time resource information and the requested resource information into a node determination model to obtain a target node, wherein the node determination model is used to determine the target node from a first preset number of bindable nodes according to the real-time resource information and the requested resource information.

[0088] Specifically, the above step S402 includes:

[0089] Step S4021, initialize the number of training iterations.

[0090] The number of training iterations is denoted as i, and initializing the number of training iterations is to set i=0.

[0091] Step S4022, determining whether the number of training iterations is less than a first preset threshold.

[0092] The first preset threshold is represented by numepisode , determined according to actual needs, by judging whether the number of training iterations is less than the first preset threshold, the training node can be controlled to determine the number of training rounds of the model process, num episode The value is set according to actual needs, for example: 500, 600.

[0093] Step S4023, when the number of training iterations is less than the first preset threshold, initialize the number of execution steps.

[0094] The execution step is represented by step. When the number of training iterations is less than the first preset threshold, the execution step is initialized, that is, step=0 is set. The execution step is used to control the number of repeated operations in a training iteration process to avoid the training iteration process being repeated and wasting computing resources.

[0095] Step S4024: extract a second preset number of experience data groups from the experience replay pool as training samples, wherein the number of experience data groups in the experience replay pool is greater than or equal to the second preset number.

[0096] The second preset number is batch_size, and the value of batch_size is set according to actual needs. A second preset number of experience data groups are extracted from the experience replay pool as training samples, and the experience data group is, for example: (s, a, r, s_), where s represents the current state of the cluster, a represents the executed action, r represents the reward value of executing action a when the state is s, and s_ represents the state of the cluster after executing action a. The experience replay pool can be determined based on the historical operations of the cluster, or an action can be randomly determined using the node determination model, and the cluster is used to execute the action, and the four-tuple (s, a, r, s_) is added to the experience replay pool.

[0097] Step S4025: train the node determination model using a second preset number of training samples to obtain a trained node determination model.

[0098] The node determination model is trained using a second preset number of training samples, the loss function is calculated and the DQN network parameters in the node determination model are updated, so that the optimal scheduling strategy network model approaches the target Q value, and the trained node determination model is obtained.

[0099] Step S4026, determining the model based on the trained node, and determining whether there is a candidate node that meets the preset conditions among the first preset number of bindable nodes.

[0100] The action space can be described as an action list A = (a1, a2, a3, ..., a pre), each action represents binding the container to be scheduled to a node in the pre-selected Node list. According to the trained node determination model, the action with the preset condition can be determined from the action list A. For example, the preset condition is that the Q value of the action is the largest. The trained node determination model determines the Q value of each action in the action list A, and the node to be bound to the action with the largest Q value is used as the candidate node.

[0101] Step S4027, if there is no candidate node and the execution step number is less than the second preset threshold, the execution step number is increased by the first preset step length, and a second preset number of experience data groups are extracted from the experience replay pool as training samples to start executing subsequent steps.

[0102] Specifically, the second preset threshold is max steps , for example: 50, 100, the specific value is set according to actual needs. If the candidate node cannot be determined in the action list A, for example: the Q value of some actions is wrong, the node determination model reports an error, etc., which leads to the inability to determine the candidate node. In addition, the current execution step is compared with the second preset threshold. If the current execution step is less than the second preset threshold, the execution step is increased by the first preset step length. The first preset step length is for example: 1, 2. The specific value is set according to actual needs. The execution step is increased by the first preset step length, for example: step=step+1. Re-execute the above steps S4024 to S4027.

[0103] Step S4028, if there is a candidate node, or the execution step number is greater than or equal to the second preset threshold, the number of training iterations is increased by the second preset step size, the trained node determination model is used as the node determination model, and subsequent steps are executed from the determination of whether the number of training iterations is less than the first preset threshold, until the number of training iterations is equal to the first preset threshold, then the process ends and a third preset number of candidate nodes is obtained.

[0104] Specifically, two cutoff conditions for a training iteration process of the node determination model are: a candidate node is determined during the training iteration process; and the number of execution steps is greater than or equal to a second preset threshold.

[0105] If there is a candidate node, or the number of execution steps is greater than or equal to the second preset threshold, the current training iteration process is stopped, the number of training iterations is increased by the second preset step size, the next round of training iteration process is started, and the above steps S4022 to S4028 are re-executed. The second preset step size is, for example, 1 or 2. The specific value is set according to actual needs. The number of training iterations is increased by the second preset step size, for example, i=i+1.

[0106] According to the above cutoff conditions, a training iteration process may stop because a candidate node is determined, or it may stop because the number of execution steps is greater than or equal to the second preset threshold. Therefore, executing the first preset threshold training iteration process can obtain one or more candidate nodes. Therefore, the third preset number, for example: 1, 2, ..., 100, ... the specific value is determined according to the actual training situation.

[0107] Step S4029: determine a target node from a third preset number of candidate nodes.

[0108] Specifically, if the third preset number is 1, there is only one candidate node, and the candidate node is directly used as the target node. If the third preset number is greater than 1, the target node is determined from the third preset number of candidate nodes, for example: the Q value of the action corresponding to each candidate node is obtained, and the candidate node corresponding to the action with the largest Q value is used as the target node.

[0109] In this embodiment, training samples are extracted from the experience replay pool to train the node determination model to improve the accuracy of the model in determining the target node. By setting the number of training iterations and the number of execution steps, the node determination model is trained for multiple rounds, and a third preset number of candidate nodes are obtained during the training process, and the target node is determined from the third preset number of candidate nodes, further improving the accuracy of the model.

[0110] In some optional implementations, before step S2024, the another container scheduling method further includes:

[0111] Step a1, obtain model configuration parameters and reward function.

[0112] Step a2: determining the first cluster state according to the real-time resource information.

[0113] Step a3: determining a candidate action according to the model configuration parameters and the preset strategy, wherein the candidate action is used to bind the bindable node to the intermediate node.

[0114] Step a4, execute the candidate action, and obtain the reward value and the second cluster state according to the model configuration parameters, the reward function, the real-time resource information and the requested resource information.

[0115] Step a5: taking the first cluster state, the candidate actions, the reward value and the second cluster state as an experience data group, and adding the experience data group to the experience replay pool.

[0116] Step a6, starting from determining the candidate action according to the model configuration parameters and the preset strategy, executing subsequent steps until the number of experience data groups in the experience replay pool is greater than or equal to a third preset threshold, wherein the third preset threshold is greater than or equal to the second preset number.

[0117] Specifically, obtain the model configuration parameters and reward function R, such as the experience replay pool size max_size, batch size batch_size, greedy strategy parameter ε, discount factor γ, learning rate lr, number of training rounds num episode , the maximum step length of a single round max steps .

[0118] Real-time resource information includes the remaining resources of each Node, which can be expressed as a resource tuple vector (rest cpu , rest mem , rest disk , rest network ), rest cpu , rest mem , rest disk , rest network Respectively represent the remaining status of the current node CPU, memory, disk, and network. In the optimization stage, the cluster state can be described as the Node set N = (N1, N2, N3, ..., N i , ..., N pre ), the Node set N is the Node list. The remaining resources of each Node are summarized to obtain the first cluster state s.

[0119] The preset strategy is, for example, the ε-greddy (greedy) strategy. The first cluster state and the requested resource information of the container to be scheduled are input into the node determination model. The node determination model uses the greedy strategy parameter ε in the model configuration parameters and the preset strategy in the action list A = (a1, a2, a3, ..., a pre )Determine candidate action a.

[0120] Execute candidate action a and bind the bindable node to the intermediate node corresponding to candidate action a. According to the model configuration parameters, reward function R, real-time resource information and requested resource information, obtain the reward value r and the new cluster state s_, i.e., the second cluster state.

[0121] The first cluster state, candidate actions, reward values ​​and second cluster state are taken as experience data groups (s, a, r, s_), and the experience data groups are added to the experience replay pool. The number of experience data groups in the experience replay pool is Cur_size+=1.

[0122] The third preset threshold is the experience replay pool size max_size in the model configuration parameters. Steps a3 to a6 are re-executed until the number of experience data groups in the experience replay pool is greater than or equal to the third preset threshold (ie, cur_size≥max_size), then the process stops.

[0123] Combining steps S4021 to S4029 with steps a1 to a6, the process of the optimal scheduling strategy algorithm can be obtained, such as Figure 4 As shown, it includes: inputting the pre-selected Node list and the requested resource information of the container to be scheduled into the node determination model. The node determination model performs an optimization stage based on deep reinforcement learning. The cluster environment s is determined according to the remaining resource situation of the nodes in the pre-selected Node list. Randomly select batch size tuples from the experience revisit pool as training samples, input the first cluster state s in the training sample into the prediction network, and input the second cluster state s into the target network. The loss function is calculated according to the Q value Q_eval(s, a|θ) output by the prediction network and the Q_target(s_, a|θ′) output by the target network, and the network parameter θ of the prediction network is updated according to the loss function. After updating the preset times (for example: 10, 20 times), the network parameter θ′ of the target network is equal to the network parameter θ of the prediction network, so that the optimal scheduling strategy network model approaches the target Q value, and the trained node determination model is obtained. The trained node determination model can determine the Q value of each action in the action list A, determine the maximum value maxQ therein, and the best action is argmaxQ(s, a|θ). The optimal Node is determined based on the best action, and the container to be scheduled is bound to the optimal Node.

[0124] Combining steps S4021 to S4029 with steps a1 to a6, another Scheduler optimization scheduling strategy algorithm design based on deep reinforcement learning can be determined, for example:

[0125] Initialize the optimal scheduling DQN model network; action set A; read N and Pod request resource data; training iteration number i = 0; current size of the experience replay pool cur_size = 0;

[0126] for i:=to num episode do

[0127] Initialize the current step number step=0

[0128] while True do

[0129] Select action a according to the ε-greddy greedy strategy

[0130] Execute action a, get reward value r and new fluid state s_

[0131] Add the quaternion (s, a, r, s_) to the experience replay pool, cur_size += 1

[0132] if(cur_size≥max_size)do

[0133] Extract batch_size experience from the experience revisit pool

[0134] Calculate the loss function and update the DQN network parameters

[0135] if(Pod m Binding successful)do

[0136] Output the preferred Node N pri

[0137] break

[0138] if(step≥max steps )do

[0139] break

[0140] s=s_

[0141] step=step+1

[0142] In this embodiment, the first cluster state, candidate actions, reward values, and second cluster state are used as experience data groups, and the experience data groups are added to the experience replay pool. Subsequently, training samples are extracted from the experience replay pool through a preset strategy, and past experience can be reused to improve the effective utilization of data and reduce the number of required interactions. In addition, using a preset strategy to break the correlation of experience can significantly improve training efficiency and stability.

[0143] In some optional implementations, the above step S2029 includes:

[0144] Step b1, obtaining the evaluation index of the candidate node, or obtaining the frequency of each candidate node in a third preset number of candidate nodes.

[0145] Step b2: taking the candidate node with the highest evaluation index or the candidate node with the highest frequency as the target node.

[0146] Specifically, the evaluation index is, for example, a Q value. The evaluation index of each candidate node output by the node determination model is obtained, and the candidate node with the highest evaluation index is used as the target node.

[0147] Alternatively, the frequency of each candidate node in the third preset number of candidate nodes is obtained, for example, the third preset number is 100, that is, there are 100 candidate nodes, including 20 nodes 1, 50 nodes 2, 16 nodes 3, and 14 nodes 4. It can be seen that the frequency of node 1 is 20, the frequency of node 2 is 50, the frequency of node 3 is 16, and the frequency of node 4 is 14. The candidate node with the highest frequency is used as the target node, for example, node 2 has the highest frequency, and node 2 is used as the target node.

[0148] In some optional implementations, after step S2028, the another container scheduling method further includes:

[0149] Step c1, taking the trained node determination model as the target model.

[0150] Step c2: when the requested resource information of the container to be bound and the real-time resource information of the available nodes are obtained, the requested resource information of the container to be bound and the real-time resource information of the available nodes are input into the target model to obtain the evaluation index of the available nodes.

[0151] Step c3: bind the container to be bound to the available node with the highest evaluation index.

[0152] In an embodiment, after the model / neural network training is completed, the trained model / neural network can be saved and preloaded. If there are matters that need to be judged later, they are input into the preloaded model / neural network to obtain the expected results.

[0153] Specifically, the trained node determination model is used as the target model, and the target model is preloaded.

[0154] The evaluation index of the available nodes is, for example, the Q value. When the requested resource information of the container to be bound and the real-time resource information of the available nodes are obtained, the requested resource information of the container to be bound and the real-time resource information of the available nodes are input into the target model, and the target model can obtain the evaluation index of the corresponding action of each available node. The container to be bound is bound to the available node with the highest evaluation index.

[0155] In this embodiment, the trained node determination model is used as the target model and saved. The target model is then directly used to determine the available nodes corresponding to the container to be bound and bind them. There is no need to spend time training the model, thereby improving the efficiency of container scheduling.

[0156] In some optional implementations, the above step a1 includes:

[0157] Step a11, obtaining a first index and a second index according to the requested resource information and the real-time resource information, wherein the first index is used to determine the load of the bindable node, and the second index is used to determine the matching degree between the bindable node and the container to be scheduled.

[0158] Step a12, obtaining a reward function according to the first index, the second index and the preset weight.

[0159] In an embodiment, the reward r can be determined by a reward function. The reward r is the reward or punishment information given by the environment for the action performed, which represents the reward value obtained by selecting action a in state s. The reward function can be expressed as R(s, a). The goal of the optimal scheduling strategy based on deep reinforcement learning is to continuously interact with the environment to maximize the cumulative reward.

[0160] Specifically, in order to avoid the problem of unreliable estimation caused by setting sparse rewards in traditional reinforcement learning methods, this embodiment fully considers the load conditions and matching degree of actual nodes to improve the convergence of the algorithm, and defines a reward function suitable for the Scheduler's optimal scheduling strategy, as shown in formula (2).

[0161] R=μ1R match +μ2R score (2)

[0162] Among them, μ1 and μ2 are loss coefficients, R match Represents the matching degree reward between Pod and node, R score Represents the node load score reward.

[0163] According to the requested resource information and the real-time resource information, the first index and the second index are obtained, and the first index is R score , the second index is R match , the preset weights are μ1 and μ2. Therefore, according to the first index, the second index and the preset weights, the reward function R shown in formula (2) can be obtained.

[0164] In this implementation, the actual node load and matching degree in the optimal scheduling phase are fully considered to determine the reward function of the optimal scheduling strategy, thereby improving the flexibility and pertinence of the scheduling strategy and avoiding the problem of unreliable estimation caused by setting sparse rewards in traditional reinforcement learning methods.

[0165] In some optional implementations, the above step a11 includes:

[0166] Step a111, determining the total amount of resources and the unallocated amount of resources of the bindable nodes according to the real-time resource information.

[0167] Step a112, obtaining a first index according to the ratio of the total amount of resources to the amount of unallocated resources.

[0168] Step a113: Determine the requested resource vector of the container to be scheduled according to the requested resource information.

[0169] Step a114, determining the remaining resource vectors of the bindable nodes according to the real-time resource information.

[0170] Step a115, determining the similarity between the requested resource vector and the remaining resource vector, and obtaining a second index according to the similarity.

[0171] In the embodiment, the current node load is determined according to the total amount of node resources and the allocated amount. In order to reduce the waste of node resources and fully consider the matching degree between the container to be scheduled and the node, the angle cosine is introduced to calculate the similarity between the remaining resource vector of the node and the requested resource vector of the container to be scheduled. The greater the similarity, the higher the matching degree between the container to be scheduled and the node, otherwise the lower the matching degree.

[0172] Specifically, according to the real-time resource information, determine the total amount of resources and the amount of resources allocated to the bindable node. For example, the total amount of resources is: cpu 、Total mem 、Total disk 、Total network , which respectively represent the total amount of resources of the node CPU, memory, disk, and network. Resource allocation amount, for example: Allocated cpu 、Allocated mem 、Allocated disk 、Allocated network , respectively represent the allocated amount of the node's CPU, memory, disk, and network resources. The unallocated amount of resources can be obtained by subtracting the allocated amount of the corresponding resources from the total amount of resources.

[0173] In order to fully consider the score of each node, the different resources of the node are scored and the score is defined separately. cpu 、score mem 、sscore disk 、score network , which respectively represent the current scores of the node’s CPU, memory, disk, and network.

[0174] The score of each resource is obtained based on the ratio of the total amount of resources to the unallocated amount of resources. For example, the score can be obtained by dividing the total amount of CPU resources by the unallocated amount of CPU resources. cpu , as shown in formula (3). The score is obtained by dividing the total amount of memory resources by the unallocated amount of memory resources. mem , as shown in formula (4). By dividing the total amount of disk resources by the unallocated amount of disk resources, we can get score disk , as shown in formula (5). By dividing the total amount of network resources by the unallocated amount of network resources, we can get score network As shown in formula (6).

[0175]

[0176] Integrate the above scores cpu sCore mem 、score disk 、score network , get the node load score reward R score , the first index is the node load score reward R score , as shown in formula (7).

[0177] R score =w cpu *score cpu +w mem *score mem +w disk *score disk +w network *score network (7)

[0178] Among them, w cpu 、w mem 、w disk 、w network They represent the weight coefficients of the node's CPU, memory, disk, and network respectively. The sum of the weight coefficients of all resources is 1. By adjusting the weight coefficients of various resources, nodes that meet the conditions can be screened out.

[0179] According to the requested resource information, determine the requested resource vector of the container to be scheduled, for example: (req cpu ,req mem ,req disk ,req network ), indicating the request of the container to be scheduled for CPU, memory, disk, and network. According to the real-time resource information, determine the remaining resource vector of the node that can be bound, for example: (rest cpu , rest mem , rest disk , rest network ), indicating the remaining status of the current node's CPU, memory, disk, and network.

[0180] The angle cosine is introduced to calculate the similarity between the remaining resource vector of the node and the requested resource vector of the container to be scheduled. The second index is obtained based on the similarity. The second index is the matching degree reward R match , the greater the similarity, the higher the matching degree between the container to be scheduled and the node, otherwise the matching degree is lower. As shown in formula (8).

[0181]

[0182] In this embodiment, the first index is determined according to the ratio of the total amount of resources of the bindable node to the unallocated amount of resources. The second index is determined according to the similarity between the request resource vector of the bindable node and the remaining resource vector of the node to be scheduled. The first index and the second index respectively characterize the load condition and matching degree of the node, thereby improving the flexibility and pertinence of the scheduling strategy.

[0183] Step S403: Bind the container to be scheduled to the target node. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.

[0184] Another container scheduling method provided in this embodiment models the Scheduler optimization strategy based on deep reinforcement learning to solve the problem that the existing optimization strategy is difficult to consider the real-time status of the node. The load and matching degree of the actual node in the optimization scheduling stage are fully considered, and the optimization scheduling strategy reward function is defined to improve the flexibility and pertinence of the scheduling strategy.

[0185] In this embodiment, a container scheduling device is also provided, which is used to implement the above embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0186] This embodiment provides a container scheduling device, such as Figure 5 As shown, including:

[0187] An acquisition module 501 is used to acquire requested resource information of a container to be scheduled and real-time resource information of a first preset number of bindable nodes;

[0188] An input module 502 is used to input the real-time resource information and the requested resource information into a node determination model to obtain a target node, wherein the node determination model is used to determine the target node from a first preset number of bindable nodes according to the real-time resource information and the requested resource information;

[0189] The binding module 503 is used to bind the container to be scheduled with the target node.

[0190] In some optional implementations, the input module 502 includes:

[0191] A first initialization unit, used to initialize the number of training iterations;

[0192] A judging unit, used to judge whether the number of training iterations is less than a first preset threshold;

[0193] A second initialization unit, used to initialize the number of execution steps when the number of training iterations is less than a first preset threshold;

[0194] An extraction unit, used to extract a second preset number of experience data groups from the experience replay pool as training samples, wherein the number of experience data groups in the experience replay pool is greater than or equal to the second preset number;

[0195] A training unit, used to train the node determination model using a second preset number of training samples to obtain a trained node determination model;

[0196] A first determination unit, configured to determine, based on a trained node determination model, whether there is a candidate node that meets a preset condition among a first preset number of bindable nodes;

[0197] A first loop unit is used for increasing the number of execution steps by a first preset step length if there is no candidate node and the number of execution steps is less than a second preset threshold, and starting to execute subsequent steps by extracting a second preset number of experience data groups from the experience replay pool as training samples;

[0198] The second loop unit is used for increasing the number of training iterations by a second preset step length if there is a candidate node, or the number of execution steps is greater than or equal to a second preset threshold, using the trained node determination model as the node determination model, and executing subsequent steps from judging whether the number of training iterations is less than the first preset threshold, until the number of training iterations is equal to the first preset threshold, then ending, and obtaining a third preset number of candidate nodes;

[0199] The second determining unit is used to determine the target node from a third preset number of candidate nodes.

[0200] In some optional implementations, the input module 502 includes:

[0201] An acquisition unit, used to obtain model configuration parameters and reward functions;

[0202] A third determining unit, configured to determine a first cluster state according to the real-time resource information;

[0203] a fourth determining unit, configured to determine a candidate action according to the model configuration parameters and a preset strategy, wherein the candidate action is used to bind the bindable node to the intermediate node;

[0204] An execution unit, used to execute the candidate action and obtain a reward value and a second cluster state according to the model configuration parameters, the reward function, the real-time resource information, and the requested resource information;

[0205] A storage unit, used to take the first cluster state, the candidate action, the reward value, and the second cluster state as an experience data group, and add the experience data group to an experience replay pool;

[0206] The third loop unit is used to execute subsequent steps starting from determining candidate actions based on model configuration parameters and preset strategies until the number of experience data groups in the experience replay pool is greater than or equal to a third preset threshold, wherein the third preset threshold is greater than or equal to the second preset number.

[0207] In some optional implementations, the second determining unit includes:

[0208] An acquisition submodule, used to acquire an evaluation index of a candidate node, or to acquire a frequency of each candidate node among a third preset number of candidate nodes;

[0209] A submodule is set to select the candidate node with the highest evaluation index or the candidate node with the highest frequency as the target node.

[0210] In some optional implementations, the input module 502 includes:

[0211] A setting unit, used for taking the trained node determination model as a target model;

[0212] An input unit, configured to input the requested resource information of the container to be bound and the real-time resource information of the available nodes into a target model to obtain an evaluation index of the available nodes when the requested resource information of the container to be bound and the real-time resource information of the available nodes are obtained;

[0213] The binding unit is used to bind the container to be bound to the available node with the highest evaluation index.

[0214] In some optional implementations, the acquiring unit includes:

[0215] A first determination submodule is used to obtain a first index and a second index according to the requested resource information and the real-time resource information, wherein the first index is used to determine the load of the bindable node, and the second index is used to determine the matching degree between the bindable node and the container to be scheduled;

[0216] The second determination submodule is used to obtain a reward function according to the first index, the second index and the preset weight.

[0217] In some optional implementations, the first determining submodule includes:

[0218] A first determining subunit is used to determine the total amount of resources and the unallocated amount of resources of the bindable node according to the real-time resource information;

[0219] A second determining subunit is used to obtain a first index according to a ratio of the total amount of resources to the unallocated amount of resources;

[0220] A third determining subunit is used to determine a request resource vector of the container to be scheduled according to the request resource information;

[0221] A fourth determining subunit, configured to determine a remaining resource vector of a bindable node according to the real-time resource information;

[0222] The fifth determining subunit is used to determine the similarity between the requested resource vector and the remaining resource vector, and obtain a second index according to the similarity.

[0223] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0224] The container scheduling device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0225] The embodiment of the present invention also provides a computer device having the above Figure 5 The container scheduling device shown.

[0226] See also Figure 6 , Figure 6 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.

[0227] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include an integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0228] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0229] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0230] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0231] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0232] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0233] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0234] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined in this application.

Claims

1. A container scheduling method, characterized in that: The method comprises: Obtaining requested resource information of the container to be scheduled and real-time resource information of a first preset number of bindable nodes; Inputting the real-time resource information and the requested resource information into a node determination model to obtain a target node, wherein the node determination model is used to determine the target node from the first preset number of bindable nodes according to the real-time resource information and the requested resource information; Bind the container to be scheduled to the target node.

2. The method according to claim 1, characterized in that The step of inputting the real-time resource information and the requested resource information into a node determination model to obtain a target node includes: Initialize the number of training iterations; Determining whether the number of training iterations is less than a first preset threshold; When the number of training iterations is less than the first preset threshold, initializing the number of execution steps; Extracting a second preset number of experience data groups from the experience replay pool as training samples, wherein the number of experience data groups in the experience replay pool is greater than or equal to the second preset number; Using a second preset number of the training samples to train the node determination model to obtain a trained node determination model; Determine, according to the trained node determination model, whether there is a candidate node that meets a preset condition among the first preset number of bindable nodes; If the candidate node does not exist and the number of execution steps is less than the second preset threshold, the number of execution steps is increased by the first preset step length, and a second preset number of experience data groups are extracted from the experience replay pool as training samples to start executing subsequent steps; If the candidate node exists, or the number of execution steps is greater than or equal to the second preset threshold, the number of training iterations is increased by a second preset step length, the trained node determination model is used as the node determination model, and subsequent steps are performed from the determination of whether the number of training iterations is less than the first preset threshold until the number of training iterations is equal to the first preset threshold, then the step ends, and a third preset number of candidate nodes is obtained; The target node is determined from the third preset number of candidate nodes.

3. The method according to claim 2, characterized in that Before extracting a second preset number of experience data groups from the experience replay pool as training samples, the method further includes: Get model configuration parameters and reward function; Determining a first cluster state according to the real-time resource information; Determining a candidate action according to the model configuration parameters and a preset strategy, wherein the candidate action is used to bind the bindable node to an intermediate node; Execute the candidate action, and obtain a reward value and a second cluster state according to the model configuration parameters, the reward function, the real-time resource information, and the requested resource information; Taking the first cluster state, the candidate action, the reward value, and the second cluster state as the experience data group, and adding the experience data group to the experience replay pool; Execute subsequent steps starting from determining candidate actions according to the model configuration parameters and preset strategies until the number of experience data groups in the experience replay pool is greater than or equal to a third preset threshold, wherein the third preset threshold is greater than or equal to the second preset number.

4. The method according to claim 2, characterized in that: The determining the target node from the third preset number of candidate nodes includes: Obtaining the evaluation index of the candidate node, or obtaining the frequency of each of the candidate nodes in the third preset number of candidate nodes; The candidate node with the highest evaluation index or the candidate node with the highest frequency is used as the target node.

5. The method according to claim 2, characterized in that: After obtaining the third preset number of candidate nodes, the method further includes: Using the trained node determination model as the target model; When the requested resource information of the container to be bound and the real-time resource information of the available nodes are obtained, the requested resource information of the container to be bound and the real-time resource information of the available nodes are input into the target model to obtain an evaluation index of the available nodes; The container to be bound is bound to the available node with the highest evaluation index.

6. The method according to claim 3, characterized in that Obtaining the reward function includes: Obtaining a first index and a second index according to the requested resource information and the real-time resource information, wherein the first index is used to determine the load of the bindable node, and the second index is used to determine the matching degree between the bindable node and the container to be scheduled; The reward function is obtained according to the first index, the second index and the preset weight.

7. The method according to claim 6, characterized in that The obtaining a first index and a second index according to the requested resource information and the real-time resource information comprises: Determine the total amount of resources and the unallocated amount of resources of the bindable node according to the real-time resource information; Obtaining the first index according to a ratio of the total amount of resources to the unallocated amount of resources; Determining a requested resource vector of the container to be scheduled according to the requested resource information; Determining a remaining resource vector of the bindable node according to the real-time resource information; Determine the similarity between the requested resource vector and the remaining resource vector, and obtain the second index according to the similarity.

8. A container scheduling device, characterized in that: The device comprises: An acquisition module, used to acquire requested resource information of the container to be scheduled and real-time resource information of a first preset number of bindable nodes; An input module, used for inputting the real-time resource information and the requested resource information into a node determination model to obtain a target node, wherein the node determination model is used for determining the target node from the first preset number of bindable nodes according to the real-time resource information and the requested resource information; A binding module is used to bind the container to be scheduled with the target node.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the container scheduling method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the container scheduling method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Container scheduling method and apparatus

    CN122526837A

  • Container scheduling method and apparatus

    CN122526837B