Pod scheduling model construction method and device, equipment and medium

By building a pod scheduling model in the K8s cluster, using iterative training and joint reward mechanisms to optimize container scheduling decisions, the problems of insufficient scheduling efficiency and resource utilization in the existing technology are solved, and more efficient resource allocation and Pod scheduling are achieved.

CN119987992AActive Publication Date: 2025-05-13CHINA TELECOM CLOUD TECH CO LTD

Patent Information

Application Number
CN202411797690.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-13
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

The existing container scheduling methods have insufficient scheduling efficiency and resource utilization, resulting in unbalanced resource supply and demand, which in turn affects the load balancing and resource utilization of the cluster.

Method used

A pod scheduling model construction method is proposed. By considering the remaining resources of each node in the K8s cluster, the pod scheduling model is used to continuously learn how to make the best scheduling decisions based on the current cluster state and the needs of the pod to be scheduled. This model optimizes scheduling decisions through a joint reward mechanism, combining the target Pod request resource reward, the degree of matching between the Pod and the candidate node, and the impact reward of the candidate node on other nodes.

Benefits of technology

It improves the scheduling efficiency and resource utilization of Pods, reduces resource waste, quickly finds the right node to schedule Pods, reduces scheduling delay, and increases the chance of Pods being successfully started and run. The model is able to adapt to different workloads and cluster states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987992A_ABST
    Figure CN119987992A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a pod scheduling model construction method and device, equipment and a medium, and the method comprises the steps: obtaining joint state data and environment data in a current training period in a current cycle; inputting the joint state data and the environment data into a pod scheduling model for iterative training; when the joint reward corresponding to the current training period is generated, according to the joint reward and / or a training period statistical value corresponding to the current training period, determining whether the current cycle stops iterative training or not; and when it is determined that the current cycle stops iterative training, entering the next cycle, and obtaining joint state data and environment data from the sample data pool corresponding to the next cycle to perform iterative training on the pod scheduling model. According to the method, the model can be ensured to better allocate resources, resource waste is reduced, and scheduling delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a pod scheduling model construction method, device, equipment and medium. Background Art

[0002] Kubernetes is the current mainstream container orchestration technology that simplifies container deployment, expansion, and management. Container scheduling, as an important function of Kubernetes, can realize the automated deployment of containerized applications, improve the reliability and scalability of programs, and ensure cluster load balancing. Therefore, the quality of container scheduling strategies directly affects the availability and resource utilization of clusters.

[0003] Current container scheduling methods mostly calculate the status information of Pods and Nodes, and select the most suitable Node at the moment to complete the binding according to the scheduling strategy. However, they ignore the mutual influence between each Node, resulting in low scheduling efficiency, and there is also the problem of low resource utilization due to imbalance in resource supply and demand. Summary of the invention

[0004] In view of this, the present invention provides a pod scheduling model construction method, device, equipment and medium to solve the problems of low scheduling efficiency and low utilization rate caused by imbalance in resource supply and demand in the container scheduling method in the related art.

[0005] In a first aspect, the present invention provides a method for constructing a pod scheduling model, the method comprising:

[0006] In the current training cycle in the current loop, obtain joint state data and environment data from the sample data pool corresponding to the current loop, wherein the joint state data is the remaining resources of each of the multiple nodes in the K8s cluster, and the environment data includes the multiple pods currently to be scheduled;

[0007] The joint state data and environment data obtained from the sample data pool corresponding to the current cycle are input into the pod scheduling model to iteratively train the pod scheduling model in the current cycle;

[0008] When the pod scheduling model generates a joint reward corresponding to the current training cycle, determine whether to stop iterative training in the current loop based on the joint reward and / or the training cycle statistics corresponding to the current training cycle, where the joint reward includes a target pod request resource reward, a match degree reward between the target pod and the candidate node, and an inter-node scheduling similarity reward corresponding to the influence of the candidate node and nodes other than the candidate node in the K8s cluster on the candidate node. The target pod is any one of the multiple pods, and the candidate node is any one of the multiple nodes.

[0009] When it is determined that the current cycle stops iterative training based on the joint reward and / or the training cycle statistics corresponding to the current training cycle, the next cycle is entered, and the joint status data and environment data are obtained from the sample data pool corresponding to the next cycle to iteratively train the pod scheduling model until the number of cycles reaches the preset threshold, then the iterative training is stopped to obtain the final pod scheduling model.

[0010] The pod scheduling model construction method provided by the present invention has the following advantages:

[0011] By considering the remaining resources of each node in the K8s cluster, the model can allocate Pods more effectively, thereby maximizing the utilization of cluster resources. Through iterative training, the model continuously learns how to make the best scheduling decision based on the current cluster status and the needs of the Pod to be scheduled, thereby improving the scheduling efficiency of the Pod. In particular, in this application document, two nested loops are included, one of which is to set different iterative training cycles in each loop, and the other is to include multiple loops for iterative training, so that the iterative training process is richer, and the model finally obtained will be able to adapt to different situations. The joint reward mechanism takes into account multiple factors, including the target Pod request resource reward, the Pod and the candidate node matching degree reward, and the candidate node’s impact reward on other nodes, which helps the model make more comprehensive and reasonable scheduling decisions. In particular, the method of this application takes into account the impact of other nodes on a certain node, and maximizes the scheduling efficiency and resource utilization. In addition, the model can adjust its behavior according to the reward corresponding to the current training cycle and the training cycle statistics, so that the model can adapt to the changing environment and needs. Through cyclic iterative training, the model can continuously learn new scheduling strategies from the sample data pool until the preset training number threshold is reached, ensuring that the model fully absorbs the information in the data during the training process. Through multiple iterations of training and acquisition of the final model, it can be ensured that the trained Pod scheduling model has high accuracy and reliability.

[0012] In summary, this method can ensure that the model allocates resources better and reduces resource waste. It can quickly find the right node to schedule the Pod and reduce scheduling delays. Through precise matching, it can increase the chances of successful Pod startup and operation. The model can adapt to different workloads and cluster states.

[0013] In an optional implementation, the method for acquiring sample data in the sample data pool corresponding to each cycle includes:

[0014] In the current loop, the initial joint state data and environment data are obtained, wherein the initial joint state data is the initial remaining resources of each of the multiple nodes in the k8s cluster in the current loop;

[0015] In the initial data generation cycle, the initial joint state data and environment data are input into the initial pod scheduling model to generate the joint action corresponding to the next data generation cycle and the joint reward corresponding to the current data generation cycle;

[0016] The initial joint state data corresponding to the initial data generation cycle, the environment data, the joint action corresponding to the next data generation cycle, and the joint reward corresponding to the current data generation cycle constitute a four-tuple as a sample data, and add it to the sample data pool corresponding to the current cycle;

[0017] In the i-th data generation cycle, based on the joint action corresponding to the i-th generation cycle and the joint state data corresponding to the i-1-th data generation cycle, generate joint state data corresponding to the i-th data generation cycle, where i is a positive integer greater than or equal to 2 and i is incremented by 1;

[0018] Input the joint state data and environment data corresponding to the i-th data generation cycle into the initial pod scheduling model, obtain the joint action corresponding to the i+1-th data generation cycle, and the joint reward corresponding to the i-th data generation cycle;

[0019] The four-tuple consisting of the joint state data and environment data corresponding to the i-th data generation cycle, the joint reward corresponding to the i-th data generation cycle, and the joint action corresponding to the i+1-th data generation cycle is added as a sample data to the sample data pool corresponding to the current cycle;

[0020] The iteration stops when i is equal to the preset threshold, and all sample data in the sample data pool corresponding to the current loop are obtained.

[0021] Specifically, by dynamically determining the joint state in each data generation cycle, the model can capture the dynamic characteristics of resource allocation and Pod scheduling in the K8s cluster over time. The joint actions and joint rewards generated in each data generation cycle simulate the decisions and results in the Pod scheduling process, which helps to train a scheduling model that is more in line with the actual situation. Through continuous data generation and iteration, the collected sample data is richer, including multiple situations under different states and decisions, which is conducive to the generalization ability of the model. Using pre-generated sample data for training can improve the efficiency of model training because there is no need to interact with the K8s cluster in real time to obtain data. Through sample data containing multiple states and decisions, the model can learn more accurate and comprehensive scheduling strategies. Because the training data covers a variety of possible scheduling situations, it adapts to different workloads and cluster states. In this method, multiple sample data can be generated in the above manner to avoid the situation where the model training cannot meet the requirements due to insufficient sample data. Moreover, the generated sample data can help the model learn how to use cluster resources more effectively and reduce resource waste. In addition, the pre-generated sample data can reduce the need for real-time data collection, thereby shortening the training time. And it can ensure that the final trained model can make scheduling decisions more quickly and improve the scheduling efficiency of Pod. Sample data can help the Pod scheduling model learn more accurately how to efficiently schedule Pod in the K8s cluster, thereby improving the overall performance and resource utilization of the cluster.

[0022] In an optional implementation, when the pod scheduling model generates a joint reward corresponding to the current training cycle, determining whether to stop iterative training in the current cycle according to the joint reward and / or the training cycle statistics corresponding to the current training cycle specifically includes:

[0023] When the reward value of the joint reward is greater than or equal to the preset reward threshold, determining to stop iterative training in the current cycle;

[0024] and / or,

[0025] When it is determined that the training cycle statistics corresponding to the current training cycle reaches a preset cycle quantity value, it is determined that the current cycle stops iterative training.

[0026] Specifically, when the reward value of the joint reward generated by the model reaches or exceeds the preset reward threshold, the model can be considered to have learned an effective scheduling strategy. Stopping iterative training at this time can avoid overfitting and save computing resources. When the training cycle reaches the preset number, the iterative training is stopped regardless of the performance of the model. This ensures that the model completes training in a reasonable time and prevents unnecessary long-term training. When the training cycle reaches a certain number, the model has accumulated enough experience, and continuing training may not significantly improve performance, so stopping training can ensure model stability. Setting clear stopping conditions in the above way helps to understand when the model meets performance requirements and improves the transparency and explainability of the training process.

[0027] In an optional implementation manner, the resource types of the remaining resources in each node include one or more of the following:

[0028] CPU resources, memory resources, disk resources, network resources, video memory resources, and GPU resources.

[0029] In an optional implementation, when the resource types include CPU resources, memory resources, disk resources, network resources, video memory resources, and GPU resources, the expression of the target pod request resource reward is as follows:

[0030] R score =w cpu *req cpu +w mem *req mem +w disk *req disk +w net *req net +w vimem *req vimem +w gpu *req gpu Formula 1)

[0031] w cpu +w mem +w disk +w net +w vimem +w gpu =1 (Formula 2)

[0032] Among them, R score The reward value for requesting resources for the target pod, req cpu The remaining CPU resources, req mem Remaining memory resources, req disk The remaining disk resources, req net is the remaining network resources, req vimemFor the remaining resources of video memory, req gpu is the remaining GPU resources, w cpu 、w mem 、w disk 、w net 、w vimem , and w gpu are the weight coefficients corresponding to each resource.

[0033] In an optional implementation, when the resource types include CPU resources, memory resources, disk resources, network resources, video memory resources, and GPU resources, the matching degree reward between the target pod and the candidate node is expressed by the following expression:

[0034]

[0035] Among them, R similar The reward value for the matching degree between the target pod and the candidate node. For joint actions, Pod i Schedule to N j Run the containerized application on the node, i is the i-th pod, j is the j-th node, and They represent the remaining CPU resources on the i-th pod and the remaining CPU resources on the j-th node respectively; and Represent the remaining memory resources on the i-th pod and the remaining memory resources on the i-th node respectively; and Represent the remaining disk resources on the i-th pod and the remaining disk resources on the i-th node respectively; and They represent the remaining network resources on the i-th pod and the remaining network resources on the j-th node respectively; and They represent the remaining video memory resources on the i-th pod and the remaining video memory resources on the j-th node respectively; and They represent the remaining GPU resources on the i-th pod and the remaining GPU resources on the j-th node respectively.

[0036] In an optional implementation, the scheduling similarity reward between nodes is expressed by the following expression:

[0037]

[0038] Among them, R attention is the scheduling similarity reward between the nodes. For the i-th node, Q is the scheduling similarity query vector, K TSchedule similarity feature representation for each node, which is used to calculate similarity with Q, and V scores the actual similarity of each node for the final weighted output, d k The dimension size of the scheduling similarity query vector.

[0039] In an optional implementation, the joint reward expression is as follows:

[0040] R=μ1R score +μ2R similar +μ3R attention (Formula 6)

[0041] Wherein, R is the reward value of the joint reward, μ1, μ2, and μ3 are pre-configured weight coefficients respectively.

[0042] In a second aspect, the present invention provides a pod scheduling model construction device, the device comprising:

[0043] An acquisition module is used to acquire joint state data and environment data from a sample data pool corresponding to the current cycle in the current training cycle in the current cycle, wherein the joint state data is the remaining resources of each of the multiple nodes in the K8s cluster, and the environment data includes the multiple pods currently to be scheduled;

[0044] The training module is used to obtain the joint state data and environment data from the sample data pool corresponding to the current cycle and input them into the pod scheduling model to iteratively train the pod scheduling model in the current cycle;

[0045] A processing module is used to determine whether to stop iterative training in the current loop according to the joint reward and / or the training cycle statistics corresponding to the current training cycle when the pod scheduling model generates a joint reward corresponding to the current training cycle, wherein the joint reward includes a target pod request resource reward, a match degree reward between the target pod and the candidate node, and an inter-node scheduling similarity reward corresponding to the influence of the candidate node and nodes other than the candidate node in the K8s cluster on the candidate node, the target pod is any one of the multiple pods, and the candidate node is any one of the multiple nodes; when it is determined that the current loop stops iterative training, enter the next loop;

[0046] The acquisition module is also used to acquire the joint state data and the environment data from the sample data pool corresponding to the next cycle;

[0047] The training module is also used to iteratively train the pod scheduling model by obtaining joint status data and environment data from the sample data pool corresponding to the next cycle, until the number of cycles reaches a preset threshold, stop the iterative training, and obtain the final pod scheduling model.

[0048] The pod scheduling model construction device provided by the present invention has the following advantages:

[0049] By considering the remaining resources of each node in the K8s cluster, the model can allocate Pods more effectively, thereby maximizing the utilization of cluster resources. Through iterative training, the model continuously learns how to make the best scheduling decision based on the current cluster status and the needs of the Pod to be scheduled, thereby improving the scheduling efficiency of the Pod. In particular, in this application document, two nested loops are included, one of which is to set different iterative training cycles in each loop, and the other is to include multiple loops for iterative training, so that the iterative training process is richer, and the model finally obtained will be able to adapt to different situations. The joint reward mechanism takes into account multiple factors, including the target Pod request resource reward, the Pod and the candidate node matching degree reward, and the candidate node’s impact reward on other nodes, which helps the model make more comprehensive and reasonable scheduling decisions. In particular, in this application, the impact of other nodes on a certain node is taken into account to maximize the scheduling efficiency and resource utilization. In addition, the model can adjust its behavior according to the reward corresponding to the current training cycle and the training cycle statistics, so that the model can adapt to the changing environment and needs. Through cyclic iterative training, the model can continuously learn new scheduling strategies from the sample data pool until the preset training number threshold is reached, ensuring that the model fully absorbs the information in the data during the training process. Through multiple iterations of training and acquisition of the final model, it can be ensured that the trained Pod scheduling model has high accuracy and reliability.

[0050] In summary, this method can ensure that the model allocates resources better and reduces resource waste. It can quickly find the right node to schedule the Pod and reduce scheduling delays. Through precise matching, it can increase the chances of successful Pod startup and operation. The model can adapt to different workloads and cluster states.

[0051] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the pod scheduling model construction method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0052] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the pod scheduling model construction method of the first aspect or any corresponding embodiment thereof.

[0053] In a fifth aspect, the present invention provides a computer program product, including computer instructions, which are used to enable a computer to execute the pod scheduling model construction method of the above-mentioned first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0055] Figure 1 It is a flow chart of a method for constructing a pod scheduling model provided by an embodiment of the present invention;

[0056] Figure 2 It is a schematic diagram of the structural principle of a pod scheduling model provided by the present invention;

[0057] Figure 3 is a schematic diagram of environmental data provided by the present invention;

[0058] Figure 4 is a schematic diagram of the joint status data provided by the present invention;

[0059] Figure 5 The present invention provides a Pod i Schedule to N j Schematic diagram of the joint actions of running containerized applications on;

[0060] Figure 6 It is a flow chart of another pod scheduling model construction method provided by an embodiment of the present invention;

[0061] Figure 7 It is a schematic diagram of the overall process of producing sample data provided by the present invention;

[0062] Figure 8 It is a schematic diagram of the overall process of a pod scheduling model construction method provided by the present invention;

[0063] Fig. 9 It is a structural block diagram of a pod scheduling model building device provided by an embodiment of the present invention;

[0064] Fig.10 It is a schematic diagram of the hardware structure of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0066] To solve the above problems, an embodiment of the present invention provides a pod scheduling embodiment. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system (computer device) including, for example, a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0067] In this embodiment, a pod scheduling model construction method is provided, which can be used for the above-mentioned terminal devices, such as mobile phones, tablet computers, etc. Figure 1 It is a flow chart of a pod scheduling model construction method provided by an embodiment of the present invention.

[0068] Before introducing the method steps of the embodiments of the present application, the terms that may be involved in the embodiments of the present application are first explained, and the details are as follows:

[0069] Reinforcement Learning (RL): The agent interacts with the environment to obtain reward and punishment information to adjust the search strategy to maximize the cumulative reward and ultimately achieve the learning goal.

[0070] Multi-Agent Reinforcement Learning (MARL): is an extension of reinforcement learning. Multiple agents obtain information and optimize strategies by observing other agents and environmental dynamics to maximize the performance of the overall system.

[0071] Container scheduling: The process of allocating containers to appropriate host nodes based on factors such as container resource requirements, scheduling policies, and node resource conditions, ensuring high availability and load balancing of the nodes.

[0072] The following describes the method steps in the embodiments of the present application. Figure 1 As shown, the process includes the following steps:

[0073] Step S101, in a current training cycle in a current cycle, obtaining joint state data and environment data from a sample data pool corresponding to the current cycle.

[0074] The joint status data is the remaining resources of each of the multiple nodes in the K8s cluster, and the environment data includes the multiple pods currently to be scheduled.

[0075] Specifically, the container scheduling process consists of the list of Pods to be scheduled P = (Pod1, Pod2, Pod3, ..., Pod i ,…,Pod m ), the remaining resources of the currently available Node nodes in the cluster, etc. form a complex dynamic multi-agent system. Therefore, the process of binding a "single-agent" Node node to a Pod to run a containerized application can be expanded to multi-agent collaboration. In this process, each agent (Node node) needs to consider the interaction with the dynamic environment and the mutual influence between other agents, which can generally be described as (N, S, A, P sa ,R,γ), where N=(N1,N2,N3,…,N i ,…,N n ) represents N nodes; S, A, and R represent the joint state, action, and reward of the current multi-agent system respectively; P sa It represents the probability of the multi-agent system executing an action from the joint state s to s_, which can be recorded as p(s_|s,a); γ is the reward attenuation coefficient, which is between 0 and 1. The larger it is, the more importance is attached to future rewards.

[0076] In the embodiment of the present application, multiple loops may be included, and multiple cycles of iterative training are performed in each loop.

[0077] Before executing iterative training, you need to obtain joint state data and environment data from the sample data pool corresponding to the current cycle. The environment data is multiple pods in the list of pods to be scheduled. The number of pods corresponding to each cycle can be the same or different. The specific pods can also be the same or different.

[0078] Step S102, obtaining joint state data and environment data from a sample data pool corresponding to the current cycle and inputting them into a pod scheduling model, so as to iteratively train the pod scheduling model in the current cycle.

[0079] Step S103, when the pod scheduling model generates a joint reward corresponding to the current training cycle, it is determined whether to stop iterative training in the current cycle according to the joint reward and / or the training cycle statistics corresponding to the current training cycle.

[0080] Specifically, in an optional example, the pod scheduling model can be a reinforcement learning model, which is a model built on the basis of the actor-critic algorithm (AC) based on the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, in which the action state information of all Node agents is input into the evaluation network, and each Node agent is executed separately through its own policy network, so that the model is more robust. However, considering that the model cannot distribute the joint reward to each agent according to the credibility, there is a problem of "credibility distribution". Therefore, the attention mechanism is also introduced in the model, and the influence between agents is differentiated through the attention mechanism, so as to achieve the purpose of distributing the joint reward to different agents according to the credibility. Therefore, the final model is MADADPG (Multi-Agent Deep Attention Deterministic Policy Gradient).

[0081] For more information about the pod scheduling model, see Figure 2 As shown, Figure 2 It is shown in Figure 1 that the scheduling model includes multiple sub-models, and the multiple sub-models obtain the influence of each other through the attention mechanism and add it to the reward value of the joint reward generated by each sub-model. Figure 2 In the example, a sub-model is set up as a sub-model corresponding to the N1 node and a sub-model corresponding to the Nn node. The sub-model corresponding to N1 is used as an example to illustrate that Actor1 generates the action corresponding to the current moment according to the current reward value and updates the state of the N1 node. Critic1 generates a reward value by calculating the state of all nodes and the action similarity information to evaluate the pros and cons of the current scheduling strategy. All sub-models update the corresponding joint state and joint action at the current moment based on the above principle. The joint reward includes the target pod request resource reward, the match degree reward between the target pod and the candidate node, and the influence degree reward corresponding to the influence of the candidate node and the nodes in the K8s cluster other than the candidate node on the candidate node. The target pod is any one of the multiple pods, and the candidate node is any one of the multiple nodes.

[0082] For details, see Figure 3 As shown, Figure 3 A schematic diagram of the environment data is shown in FIG. In an optional example, the Pod list P = (Pod1, Pod2, Pod3, ..., Pod i , ..., Pod m), the resource types that a pod can request can include, for example, CPU, memory, disk, network, etc. However, in order to meet the diversity of cluster resource requirements, GPU and video memory can also be included in the scheduling criteria. Among them, the resource requested by the i-th container Pod application can be described as a resource tuple vector in Indicates the request status of each resource of the i-th Pod. cpu, mem, disk, net, vimem, and gpu represent CPU, memory, disk, network, video memory, and GPU indicators respectively.

[0083] Figure 4 The schematic diagram of the joint state data is shown in FIG. 1 . The state of container scheduling can be described as the current idle Node node N of the cluster = (N1, N2, N3, ..., N j , ..., N n )'s joint state, such as Figure 4 As shown, the remaining resources of the jth Node are available as a tuple vector Indicates that Indicates the remaining status of each resource on the j-th Node.

[0084] The container scheduling multi-agent system selects joint actions based on the available Node nodes N in the current cluster to achieve efficient container resource scheduling, such as Figure 5 As shown in , the action space of the joint action can be represented as a list of Pod containers to be scheduled, such as Figure 5 As shown, the joint action Where i and j are both integers, representing Pod i Schedule to N j Run containerized applications on it.

[0085] Step S104, when it is determined that the current cycle stops iterative training based on the joint reward and / or the training cycle statistics corresponding to the current training cycle, enter the next cycle, and obtain the joint state data and environment data from the sample data pool corresponding to the next cycle to iteratively train the pod scheduling model until the number of cycles reaches the preset number threshold, stop the iterative training, and obtain the final pod scheduling model.

[0086] Specifically, when the number of joint rewards and / or training cycles generated in the current cycle has reached a preset number, that is, the training cycle statistics value (maximum number of training steps), it can be determined that the current cycle stops training and then enters the next cycle.

[0087] After entering the next cycle, the iterative training of the next cycle is performed with reference to the above-mentioned method steps, wherein the sample data used in the training process is the joint state data and environment data obtained in the sample delay corresponding to the next cycle. When the number of cycles reaches the preset number threshold, the iterative training is stopped to obtain the final pod scheduling model.

[0088] The pod scheduling model construction method provided in this embodiment, by considering the remaining resources of each node in the K8s cluster, the model can allocate Pod more effectively, thereby maximizing the utilization of cluster resources. The model continuously learns how to make the best scheduling decision according to the current cluster state and the needs of the Pod to be scheduled through iterative training, thereby improving the scheduling efficiency of the Pod. In particular, in this application document, two nested loops are included, one of which is to set different iterative training cycles in each loop, and the other is to include multiple loops for iterative training, so that the iterative training process is richer, and the model finally obtained will be able to adapt to different situations. The joint reward mechanism takes into account multiple factors, including the target Pod request resource reward, the matching degree reward of the Pod and the candidate node, and the influence reward of the candidate node on other nodes, which helps the model to make more comprehensive and reasonable scheduling decisions. In particular, in the present application method, the influence of other nodes on a certain node is taken into account to improve the scheduling efficiency and resource utilization as much as possible. In addition, the model can adjust its behavior according to the reward corresponding to the current training cycle and the training cycle statistics, so that the model can adapt to the changing environment and needs. Through iterative training, the model can continuously learn new scheduling strategies from the sample data pool until the preset training times threshold is reached, ensuring that the model fully absorbs the information in the data during the training process. Through multiple iterative training and the acquisition of the final model, it can be ensured that the trained Pod scheduling model has high accuracy and reliability.

[0089] In summary, this method can ensure that the model allocates resources better and reduces resource waste. It can quickly find the right node to schedule the Pod and reduce scheduling delays. Through precise matching, it can increase the chances of successful Pod startup and operation. The model can adapt to different workloads and cluster states.

[0090] In this embodiment, a pod scheduling model construction method is provided, which can be used for the above-mentioned mobile terminals, such as mobile phones, tablet computers, etc. Figure 2 is a flow chart of a method for constructing a pod scheduling model provided by an embodiment of the present invention, such as Figure 2 As shown, based on the above embodiment, the method for obtaining sample data in the sample data pool corresponding to each cycle includes the following method steps, see Figure 6 As shown:

[0091] Step S601, in the current loop, initial joint state data and environment data are obtained.

[0092] The initial joint state data is the initial remaining resources of each of the multiple nodes in the k8s cluster in the current cycle.

[0093] Step S602, in the initial data generation cycle, the initial joint state data and the environment data are input into the initial pod scheduling model to generate a joint action corresponding to the next data generation cycle and a joint reward corresponding to the current data generation cycle.

[0094] Specifically, as introduced above, the pod scheduling model is the MADADPG model. Specifically, after the initial joint state data and environment data are input into the initial pod scheduling model, corresponding joint actions and joint rewards are generated.

[0095] Specifically, for n nodes N=(N1, N2, N3, ..., N j , ..., N n ), the joint strategy function can be expressed as π = (v(θ1), π(θ2),, ..., v(θ N )), where θ i , i∈[1,n] is the policy function parameter of the ith agent. The input of the evaluation network (Critic) in the model training phase is the observed state S=(s1, s2, ..., s i , ..., s n ) and action set A = (a1, a2, ..., a i , ..., a n ), then the Q value of each agent can be expressed as:

[0096]

[0097] in, For the target network, FC1 i , FC2 i is the fully connected layer of the ith agent, V i is the value weight sum of all other agents, V i It can be expressed as:

[0098]

[0099] Among them, j∈[1,n] and j≠i, ReLU is the activation function, W is a fixed matrix, and the attention weight α j The embedding e i and e j For comparison, we can prevent the gradient from disappearing by scaling the matrix dimensions.j It can be expressed as:

[0100]

[0101] in W i e i and e j The weight coefficient of , then the loss function of the evaluation network can be expressed as:

[0102]

[0103] in (S, A, R, S_)~D are the sample data in the sample data pool corresponding to the current cycle in the current training stage, and E is the expectation.

[0104] At this time, the decision network parameter update can be expressed as:

[0105]

[0106] Among them, θ μ is the policy network parameter, α u is the policy network learning rate, represents the gradient descent method, and J represents the expected reward function.

[0107] After training, the agent only needs to consider part of the observation input (the input in the evaluation network itself includes part of the observation data plus the input of other agents, and part of the observation input includes joint rewards), and input it into the decision network to obtain the optimal decision action and corresponding joint reward.

[0108] Step S603, the initial joint state data corresponding to the initial data generation cycle, the environmental data, the joint action corresponding to the next data generation cycle, and the joint reward corresponding to the current data generation cycle constitute a four-tuple as a sample data, and add it to the sample data pool corresponding to the current cycle.

[0109] Step S604, in the ith data generation cycle, based on the joint action corresponding to the ith generation cycle and the joint state data corresponding to the (i-1)th data generation cycle, generates joint state data corresponding to the ith data generation cycle.

[0110] Here, i is a positive integer greater than or equal to 2, and i increases by 1.

[0111] Specifically, the joint action means that the scheduled pod is scheduled to a certain node according to the current situation. At this time, if it is scheduled to the node, it will occupy certain resources. Therefore, the joint status data in the current cycle, that is, the i-th cycle, can be generated based on the scheduling action and the joint status data in the i-1th cycle. Among them, i increases with the increase of the cycle, and the step length of each increase is 1.

[0112] Step S605, input the joint state data and environment data corresponding to the i-th data generation cycle into the initial pod scheduling model, obtain the joint action corresponding to the i+1-th data generation cycle, and the joint reward corresponding to the i-th data generation cycle.

[0113] Step S606, take the four-tuple consisting of the joint status data and environmental data corresponding to the i-th data generation cycle, the joint reward corresponding to the i-th data generation cycle, and the joint action corresponding to the (i+1)-th data generation cycle as a sample data, and add it to the sample data pool corresponding to the current cycle, until i is equal to the preset threshold, stop the iteration, and obtain all the sample data in the sample data pool corresponding to the current cycle.

[0114] Figure 7 The overall process diagram is shown in FIG. 1 , including different nodes respectively generating joint scheduling actions A1, A2, ..., An based on the input joint reward R and joint state S through the above model, and then scheduling the pods in the list of containers to be scheduled according to the scheduling action A, and then obtaining the joint reward and the joint state S of the next cycle. Further, based on the joint state and joint reward of the next cycle, a new scheduling action is generated... This cycle repeats to obtain the aforementioned sample data and add it to the sample data pool.

[0115] By dynamically determining the joint state in each data generation cycle, the model can capture the dynamic characteristics of resource allocation and Pod scheduling in the K8s cluster over time. The joint actions and joint rewards generated in each data generation cycle simulate the decisions and results in the Pod scheduling process, which helps to train a scheduling model that is more in line with the actual situation. Through continuous data generation and iteration, the collected sample data is richer, including multiple situations under different states and decisions, which is conducive to the generalization ability of the model. Using pre-generated sample data for training can improve the efficiency of model training because there is no need to interact with the K8s cluster in real time to obtain data. With sample data containing multiple states and decisions, the model can learn more accurate and comprehensive scheduling strategies. Because the training data covers a variety of possible scheduling situations, it adapts to different workloads and cluster states. In this method, multiple sample data can be generated in the above way to avoid the situation where the model training cannot meet the requirements due to insufficient sample data. Moreover, the generated sample data can help the model learn how to use cluster resources more effectively and reduce resource waste. In addition, pre-generated sample data can reduce the need for real-time data collection, thereby shortening the training time. And it can ensure that the final trained model can make scheduling decisions more quickly and improve the scheduling efficiency of Pod. Sample data can help the Pod scheduling model learn more accurately how to efficiently schedule Pod in the K8s cluster, thereby improving the overall performance and resource utilization of the cluster.

[0116] In an optional example, when the pod scheduling model generates a joint reward corresponding to the current training cycle, determining whether to stop iterative training in the current cycle according to the joint reward and / or the training cycle statistics corresponding to the current training cycle specifically includes:

[0117] When the reward value of the joint reward is greater than or equal to the preset reward threshold, determining to stop iterative training in the current cycle;

[0118] and / or,

[0119] When it is determined that the training cycle statistics corresponding to the current training cycle reaches a preset cycle quantity value, it is determined that the current cycle stops iterative training.

[0120] Specifically, when the reward value of the joint reward generated by the model reaches or exceeds the preset reward threshold, it can be considered that the model has learned an effective scheduling strategy. Stopping iterative training at this time can avoid overfitting and save computing resources. When the training cycle reaches the preset number, the iterative training is stopped regardless of the performance of the model. It can ensure that the model completes the training in a reasonable time and prevent unnecessary long-term training. When the training cycle reaches a certain number, the model has accumulated enough experience, and continuing training may not significantly improve performance, so stopping training can ensure the stability of the model. Setting clear stopping conditions in the above way helps to understand when the model meets the performance requirements and improves the transparency and explainability of the training process.

[0121] Based on any of the foregoing embodiments, when the resource types include CPU resources, memory resources, disk resources, network resources, video memory resources, and GPU resources, the expression of the target pod request resource reward is as follows:

[0122] R score =w cpu *req cpu +w mem *req mem +w disk *req disk +w net *req net +w vimem *req vimem +w gpu *req gpu (Formula 12)

[0123] w cpu +w mem +w disk +w net +w vimem +w gpu =1 (Formula 13)

[0124] Among them, R score The reward value for requesting resources for the target pod, req cpu The remaining CPU resources, req mem Remaining memory resources, req disk The remaining disk resources, req net is the remaining network resources, req vimem For the remaining resources of video memory, req gpu is the remaining GPU resources, w cpu 、w mem 、w disk 、w net 、w vimem , and w gpuare the weight coefficients corresponding to each resource.

[0125] Based on any of the foregoing embodiments, when the resource types include CPU resources, memory resources, disk resources, network resources, video memory resources, and GPU resources, the matching degree reward between the target pod and the candidate node is expressed by the following expression:

[0126]

[0127] Among them, R similar The reward value for the matching degree between the target pod and the candidate node. For joint actions, Pod i Schedule to N j Run the containerized application on the node, i is the i-th pod, j is the j-th node, and They represent the remaining CPU resources on the i-th pod and the remaining CPU resources on the j-th node respectively; and Represent the remaining memory resources on the i-th pod and the remaining memory resources on the i-th node respectively; and Represent the remaining disk resources on the i-th pod and the remaining disk resources on the i-th node respectively; and They represent the remaining network resources on the i-th pod and the remaining network resources on the j-th node respectively; and They represent the remaining video memory resources on the i-th pod and the remaining video memory resources on the j-th node respectively; and They represent the remaining GPU resources on the i-th pod and the remaining GPU resources on the j-th node respectively.

[0128] The scheduling similarity reward between nodes is expressed as follows:

[0129]

[0130] Among them, R attention is the scheduling similarity reward between the nodes. For the i-th node, Q is the scheduling similarity query vector, K T Schedule similarity feature representation for each node, which is used to calculate similarity with Q, and V scores the actual similarity of each node for the final weighted output, d k The dimension size of the scheduling similarity query vector.

[0131] In an optional implementation, the joint reward expression is as follows:

[0132] R=μ1R score +μ2R similar +μ3R attention (Formula 17)

[0133] Wherein, R is the reward value of the joint reward, μ1, μ2, and μ3 are pre-configured weight coefficients respectively.

[0134] It should be noted that for all resource types, All need to be satisfied

[0135]

[0136] Figure 8 The overall flow chart of the above method is shown in FIG. Figure 8 As shown, the specific implementation process has been described in detail in the previous article, so it will not be repeated here.

[0137] In addition, the embodiment of the present application also provides a pod scheduling method. Corresponding environment data and joint state data are obtained, and then input into the aforementioned trained model, so that corresponding scheduling actions can be generated.

[0138] In this embodiment, a pod scheduling model construction device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated here. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0139] This embodiment provides a pod scheduling model construction device, such as Fig. 9 As shown, it includes: an acquisition module 901, a training module 902 and a processing module 903.

[0140] An acquisition module 901 is used to acquire joint state data and environment data from a sample data pool corresponding to the current cycle in a current training cycle in the current cycle, wherein the joint state data is the remaining resources of each of the multiple nodes in the K8s cluster, and the environment data includes multiple pods currently to be scheduled;

[0141] The training module 902 is used to obtain the joint state data and the environment data from the sample data pool corresponding to the current cycle and input them into the pod scheduling model to iteratively train the pod scheduling model in the current cycle;

[0142] Processing module 903 is used to determine whether to stop iterative training in the current loop according to the joint reward and / or the training cycle statistics corresponding to the current training cycle when the pod scheduling model generates a joint reward corresponding to the current training cycle, wherein the joint reward includes a target pod request resource reward, a match degree reward between the target pod and the candidate node, and an inter-node scheduling similarity reward corresponding to the influence of the candidate node and the nodes other than the candidate node in the K8s cluster on the candidate node, the target pod is any one of the multiple pods, and the candidate node is any one of the multiple nodes; when it is determined that the current loop stops iterative training, enter the next loop;

[0143] The acquisition module 901 is also used to acquire the joint state data and the environment data from the sample data pool corresponding to the next cycle;

[0144] The training module 902 is also used to iteratively train the pod scheduling model using the joint state data and environment data obtained from the sample data pool corresponding to the next cycle, until the number of cycles reaches a preset threshold, stop the iterative training, and obtain the final pod scheduling model.

[0145] In an optional implementation, the acquisition module 901 is specifically configured to:

[0146] In the current loop, the initial joint state data and environment data are obtained, wherein the initial joint state data is the initial remaining resources of each of the multiple nodes in the k8s cluster in the current loop;

[0147] In the initial data generation cycle, the initial joint state data and environment data are input into the initial pod scheduling model to generate the joint action corresponding to the next data generation cycle and the joint reward corresponding to the current data generation cycle;

[0148] The initial joint state data corresponding to the initial data generation cycle, the environment data, the joint action corresponding to the next data generation cycle, and the joint reward corresponding to the current data generation cycle constitute a four-tuple as a sample data, and add it to the sample data pool corresponding to the current cycle;

[0149] In the i-th data generation cycle, based on the joint action corresponding to the i-th generation cycle and the joint state data corresponding to the i-1-th data generation cycle, generate joint state data corresponding to the i-th data generation cycle, where i is a positive integer greater than or equal to 2 and i is incremented by 1;

[0150] Input the joint state data and environment data corresponding to the i-th data generation cycle into the initial pod scheduling model, obtain the joint action corresponding to the i+1-th data generation cycle, and the joint reward corresponding to the i-th data generation cycle;

[0151] The four-tuple consisting of the joint state data and environment data corresponding to the i-th data generation cycle, the joint reward corresponding to the i-th data generation cycle, and the joint action corresponding to the i+1-th data generation cycle is added as a sample data to the sample data pool corresponding to the current cycle;

[0152] The iteration stops when i is equal to the preset threshold, and all sample data in the sample data pool corresponding to the current loop are obtained.

[0153] In an optional implementation, the processing module 903 is specifically configured to: when the reward value of the joint reward is greater than or equal to a preset reward threshold, determine to stop iterative training in the current cycle;

[0154] and / or,

[0155] When it is determined that the training cycle statistics corresponding to the current training cycle reaches a preset cycle quantity value, it is determined that the current cycle stops iterative training.

[0156] In an optional implementation, the resource types of the remaining resources in each node include one or more of the following:

[0157] CPU resources, memory resources, disk resources, network resources, video memory resources, and GPU resources.

[0158] In an optional implementation, when the resource types include CPU resources, memory resources, disk resources, network resources, video memory resources, and GPU resources, the expression of the target pod request resource reward is as follows:

[0159] R score =w cpu *req cpu +w mem *req mem +w disk *req disk +w net *req net +w vimem *req vimem +w gpu *req gpu (Formula 18)

[0160] w cpu +w mem +w disk +w net +w vimem +w gpu =1 (Formula 19)

[0161] Among them, R scoreThe reward value for requesting resources for the target pod, req cpu Remaining CPU resources, reqme m For the remaining memory resources, reqdis k The remaining disk resources, t is the remaining network resources, req vimem For the remaining resources of video memory, req gpu is the remaining GPU resources, w cpu 、w mem 、w disk 、w net 、w vimem , and w gpu are the weight coefficients corresponding to each resource.

[0162] In an optional implementation, when the resource types include CPU resources, memory resources, disk resources, network resources, video memory resources, and GPU resources, the matching degree reward between the target pod and the candidate node is expressed by the following expression:

[0163]

[0164] Among them, R similar The reward value for the matching degree between the target pod and the candidate node. For joint actions, Pod i Schedule to N j Run the containerized application on the node, i is the i-th pod, j is the j-th node, and They represent the remaining CPU resources on the i-th pod and the remaining CPU resources on the j-th node respectively; and Represent the remaining memory resources on the i-th pod and the remaining memory resources on the i-th node respectively; and Represent the remaining disk resources on the i-th pod and the remaining disk resources on the i-th node respectively; and They represent the remaining network resources on the i-th pod and the remaining network resources on the j-th node respectively; and They represent the remaining video memory resources on the i-th pod and the remaining video memory resources on the j-th node respectively; and They represent the remaining GPU resources on the i-th pod and the remaining GPU resources on the j-th node respectively.

[0165] In an optional implementation, the scheduling similarity reward between nodes is expressed by the following expression:

[0166]

[0167] Among them, R attention is the scheduling similarity reward between the nodes. For the i-th node, Q is the scheduling similarity query vector, K T Schedule similarity feature representation for each node, which is used to calculate similarity with Q, and V scores the actual similarity of each node for the final weighted output, d k The dimension size of the scheduling similarity query vector.

[0168] In an optional implementation, the joint reward expression is as follows:

[0169] R=μ1R score +μ2R similar +μ3R attention (Formula 23)

[0170] Wherein, R is the reward value of the joint reward, μ1, μ2, and μ3 are pre-configured weight coefficients respectively.

[0171] The pod scheduling model building device in this embodiment is presented in the form of a functional module, where the module refers to an application specific integrated circuit (ASIC), a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0172] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0173] A pod scheduling model construction device provided by an embodiment of the present invention can allocate Pods more effectively by considering the remaining resources of each node in the K8s cluster, thereby maximizing the utilization of cluster resources. The model continuously learns how to make the best scheduling decision based on the current cluster state and the needs of the Pod to be scheduled through iterative training, thereby improving the scheduling efficiency of the Pod. In particular, in this application document, two nested loops are included, one of which is to set different iterative training cycles in each loop, and the other is to include multiple loops for iterative training, so that the iterative training process is richer, and the model finally obtained will be able to adapt to different situations. The joint reward mechanism takes into account multiple factors, including the target Pod request resource reward, the matching degree reward between the Pod and the candidate node, and the influence reward of the candidate node on other nodes, which helps the model to make more comprehensive and reasonable scheduling decisions. In particular, in this application, the influence of other nodes on a certain node is taken into account to improve the scheduling efficiency and resource utilization as much as possible. In addition, the model can adjust its behavior according to the reward corresponding to the current training cycle and the training cycle statistics, so that the model can adapt to the changing environment and needs. Through iterative training, the model can continuously learn new scheduling strategies from the sample data pool until the preset training times threshold is reached, ensuring that the model fully absorbs the information in the data during the training process. Through multiple iterative training and the acquisition of the final model, it can be ensured that the trained Pod scheduling model has high accuracy and reliability.

[0174] In summary, this method can ensure that the model allocates resources better and reduces resource waste. It can quickly find the right node to schedule the Pod and reduce scheduling delays. Through precise matching, it can increase the chances of successful Pod startup and operation. The model can adapt to different workloads and cluster states.

[0175] The embodiment of the present invention also provides a computer device having the above Fig. 9 The pod scheduling model shown builds the device.

[0176] See also Fig.10 , Fig.10 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Fig.10As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Fig.10 A processor 10 is taken as an example.

[0177] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include an integrated circuit. The integrated circuit may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0178] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0179] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of a computer device based on the presentation of a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0180] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0181] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Fig.10 The example of connecting through bus is taken in the following.

[0182] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0183] The embodiment of the present invention also provides a computer-readable storage medium. The method provided in the above embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or is implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0184] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0185] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for constructing a pod scheduling model, characterized in that: The method comprises: In the current training cycle in the current loop, obtain joint state data and environment data from the sample data pool corresponding to the current loop, wherein the joint state data is the remaining resources of each of the multiple nodes in the K8s cluster, and the environment data includes the multiple pods currently to be scheduled; Obtaining joint state data and environment data from a sample data pool corresponding to the current cycle and inputting them into a pod scheduling model, so as to iteratively train the pod scheduling model within the current cycle; When the pod scheduling model generates a joint reward corresponding to the current training cycle, determine whether to stop iterative training in the current cycle according to the joint reward and / or the training cycle statistics corresponding to the current training cycle, wherein the joint reward includes a target pod request resource reward, a matching degree reward between the target pod and the candidate node, and an inter-node scheduling similarity reward corresponding to the influence of the candidate node and nodes other than the candidate node in the K8s cluster on the candidate node, the target pod is any one of the multiple pods, and the candidate node is any one of the multiple nodes; When it is determined that the current cycle stops iterative training based on the joint reward and / or the training cycle statistics corresponding to the current training cycle, the next cycle is entered, and the joint status data and environmental data are obtained from the sample data pool corresponding to the next cycle to iteratively train the pod scheduling model until the number of cycles reaches a preset threshold, then the iterative training is stopped to obtain the final pod scheduling model.

2. The method according to claim 1, characterized in that The method of acquiring the sample data in the sample data pool corresponding to each cycle includes: In the current loop, initial joint state data and environment data are obtained, wherein the initial joint state data is the initial remaining resources of each of the multiple nodes in the k8s cluster in the current loop; In an initial data generation cycle, the initial joint state data and the environment data are input into an initial pod scheduling model to generate a joint action corresponding to a next data generation cycle and a joint reward corresponding to a current data generation cycle; The initial joint state data corresponding to the initial data generation cycle, the environmental data, the joint action corresponding to the next data generation cycle, and the joint reward corresponding to the current data generation cycle constitute a four-tuple as one of the sample data, and add it to the sample data pool corresponding to the current cycle; In the i-th data generation cycle, based on the joint action corresponding to the i-th generation cycle and the joint state data corresponding to the i-1-th data generation cycle, generate joint state data corresponding to the i-th data generation cycle, where i is a positive integer greater than or equal to 2 and i is incremented by 1; Input the joint state data and the environment data corresponding to the i-th data generation cycle into the initial pod scheduling model, obtain the joint action corresponding to the i+1-th data generation cycle, and the joint reward corresponding to the i-th data generation cycle; and adding a 4-tuple consisting of the joint state data corresponding to the i-th data generation cycle, the environmental data, the joint reward corresponding to the i-th data generation cycle, and the joint action corresponding to the i+1-th data generation cycle as one of the sample data to the sample data pool corresponding to the current cycle; The iteration is stopped when i is equal to a preset threshold, and all sample data in the sample data pool corresponding to the current cycle are obtained.

3. The method according to claim 1 or 2, characterized in that: When the pod scheduling model generates a joint reward corresponding to the current training cycle, determining whether to stop iterative training in the current cycle according to the joint reward and / or the training cycle statistics corresponding to the current training cycle specifically includes: When the reward value of the joint reward is greater than or equal to a preset reward threshold, determining that the current cycle stops iterative training; and / or, When it is determined that the training cycle statistics value corresponding to the current training cycle reaches a preset cycle quantity value, it is determined that the current cycle stops iterative training.

4. The method according to claim 1 or 2, characterized in that: The resource types of the remaining resources in each node include one or more of the following: CPU resources, memory resources, disk resources, network resources, video memory resources, and GPU resources.

5. The method according to claim 4, characterized in that When the resource types include CPU resources, memory resources, disk resources, network resources, video memory resources, and GPU resources, the expression of the target pod request resource reward is as follows: R score =w cpu *req cpu +w mem *req mem +w disk *req disk +w net * req net +w vimem *req vimem +w gpu *req gpu (Formula 1) w cpu +w mem +w disk +w net +w vimem +w gpu = 1 (Formula 2) Among them, R score The reward value for requesting resources for the target pod, req cpu The remaining CPU resources, req mem Remaining memory resources, req disk The remaining disk resources, req net is the remaining network resources, req vimem For the remaining resources of video memory, req gpu is the remaining GPU resources, w cpu 、w mem 、w disk 、w net 、w vimem , and w gpu are the weight coefficients corresponding to each resource.

6. The method according to claim 4, characterized in that When the resource types include CPU resources, memory resources, disk resources, network resources, video memory resources, and GPU resources, the matching degree reward between the target pod and the candidate node is expressed by the following expression: Among them, R similar The reward value for the matching degree between the target pod and the candidate node, For joint actions, Pod i Schedule to N j Run the containerized application on the node, i is the i-th pod, j is the j-th node, and They represent the remaining CPU resources on the i-th pod and the remaining CPU resources on the j-th node respectively; and Represent the remaining memory resources on the i-th pod and the remaining memory resources on the j-th node respectively; and They represent the remaining disk resources on the i-th pod and the remaining disk resources on the j-th node respectively; and They represent the remaining network resources on the i-th pod and the remaining network resources on the j-th node respectively; and They represent the remaining video memory resources on the i-th pod and the remaining video memory resources on the j-th node respectively; and They represent the remaining GPU resources on the i-th pod and the remaining GPU resources on the j-th node respectively.

7. The method according to claim 1 or 2, characterized in that: The scheduling similarity reward between nodes is expressed by the following expression: Among them, R attention is the scheduling similarity reward between the nodes. For the i-th node, Q is the scheduling similarity query vector, K T Schedule similarity feature representation for each node, which is used to calculate similarity with Q, and V scores the actual similarity of each node for the final weighted output, d k The dimension size of the scheduling similarity query vector.

8. A pod scheduling model construction device, characterized in that: The device comprises: An acquisition module, used to acquire joint state data and environment data from a sample data pool corresponding to the current cycle in a current training cycle in the current cycle, wherein the joint state data is the remaining resources of each of the multiple nodes in the K8s cluster, and the environment data includes multiple pods currently to be scheduled; A training module, used to obtain joint state data and environment data from a sample data pool corresponding to the current cycle and input them into a pod scheduling model, so as to iteratively train the pod scheduling model within the current cycle; A processing module is used to determine whether to stop iterative training in the current loop according to the joint reward and / or the training cycle statistics corresponding to the current training cycle when the pod scheduling model generates a joint reward corresponding to the current training cycle, wherein the joint reward includes a target pod request resource reward, a matching degree reward between the target pod and the candidate node, and an inter-node scheduling similarity reward corresponding to the influence of the candidate node and the nodes other than the candidate node in the K8s cluster on the candidate node, the target pod is any one of the multiple pods, and the candidate node is any one of the multiple nodes; when it is determined that the current loop stops iterative training, enter the next loop; The acquisition module is further used to acquire the joint state data and the environment data from the sample data pool corresponding to the next cycle; The training module is also used to iteratively train the pod scheduling model by obtaining joint status data and environment data from a sample data pool corresponding to the next cycle, until the number of cycles reaches a preset number threshold, then stop the iterative training and obtain the final pod scheduling model.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the pod scheduling model construction method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the pod scheduling model construction method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Container cloud cluster resource utilization optimization method based on deep reinforcement learning

    CN112416578A

  • Modeling method and device of resource scheduling model, equipment and medium

    CN118672745A

  • Cluster scheduling method and device, storage medium and electronic equipment

    CN118672771A

  • Resource scheduling method and device, electronic equipment and storage medium

    CN118819806A

  • Container CPU resource scheduling and isolation method and apparatus, and storage medium and electronic device

    WO2023045467A1

Cited By

  • Model training method and device, container preheating method and device and electronic equipment

    CN120850052A