One-cloud multi-core heterogeneous resource hybrid scheduling method
By introducing resource pressure index and task affinity score, combined with Prometheus monitoring, and using reinforcement learning and multi-objective optimization algorithms to optimize resource scheduling in cloud computing environments, the problems of low resource utilization and opaque decision-making in traditional schedulers are solved, and efficient and transparent resource management is achieved.
Patent Information
- Application Number
- CN202510246738.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-04
AI Technical Summary
In traditional cloud computing environments, the scheduler schedules based on static resource indicators, resulting in low resource utilization, unable to adapt to different business needs, and scheduling decisions are not transparent and accurate enough.
Introduce resource pressure index and resource pressure trend index, combine node resource availability scores and task type affinity scores, dynamic scheduling is performed through reinforcement learning and multi-objective optimization algorithms, and use Prometheus monitoring solution to monitor resource status in real time to optimize resource allocation.
It improves resource utilization, enhances the transparency and verifiability of scheduling decisions, and improves the stability of the cluster and task execution efficiency.
Smart Images

Figure CN120263792A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of heterogeneous data processing, and in particular to a heterogeneous resource hybrid scheduling method for one cloud with multiple cores. Background Art
[0002] "One cloud with multiple cores" in the prior art refers to an architecture in a cloud computing environment that can support multiple different types of processor cores (i.e., "cores") to work together, which generally includes but is not limited to different types of computing resources such as CPUs, GPUs, FPGAs, and ASICs. The advantage of this architecture is that it can utilize the advantages of different types of processor cores for different computing tasks to achieve efficient computing.
[0003] The following are the disadvantages of the heterogeneous resource hybrid scheduling method of "one cloud with multiple cores" in the prior art:
[0004] Traditional schedulers may decide the scheduling of Pods only based on static resource metrics (CPU and memory), ignoring the resource usage trends and load conditions; traditional scheduling strategies lack flexibility and cannot adapt to different business requirements and resource environments; due to inaccurate scheduling decisions, some nodes in the cluster have excessive resources while other nodes have insufficient resources, thus reducing resource utilization; the traditional scheduling decision-making process may not be transparent enough to verify its accuracy.
[0005] The solution proposed by the present invention for the above-mentioned disadvantages is as follows:
[0006] By introducing the Resource Pressure Index (RPI) and the Resource Pressure Trend Index (RTI), and comprehensively considering the resource matching score and the load balancing score, the resource availability and load conditions of nodes can be evaluated more accurately; by implementing a scheduling strategy based on reinforcement learning and a multi-objective optimization algorithm, the node weights can be dynamically adjusted to balance the utilization of different resource types and business priorities, thereby improving the adaptability and flexibility of the scheduling strategy; by introducing the Prometheus monitoring solution, the resource status of nodes can be monitored in real time, and resources can be intelligently allocated according to resource requirements and cluster status, thereby improving resource utilization; by recording key information, including the Resource Availability Score (RAS) of nodes, the Task Type Affinity Score (TAS), and the weight factor, and providing traceability of the calculation process, the transparency and verifiability of scheduling decisions are ensured. Summary of the Invention
[0007] The present invention provides a heterogeneous resource hybrid scheduling method for one cloud with multiple cores to solve the above-mentioned existing technical problems.
[0008] The technical solution of the present invention is realized as follows: A heterogeneous resource hybrid scheduling method for one cloud with multiple cores, including:
[0009] S1. In the Kubernetes container orchestration system, the scheduler reorders the Pod queue waiting for scheduling according to the task types of the Pod container groups. When the scheduler scores the nodes to determine which node is most suitable for running the Pod, it will preferentially schedule tasks of the same type to the same node;
[0010] S2. TOPSIS is used to rank alternatives according to multiple criteria and is used to score the nodes in the cluster to determine which node is most suitable for running a new Pod. The scoring takes into account various attributes of the nodes, including resource availability and load;
[0011] S3. The job scheduler of "optimus" is used to optimize the scheduling of deep learning training tasks. The purpose of this scheduler is to minimize the time required for training tasks by intelligently allocating resources;
[0012] S4. During the Kubernetes scheduling process, in addition to the standard resource metrics CPU and memory, disk I / O, network I / O, and GPU resources are also introduced as additional evaluation metrics for scheduling decisions, during the node filtering (i.e., screening out nodes that are not suitable for running Pods) and priority calculation phases;
[0013] S5. The Prometheus open-source monitoring solution is used to collect and store metric data. It is mentioned that Prometheus is deployed in the cluster to monitor and collect resource utilization data of nodes and Pods. This data is used for scheduling decisions and performance analysis. The scheduling priority of each node is quantified according to the resource pressure index and the resource pressure trend index, and scheduling decisions are made accordingly.
[0014] Furthermore, in step S1, the score of the node Pod is calculated to verify that in step S1, the Kubernetes scheduler reorders and scores the nodes according to the Pod task type. The resource availability score (RAS) of the node, the task type affinity score (TAS) of the node, and the weight factors (α and β). Through this information, the score can be recalculated and verified whether it is consistent with the previous calculation result, ensuring the transparency and verifiability of the scheduling strategy during the calculation process.
[0015] Furthermore, in step S2, the nodes in the cluster are scored to determine which node is most suitable for running a new Pod. The TOPSIS multi-criteria decision-making method is adopted. The alternatives are ranked by calculating the distance between each alternative and the ideal solution. By determining the evaluation criteria and weights, constructing the decision matrix, standardizing the decision matrix, weighting the standardized decision matrix, determining the ideal solution and the negative ideal solution, calculating the distance from each node to the ideal solution and the negative ideal solution, calculating the relative closeness of each node to the ideal solution, and then ranking.
[0016] Furthermore, in step S3, resources are allocated to minimize the time required for training tasks. The intelligent scheduler allocates resources, uses the model to calculate the resource matching score for each node considering two main factors: resource requirements and node load, calculates the load balancing score, and the comprehensive score is the product of the resource matching score and the load balancing score, representing the comprehensive fitness of the node. The scheduling policy schedules according to the comprehensive score, updates the node status, and repeats the steps until all tasks are scheduled. The intelligent scheduler makes scheduling decisions by calculating the resource matching score and the load balancing score.
[0017] Furthermore, in step S4, disk I / O, network I / O, and GPU resources are introduced as additional evaluation metrics for scheduling decisions. By calculating the comprehensive score of each node, it is determined which node is more suitable for running the new Pod, ensuring that both nodes can meet the resource requirements of the Pod. This is done by comparing the requirements of the Pod and the remaining resources of the node. Then, the priority of each node is determined by calculating the resource matching score and the load balancing score. The final score indicates a more suitable choice for running the Pod, demonstrating that the scheduling decision framework can select the best node according to resource requirements and cluster status.
[0018] Furthermore, in step S5, the Prometheus open-source monitoring solution is used to enhance the resource scheduling policy of the Kubernetes cluster. The comprehensive score Si of all nodes is calculated, and then the node with the highest comprehensive score is selected for Pod scheduling.
[0019] Beneficial Effects
[0020] Improvement in resource utilization: By comprehensively considering various resource types such as CPU, memory, disk I / O, network I / O, and GPU, the present invention ensures the optimal allocation of resources, thereby improving the overall resource utilization and reducing resource waste.
[0021] Optimization of scheduling decisions: The scheduling policy based on reinforcement learning and multi-objective optimization algorithms can dynamically adjust the node weights according to real-time data and business requirements, making the scheduling decisions more accurate and intelligent.
[0022] Enhancement of system transparency and verifiability: The method of the present invention records detailed scheduling decision information, including resource scores and weight factors, making the scheduling process more transparent and traceable, facilitating monitoring and troubleshooting.
[0023] Improvement in cluster stability: By real-time monitoring and predicting resource usage trends, the present invention can prevent resource bottlenecks and overload situations, thereby improving the stability and reliability of the cluster.
[0024] Improvement in task execution efficiency: Through intelligent scheduling and resource optimization, the present invention can reduce the task execution time, enabling deep learning training tasks to be completed faster, thereby improving the efficiency of data processing and analysis.
[0025] Working principle:
[0026] Evaluation of nodes and Pods First, according to the task type of the Pod and the resource availability of the node, the scores of the nodes are calculated using the preset weight factors α and β to evaluate the suitability of the nodes for specific Pods.
[0027] TOPSIS multi-criteria decision-making After determining the evaluation criteria and weights, the TOPSIS method is used to construct a decision matrix. By calculating the distances of each node to the ideal solution and the negative ideal solution, the nodes are scored, and the node most suitable for running the new Pod is selected.
[0028] Intelligent scheduling of resources By analyzing the characteristics of the training tasks, a resource demand model is established, the resource status is monitored in real time, and based on factors such as task priority, resource efficiency, and load balancing, an intelligent scheduling algorithm is used for resource allocation.
[0029] Calculation of comprehensive scores Disk I / O, network I / O, and GPU resources are introduced as additional evaluation indicators for scheduling decisions. By calculating the comprehensive scores of each node, the node most suitable for running the new Pod is selected.
[0030] Prometheus monitoring integration Utilize the Prometheus monitoring solution, deploy and configure the Prometheus server, achieve monitoring and data collection of cluster resources, and integrate the data into the scheduler to support scheduling decisions.
[0031] Testing and verification Through simulation testing and performance analysis, verify the robustness of the scheduler and the impact of the scheduling strategy on the cluster performance.
[0032] Resource pressure assessment and prediction Use a time series analysis model to predict resource utilization, evaluate the resource pressure index (RPI) and the resource pressure trend index (RTI) to quantify the scheduling priority of nodes.
[0033] Scheduling decision Select the optimal node for Pod scheduling according to the comprehensive score to ensure that the Pod can run on the most suitable node, optimizing the resource utilization and performance of the cluster. Description of the drawings
[0034] Figure 1 It is a structural block diagram of a heterogeneous resource hybrid scheduling method with one cloud and multiple cores in an embodiment of the present invention.
[0035] Figure 2It is a step block diagram of a heterogeneous resource hybrid scheduling method with one cloud and multiple cores in an embodiment of the present invention. Detailed implementation manners
[0036] In order to make the objectives and advantages of the present invention clearer, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0037] The preferred implementation methods of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these implementation methods are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0038] It should be noted that when an element is considered to be "connected" to another element, it can be directly connected to the other element or connected to the other element through an intermediate element. In addition, in the following embodiments, "connection", if there is a transmission of electrical signals or data between the connected objects, should be understood as "electrical connection", "communication connection", etc.
[0039] As used herein, the singular forms "a", "an" and "the" may also include the plural forms unless the context clearly dictates otherwise. It should also be understood that the terms "comprises / include" or "has" etc. specify the presence of the stated features, wholes, steps, operations, components, parts or combinations thereof, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, components, parts or combinations thereof. At the same time, the term "and / or" used in this specification includes any and all combinations of the related listed items.
[0040] Please refer to Figure 1 - Figure 2 As shown, a heterogeneous resource hybrid scheduling method with one cloud and multiple cores includes:
[0041] S1. In the Kubernetes container orchestration system, the scheduler reorders the waiting Pod queue according to the task types of the Pod container groups. When the scheduler scores the nodes to determine which node is most suitable for running the Pod, it will preferentially schedule tasks of the same type to the same node;
[0042] S2. TOPSIS is used to rank alternatives according to multiple criteria and is used to score the nodes in the cluster to determine which node is most suitable for running a new Pod. The scoring takes into account various attributes of the nodes, including resource availability and load;
[0043] S3. The job scheduler of "optimus" is used to optimize the scheduling of deep learning training tasks. The purpose of this scheduler is to minimize the time required for training tasks by intelligently allocating resources;
[0044] S4. During the Kubernetes scheduling process, in addition to the standard resource metrics of CPU and memory, disk I / O, network I / O, and GPU resources are also introduced as additional evaluation metrics for scheduling decisions during the node filtering (i.e., screening out nodes that are not suitable for running Pods) and priority calculation phases;
[0045] S5. Use the Prometheus open-source monitoring solution to collect and store metric data. It is mentioned that Prometheus is deployed in the cluster to monitor and collect the resource utilization data of nodes and Pods. This data is used for scheduling decisions and performance analysis. The scheduling priority of each node is quantified according to the resource pressure index and the resource pressure trend index, and scheduling decisions are made accordingly.
[0046] Specifically, in the S1 step, calculate the score of the node Pod, and verify the calculation process in the S1 step where the Kubernetes scheduler reorders according to the Pod task type and scores the nodes:
[0047] Preset scenario parameters:
[0048] Node Resource Availability Score (RAS): between 0 and 10, the larger the value, the richer the resources;
[0049] Node Task Type Affinity Score (TAS): between 0 and 5, the larger the value, the higher the affinity;
[0050] Weight factors: resource availability weight factor (α) and task type affinity weight factor (β);
[0051] Among them, α = 0.7 and β = 0.3;
[0052] Node score calculation formula:
[0053] The node score calculation formula is as follows:
[0054] NodeScore(N,P) = α × RAS(N,P) + β × TAS(N,P);
[0055] Among them: N represents the node, and P represents the Pod;
[0056] Node Task Type Affinity Score (TAS):
[0057] Preset the task type affinity score for TypeX as follows:
[0058] The affinity score of Node 1 for Type X: TAS(Node 1,Type X) = 3;
[0059] Affinity score of Node 2 for Type X: TAS(Node 2, Type X) = 2;
[0060] Calculation of node scores for PodA(TypeX):
[0061] Calculate the scores of Node1 and Node2 for PodA(TypeX):
[0062] NodeScore(Node 1, Pod A) = α
[0063] ×RAS(Node 1) + β×TAS(Node 1, Type X)
[0064] NodeScore(Node 1, Pod A) = 0.7×8 + 0.3×3
[0065] NodeScore(Node 1, Pod A) = 5.6 + 0.9
[0066] NodeScore(Node 1, Pod A) = 6.5;
[0067] NodeScore(Node 2, Pod A) = α×RAS(Node 2) + β×TAS(Node 2, Type X)
[0068] NodeScore(Node 2, Pod A) = 0.7×6 + 0.3×2
[0069] NodeScore(Node 2, Pod A) = 4.2 + 0.6
[0070] NodeScore(Node 2, Pod A) = 4.8;
[0071] According to the above calculations, the score of Node 1 (6.5) is higher than the score of Node 2 (4.8), so Pod A will be scheduled to Node1;
[0072] Verify that the calculation results are traceable:
[0073] The following information needs to be recorded: the resource availability score (RAS) of the node, the task type affinity score (TAS) of the node, and the weight factors (α and β). With this information, the scores can be recalculated to verify whether they are consistent with the previous calculation results, ensuring the transparency and verifiability of the scheduling policy during the calculation process;
[0074] The physical meanings of the characters in the calculation formula during the above calculation process are:
[0075] N: Represents Node. In Kubernetes, a node is a physical server, virtual machine, or cloud instance, which is the computing resource for running Pods;
[0076] P: Represents Pod, which is the basic unit of scheduling in Kubernetes. It contains one or more containers and is a component of the application;
[0077] RAS: Represents Resource Availability Score, which is a quantified score indicating the abundance of available resources on a node, including CPU, memory, and storage;
[0078] TAS: Represents Task Affinity Score, which is a quantified score indicating the affinity or preference of a node for a specific type of task, based on historical running data or specific policies;
[0079] α: Represents Resource Availability Weight Factor, which is a value between 0 and 1, indicating the relative importance of resource availability in the overall score calculation;
[0080] β: Represents Task Affinity Weight Factor, which is also a value between 0 and 1, indicating the relative importance of task affinity in the overall score calculation;
[0081] In the formula NodeScore(N,P) = α × RAS(N,P) + β × TAS(N,P):
[0082] NodeScore(N,P): Represents the total score of node N for Pod P, which is calculated based on the resource availability score, task affinity score, and their respective weight factors;
[0083] This score is used to determine which node is most suitable for running a specific Pod during the Kubernetes scheduling decision-making process. In this way, the scheduler can optimize resource allocation and task scheduling based on predefined rules and weights.
[0084] Specifically, in step S2, the nodes in the cluster are scored to decide which node is most suitable for running a new Pod:
[0085] After completing step S1, the TOPSIS (Technique for Order Preference by Similarittto Ideal Solution) method is used to implement step S2 to score the nodes in the cluster. TOPSIS is a multi-criteria decision-making method that ranks alternatives by calculating the distance of each alternative from the ideal solution. Detailed steps:
[0086] 1. Determine evaluation criteria and weights. First, the criteria for evaluating nodes need to be determined. These criteria include resource availability, load, network latency, and storage capacity. For each criterion, a weight also needs to be assigned to indicate its relative importance in the decision-making. Presumably, there are m evaluation criteria, and w1, w2,..., wm are the corresponding weights, and
[0087] 2. Construct the decision matrix. Construct an n×m decision matrix A, where n is the number of nodes and m is the number of evaluation criteria. The element aij of the matrix represents the performance of the i-th node on the j-th criterion.
[0088] 3. Standardize the decision matrix. Since different criteria have different dimensions, the decision matrix needs to be standardized. Use the following formula to calculate the standardized matrix R:
[0089]
[0090] 4. Weighted standardized decision matrix. Multiply each criterion by its corresponding weight to obtain the weighted standardized decision matrix V: v ij = w j × r ij ;
[0091] 5. Determine the ideal solution and the negative ideal solution. The ideal solution A+ is the optimal value for each criterion, and the negative ideal solution A- is the worst value for each criterion;
[0092] A+ = (v1+, v2+,..., vm+)
[0093] A- = (v1-, v2-,..., vm-)
[0094] 6. Calculate the distances. Calculate the distances of each node from the ideal solution and the negative ideal solution:
[0095]
[0096] 7. Calculate the relative closeness. Finally, calculate the relative closeness Ci of each node to the ideal solution:
[0097]
[0098] 8. Sorting and selection: Nodes are sorted according to the value of Ci. The higher the value of Ci, the closer the node is to the ideal solution, and thus it is more suitable for running new Pods.
[0099] By using the TOPSIS method through the above steps to score the nodes in the cluster and determine which node is most suitable for running new Pods, this process takes into account various attributes of the nodes, including resource availability and load, making the scheduling decision more comprehensive and accurate.
[0100] In the calculation formulas in the above steps, each character and symbol has a specific physical meaning. The following are the explanations for them:
[0101] m: The number of evaluation criteria;
[0102] n: The number of nodes in the cluster;
[0103] w1, w2,..., wm: The weights corresponding to each evaluation criterion;
[0104] aij: The element in the decision matrix A, representing the performance of the i-th node on the j-th criterion;
[0105] R: The standardized decision matrix;
[0106] rij: The element in the standardized matrix R, representing the standardized performance of the i-th node on the j-th criterion;
[0107] V: The weighted standardized decision matrix;
[0108] vij: The element in the weighted standardized decision matrix V, which is the product of the standardized performance rij and the weight wj;
[0109] A+: The ideal solution, representing the optimal value for each criterion;
[0110] A-: The negative ideal solution, representing the worst value for each criterion;
[0111] vj+: The value of the j-th criterion in the ideal solution A+;
[0112] vj-: The value of the j-th criterion in the negative ideal solution A-;
[0113] Di+: The Euclidean distance from the i-th node to the ideal solution A+;
[0114] Di-: The Euclidean distance from the i-th node to the negative ideal solution A-;
[0115] Ci: The relative closeness of the i-th node to the ideal solution, a numerical value between 0 and 1, representing the degree to which the node is close to the ideal solution;
[0116] These symbols and formulas together form the basis of the TOPSIS method, which is used to rank alternatives (in this case, nodes in a cluster) according to multiple criteria to decide which node is most suitable for running a new Pod.
[0117] Specifically, step S3 allocates resources to minimize the time required for the training task:
[0118] After completing step S2, there are already results of scoring nodes in the cluster based on multiple criteria. In step S3, resources are intelligently allocated to minimize the training time. First, analyze the characteristics of the training task to determine the key characteristics of the deep learning training task, including compute-intensive, I / O-intensive, memory requirements, and data locality. Then, establish a resource demand model, including the resource requirements for CPU, GPU, memory, storage, and network. Next, monitor the resource status in real time using the cluster monitoring system to monitor the node resource status in real time, including resource utilization rate, load, and network latency. Then, the intelligent scheduling algorithm considers the following factors: task priority, allocating resources according to the importance and urgency of the task; resource efficiency, selecting the node that can complete the task fastest and has the highest resource efficiency; resource balance, maintaining the cluster load balance by avoiding overloading some nodes while other nodes are idle; data locality, scheduling tasks on the node where the data is located to reduce data transfer time; implementing scheduling strategies, including dynamic resource allocation, resource reservation, and backfill scheduling. Dynamic resource allocation dynamically adjusts resource allocation according to the task progress and resource requirements. Resource reservation reserves resources for critical tasks to ensure they can be completed in a timely manner. Backfill scheduling temporarily schedules low-priority tasks while waiting for high-priority task resources to be released. Execute the scheduling strategy. According to the decision of the intelligent scheduling algorithm, schedule the training task to the node with the highest score, ensuring that the execution of the scheduling decision does not violate the resource limits and policies of the cluster. Optimize and adjust by monitoring the task execution, collecting performance data, analyzing the scheduling results, identifying bottlenecks and optimization points, and adjusting the algorithm according to the feedback to further improve resource utilization and reduce training time. Continuously iterate by continuously experimenting and adjusting to continuously improve the strategy. Through these steps, the "Optimus" job scheduler in step S3 can intelligently allocate resources, thereby minimizing the time required for the deep learning training task.
[0119] The following calculation process illustrates how the intelligent scheduler allocates resources, using a model that considers two main factors: resource requirements and node load.
[0120] Preset three training tasks (Task A, Task B, Task C) and two cluster nodes (NodeX, Node Y). The resources required for each task are: 2 GPUs and 8GB of memory. The node resources are: each node has 4 GPUs and 16GB of memory.
[0121] Current status of nodes: Node X: current load is 60%, remaining 2 GPUs, 8GB of memory, Node Y: current load is 40%, remaining 2 GPUs, 8GB of memory;
[0122] Calculation steps:
[0123] 1. Calculate the resource matching score for each node for each task. The higher the score, the more the node can meet the requirements of the task; Resource matching score = (Remaining GPUs / GPUs required for the task) * (Remaining memory / Memory required for the task);
[0124] For Task A:
[0125] Node X score = (2 / 2) * (8 / 8) = 1
[0126] Node Y score = (2 / 2) * (8 / 8) = 1
[0127] 2. Calculate the load balancing score. Calculate the load balancing score for each node. The higher the score, the lower the load of the node; Load balancing score = (1 - Current load ratio);
[0128] For Node X and Node Y:
[0129] Node X score = 1 - 0.6 = 0.4
[0130] Node Y score = 1 - 0.4 = 0.6
[0131] 3. The comprehensive score is the product of the resource matching score and the load balancing score, representing the comprehensive suitability of the node. Comprehensive score = Resource matching score * Load balancing score;
[0132] For Task A:
[0133] Node X comprehensive score = 1 * 0.4 = 0.4
[0134] Node Y comprehensive score = 1 * 0.6 = 0.6
[0135] 4. The scheduling policy schedules Task A to the node with the highest score according to the comprehensive score. Task A is scheduled to Node Y;
[0136] 5. Update the node status. After scheduling, update the status of Node Y. Node Y: current load is 60%, remaining 0 GPUs, 8GB of memory;
[0137] 6. Repeat the steps. Repeat the above steps for Task B and Task C until all tasks are scheduled;
[0138] Through the above calculation process, the intelligent scheduler makes scheduling decisions by calculating resource matching scores and load balancing scores. In applications, the scheduler calculates based on, but is not limited to, data locality, task priority, and network latency.
[0139] Specifically, step S4 will introduce disk I / O, network I / O, and GPU resources as additional evaluation metrics for scheduling decisions. The following is the calculation and verification process:
[0140] Preset a Pod with the following resource requirements:
[0141] CPU requirement: 2 cores
[0142] Memory requirement: 4 GB
[0143] GPU requirement: 1
[0144] Disk I / O requirement: 50 MB / s
[0145] Network I / O requirement: 100 Mbps
[0146] Node resource situation:
[0147] Preset that there are two nodes (Node X and Node Y) in the cluster, and their resource situations are as follows:
[0148] Node X:
[0149] Total CPU: 8 cores, current load: 1 core
[0150] Total memory: 16 GB, current load: 2 GB
[0151] Total GPU: 2, current load: 1
[0152] Total disk I / O: 100 MB / s, current load: 20 MB / s
[0153] Total network I / O: 1 Gbps, current load: 200 Mbps
[0154] Node Y:
[0155] Total CPU: 4 cores, current load: 1 core
[0156] Total memory: 8 GB, current load: 1 GB
[0157] Total GPU: 1, current load: 0
[0158] Total disk I / O: 80 MB / s, current load: 10 MB / s
[0159] Total network I / O: 500 Mbps, current load: 50 Mbps
[0160] Node filtering:
[0161] Node X:
[0162] Total CPU: 8 cores, current load: 1 core
[0163] Total memory: 16 GB, current load: 2 GB
[0164] Total GPUs: 2, current load: 1
[0165] Total disk I / O: 100 MB / s, current load: 20 MB / s
[0166] Total network I / O: 1 Gbps, current load: 200 Mbps
[0167] Node Y:
[0168] Total CPU: 4 cores, current load: 1 core
[0169] Total memory: 8 GB, current load: 1 GB
[0170] Total GPUs: 1, current load: 0
[0171] Total disk I / O: 80 MB / s, current load: 10 MB / s
[0172] Total network I / O: 500 Mbps, current load: 50 Mbps
[0173] Priority calculation:
[0174] Node X:
[0175] Resource matching score:
[0176] CPU resource matching score = 7 cores / 2 cores = 3.5
[0177] Memory resource matching score = 14 GB / 4 GB = 3.5
[0178] GPU resource matching score = 1 / 1 = 1
[0179] Disk I / O resource matching score = 80 MB / s / 50 MB / s = 1.6
[0180] Network I / O resource matching score = 800 Mbps / 100 Mbps = 8
[0181] Resource matching score = 3.5 * 3.5 * 1 * 1.6 * 8 ≈ 193.6
[0182] Load balancing score:
[0183] Current load ratio = (1 core + 2GB + 1 GPU + 20MB / s + 200Mbos) / (8 cores + 16GB + 2 GPUs + 100MB / s + 1Gbps) ≈ 0.125
[0184] Load balancing score = 1 - 0.125 = 0.875
[0185] Composite score = Resource matching score * Load balancing score ≈ 193.6 * 0.875 ≈ 169.15
[0186] Node Y:
[0187] Resource matching score:
[0188] CPU resource matching score = 3 cores / 2 cores = 1.5
[0189] Memory resource matching score = 7GB / 4GB = 1.75
[0190] GPU resource matching score = 1 / 1 = 1
[0191] Disk I / O resource matching score = 70MB / s / 50MB / s = 1.4
[0192] Network I / O resource matching score = 450Mbps / 100Mbps
[0193] Resource matching score = 1.5 * 1.75 * 1 * 1.4 * 4.5 ≈ 16.4375
[0194] Load balancing score:
[0195] Current load ratio = (1 core + 1GB + 0 GPUs + 10MB / s + 50Mbps) / (4 cores + 8GB + 1 GPU + 80WB / s + 500Mbps) ≈ 0.125
[0196] Load balancing score = 1 - 0.125 = 0.875
[0197] Composite score = Resource matching score * Load balancing score ≈ 16.4
[0198] Furthermore, by calculating the composite score of each node to determine which node is more suitable for running the new Pod, the composite score reflects the ability of the node to maintain good load balancing while meeting the resource requirements of the Pod;
[0199] Through calculation, the following composite scores are obtained:
[0200] The composite score of Node X ≈ 169.15
[0201] The comprehensive score of Node Y ≈ 16.4375
[0202] The higher the score of resource matching, the better the matching degree between the remaining resources of the node and the Pod requirements. The matching degree of Node X in all resource types is higher than that of Node Y;
[0203] The higher the score of load balancing, the more balanced the load of the node. In the above content, the load balancing scores of the two nodes are the same because their current load ratios are the same;
[0204] The comprehensive score of Node X is much higher than that of Node Y, which means that Node X can not only better meet the resource requirements of the Pod, but also maintain a better load balance;
[0205] For the above calculation content, first, ensure that both nodes can meet the resource requirements of the Pod, which is completed by comparing the requirements of the Pod and the remaining resources of the node. Then, determine the priority of each node by calculating the resource matching score and the load balancing score, which reflects that the scheduler considers not only resource availability but also how to maintain the balanced use of resources in the whole cluster. The final score shows that Node X is a more suitable choice to run the Pod, proving that the scheduling decision framework can select the best node according to resource requirements and cluster status.
[0206] Specifically, in the S5 step, the Prometheus open-source monitoring solution is used:
[0207] The S5 step uses the Prometheus open-source monitoring solution to enhance the resource scheduling strategy of the Kubernetes cluster. The following are the detailed steps and methods:
[0208] Deploying and Configuring the Prometheus Monitoring Architecture: For the highly available deployment of the Prometheus server, a multi-replica deployment mode is adopted. The Prometheus server is deployed in the Kubernetes cluster through StatefulSets to ensure the stability of the monitoring system. Configure a long-term storage solution based on Thanos to achieve cross-cluster monitoring data aggregation and query; Define complex Recording Rules and Alerting Rules in prometheus.yml for Prometheus configuration to optimize query performance and alert efficiency. Implement file-based service discovery and automatically update monitoring targets through a configuration management tool (Ansible); Monitoring Target Configuration Node Monitoring: In addition to basic metrics, also capture kernel parameters, file system mount points, network device status metrics. Pod Monitoring: Integrate the Prometheus Adapter through a custom Metrice API to achieve HPA (Horizontal Pod Autoscaler) based on custom metrics; Deploy a monitoring agent Deploy blackbox Exporter on nodes to monitor the availability and response time of external services, and use Pushgateway to support the metric push of short-lived tasks; Alert Rule Configuration Write alert expressions using PromQL to detect abnormal patterns, such as mutations in time series and trend prediction. Configure the Silences and Inhibit Rules of Alertmanager to reduce false alarms and improve the accuracy of alerts;
[0209] Integrating Prometheus Data into the Scheduler: API Access Use Prometheus's HTTP API V2 to implement batch and streaming queries to improve the efficiency of the scheduler accessing metric data; Custom Advanced Scheduler Implement a pluggable architecture in the scheduler to support hot-swappable scheduling algorithms;
[0210] Implementing Scheduling Logic: Resource Utilization Metrics Introduce a time series analysis library (Facebook
[0211] Prophet or TensorFlow) to predict resource utilization, calculate the standard deviation and coefficient of variation based on historical data to evaluate the instability of resource usage; Dynamic Weight Adjustment and Optimization Implement a scheduling strategy based on reinforcement learning, adjust node weights through simulation and feedback, and adopt a multi-objective optimization algorithm (Pareto optimization) to balance the utilization of different resource types and business priorities
[0212] Testing and Validation: Design test cases with multiple scenarios and variables for simulation testing, inject faults through the Chaos Engineering tool (ChaosMonkey), and verify the robustness of the scheduler; Performance Analysis and Optimization Use Prometheus and Grafana for advanced data analysis, including performance bottleneck location, capacity planning, and trend analysis, and compare the impact of different scheduling strategies on cluster performance through A / B testing;
[0213] The calculation process is as follows:
[0214] Preset a cluster with N nodes, and each node has M types of resources, including CPU, memory, disk I / O, network I / O, and GPU;
[0215] Resource Pressure Index (RPI), which is used to quantify the pressure level of each type of resource on each node. For the i-th node and the j-th resource, define the Resource Pressure Index RPI ij as follows
[0216]
[0217] where U ij is the current utilization rate of resource j on node i, L ij is the minimum utilization rate threshold of resource j on node i, and C ij is the total capacity of resource j on node i;
[0218] The resource prediction model uses the time series prediction model ARIMA to predict resource utilization:
[0219]
[0220] where is the predicted utilization rate of resource j on node i at time t + T, Uij(t) is the actual utilization rate of resource j on node i at time t, f is the prediction function, and θ is the model parameter;
[0221] Data calculation process:
[0222] 1. Data preprocessing: For each type of resource of each node, use moving window average or exponential smoothing to downsample and aggregate the data:
[0223]
[0224] where U` ij (t) is the downsampled utilization rate of resource j on node i at time t, and k is the size of the moving window;
[0225] 2. Resource pressure assessment: Calculate the RPI of each type of resource of each node:
[0226]
[0227] 3. The resource prediction uses the model to calculate the resource utilization rate in the future T time:
[0228]
[0229] 4. The resource pressure trend analysis calculates the resource pressure trend index RTI:
[0230]
[0231] 5. The comprehensive score calculation defines a comprehensive score function S i to quantify the scheduling priority of the node:
[0232]
[0233] where Si is the comprehensive score of node i, used to quantify the scheduling priority of the node, wj is the weight of resource j, and a and b are adjustment coefficients used to balance the influence of RPI and RTI;
[0234] 6. The scheduling decision selects the node with the highest comprehensive score for Pod scheduling:
[0235]
[0236] where Si is the comprehensive score of node i, used to quantify the scheduling priority of the node, and Scheduled Node is the finally selected node for running the Pod;
[0237] Example calculation: Assume there are 3 nodes, considering two resources, CPU and memory. The calculation process is as follows:
[0238] Data preprocessing: For the CPU resource of node 1, the preset utilization rates in the past 5 minutes are 0.6, 0.7, 0.65, 0.8, 0.75 respectively; then the downsampled utilization rate is:
[0239]
[0240] Resource pressure assessment: Assume the lowest utilization rate threshold of the CPU resource of node 1 is 0.2 and the total capacity is 1, then the RPI is:
[0241]
[0242] Resource prediction: Assume the prediction model predicts that the CPU utilization rate of node 1 after 5 minutes is 0.85, then the RTI is:
[0243]
[0244] Comprehensive score calculation: Preset WCPU The weights for the CPU and memory are CPU = 0.6 and W MEM = 0.4, and the adjustment coefficients are a = 0.7 and b = 0.3. Using these presets, the comprehensive score is calculated as follows:
[0245] For the memory resources of Node 1, the preset utilization rates for the past 5 minutes are [0.5, 0.55, 0.6, 0.58, 0.53], then the downsampled utilization rate is:
[0246]
[0247] The minimum utilization rate threshold for the memory resources of Node 1 is 0.1, and the total capacity is 2, then the RPI is:
[0248]
[0249] The prediction model predicts that the memory utilization rate of Node 1 after 5 minutes will be 0.65, then the RTI is:
[0250]
[0251] Calculate the comprehensive score S1 of Node 1:
[0252] S1 = W CPU ·(a·RPI 1CPU + b·RTI 1CPU ) + W MEM ·(a·RPI 1MEM + b·RTI 1MEM )
[0253] S1 = 0.6·(0.7·0.5 + 0.3·0.03) + 0.4·(0.7·0.227 + 0.3·0.0114)
[0254] S1 = 0.6·(0.35 + 0.009) + 0.4·(0.1589 + 0.00342)
[0255] S1 = 0.6·0.359 + 0.4·0.16232
[0256] S1 = 0.2154 + 0.064928
[0257] S1 = 0.280328
[0258] The scheduling decision repeats the above steps to calculate the comprehensive scores Si of all nodes. Then, the node with the highest comprehensive score is selected for Pod scheduling. The score of Node 1, S1 = 0.280328, is the highest. So the scheduling decision will be: ScheduledNode = Node 1.
[0259] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easily understood by those skilled in the art that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
[0260] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and scope of the present invention shall be included within the protection scope of the present invention.
Claims
1. A heterogeneous resource hybrid scheduling method with one cloud and multiple cores, characterized in that: S1. In the Kubernetes container orchestration system, the scheduler reorders the waiting Pod queue according to the task types of the Pod container groups. When the scheduler scores the nodes to determine the most suitable node to run the Pod, it will preferentially schedule tasks of the same type to the same node; S2. TOPSIS is used to rank alternative solutions according to multiple criteria, score the nodes in the cluster, and determine the nodes suitable for running new Pods; S3. The job scheduler of optimus is used to optimize the scheduling of deep learning training tasks, and minimize the time required for training tasks by intelligently allocating resources; S4. During the Kubernetes scheduling process, disk I / O, network I / O, and GPU resources are introduced as additional evaluation metrics for scheduling decisions, during the node filtering and priority calculation phases; S5. Utilize the Prometheus open-source monitoring solution to collect and store metric data. Deploy Prometheus in the cluster to monitor and collect resource utilization data of nodes and Pods for scheduling decisions and performance analysis. Quantify the scheduling priority of each node according to the resource pressure index and the resource pressure trend index, and make scheduling decisions accordingly.
2. The heterogeneous resource hybrid scheduling method with one cloud and multiple cores according to claim 1, characterized in that: In the S1 step, calculate the score of the node Pod, verify that in the S1 step, the Kubernetes scheduler reorders and scores the nodes according to the Pod task type, the resource availability score RAS of the node, the task type affinity score TAS of the node, the weight factors α and β, to recalculate the score, and verify the calculation result.
3. A heterogeneous resource hybrid scheduling method with one cloud and multiple cores according to claim 1, characterized in that: In the S2 step, score the nodes in the cluster, adopt the TOPSIS multi-criteria decision-making method, rank the alternative solutions by calculating the distance between each alternative solution and the ideal solution, construct a decision matrix by determining the evaluation criteria and weights, standardize the decision matrix, the weighted standardized decision matrix, determine the ideal solution and the negative ideal solution, calculate the distance from each node to the ideal solution and the negative ideal solution, calculate the relative closeness of each node to the ideal solution, and finally rank them.
4. A heterogeneous resource hybrid scheduling method of one cloud with multiple cores according to claim 1, characterized in that: In the S3 step, allocate resources to minimize the time required for training tasks. The intelligent scheduler allocates resources, calculates the resource matching score of each node for each task, calculates the load balancing score, and the comprehensive score is the product of the resource matching score and the load balancing score, representing the comprehensive fitness of the node. The scheduling strategy schedules according to the comprehensive score, updates the node status, and repeats the steps until all tasks are scheduled. The intelligent scheduler makes scheduling decisions by calculating the resource matching score and the load balancing score.
5. A heterogeneous resource hybrid scheduling method of one cloud with multiple cores according to claim 1, characterized in that: Step S4 introduces disk I / O, network I / O, and GPU resources as additional evaluation metrics for scheduling decisions. It determines the node to run the new Pod by calculating the comprehensive score of each node, which is completed by comparing the requirements of the Pod and the remaining resources of the node. Then, it determines the priority of each node by calculating the resource matching score and the load balancing score. The scheduling decision framework can select the best node based on resource requirements and cluster status.
6. A heterogeneous resource hybrid scheduling method with one cloud and multiple cores according to claim 1, characterized in that: Step S5 uses the Prometheus open-source monitoring solution to enhance the resource scheduling strategy of the Kubernetes cluster, calculates the comprehensive scores of all nodes, and selects the node with the highest comprehensive score for Pod scheduling.
Citation Information
Cited By
Cloud platform-based computing power resource dynamic scheduling and monitoring method
CN121070513A