Intelligent scheduling method of container arrangement system, intelligent scheduling system, equipment, medium and product
By acquiring cluster operation data from the container orchestration system and using an AI inference engine, candidate nodes are selected and scheduling decisions are optimized, solving the problem of low scheduling efficiency in existing technologies and achieving efficient resource utilization and stable service quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-03-10
AI Technical Summary
Existing container orchestration and scheduling solutions are difficult to adapt to the complex business needs under single-cluster and multi-cluster architectures, resulting in low scheduling efficiency, fluctuating service quality, and wasted resources.
By acquiring cluster operation data from the container orchestration system, using an AI inference engine and preset scheduling strategies to filter candidate nodes, and combining the feature information of the Pods to be scheduled, the scheduling decision model is optimized to form a scheduling closed loop, achieving dynamic adaptation and optimization.
It improves scheduling efficiency, reduces resource waste, and ensures the stability and adaptability of service quality, making it suitable for both single-cluster and multi-cluster architectures.
Smart Images

Figure CN121636061A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of container orchestration technology, and more specifically, relates to an intelligent scheduling method, device, equipment, medium and product for a container orchestration system. Background Technology
[0002] With the rapid popularization of cloud computing, microservice architecture and multi-tenant deployment mode, container orchestration system has become the infrastructure supporting the efficient operation of various businesses. Its scheduling capability directly determines the efficiency of resource utilization, the stability of business operation and the level of service quality.
[0003] However, existing container orchestration and scheduling solutions still have many technical bottlenecks, making it difficult to adapt to the complex business needs under single-cluster and multi-cluster architectures, ultimately resulting in problems such as low scheduling efficiency, fluctuating service quality, and waste of resources. Summary of the Invention
[0004] The purpose of this application is to provide an intelligent scheduling method, device, equipment, medium and product for a container orchestration system, aiming to solve the technical problems that existing container orchestration and scheduling schemes are difficult to adapt to the complex business needs under single cluster and multi-cluster architectures, resulting in low scheduling efficiency, service quality fluctuations and resource waste.
[0005] To achieve the above objectives, according to the first aspect of this application, an intelligent scheduling method for a container orchestration system is provided, the method comprising: Acquire cluster operation data of member clusters in a single-cluster architecture or multi-cluster architecture deployed by a container orchestration system. The member cluster consists of multiple nodes, and each node is a server or computing unit that carries the container group Pods. The cluster operation data includes node resource utilization, Pod operation performance, application performance indicators, and load change data. In response to receiving a scheduling request from a Pod to be scheduled, multiple candidate nodes that meet the Pod's running constraints are selected from the member cluster. Based on the cluster operation data, the feature information of the Pod to be scheduled, and the preset scheduling strategy of calling the AI inference engine, the target node among the multiple candidate nodes is determined; Bind the Pod to be scheduled to the target node, and monitor the running status data of the Pod after it is bound to the target node; The decision model of the AI inference engine and the preset scheduling strategy are optimized based on the running status data to form a scheduling closed loop.
[0006] According to a second aspect of this application, an intelligent scheduling system is provided, characterized in that an intelligent scheduling method for executing the container orchestration system described in any one of the claims comprises: The processor is used to acquire cluster operation data of member clusters in a single-cluster architecture or multi-cluster architecture of a container orchestration system. The member cluster consists of multiple nodes, and the nodes are servers or computing units that carry container group Pods. The cluster operation data includes node resource utilization, Pod operation performance, application performance indicators and load change data. A scheduling plugin is used to respond to a scheduling request received from a Pod to be scheduled, to filter multiple candidate nodes from the member cluster that meet the Pod's running constraints; to determine a target node among the multiple candidate nodes based on the cluster running data, the feature information of the Pod to be scheduled, and a preset scheduling strategy invoked by the AI inference engine; to bind the Pod to be scheduled to the target node and monitor the running status data of the Pod after it is bound to the target node; and to optimize the decision model of the AI inference engine and the preset scheduling strategy based on the running status data to form a scheduling closed loop.
[0007] According to a third aspect of this application, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device causes the electronic device to perform the method as described in any one of the claims.
[0008] According to a fourth aspect of this application, a computer-readable storage medium is provided that stores a computer program, which, when executed by a processor, implements the method as described in any one of the claims.
[0009] According to a fifth aspect of this application, a computer program product is provided that, when run on an electronic device, causes the electronic device to perform the method described in any one of the first aspects above.
[0010] It is understandable that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating an intelligent scheduling method for a container orchestration system provided in an embodiment of this application; Figure 2This is a flowchart illustrating an optional intelligent scheduling method for a container orchestration system provided in an embodiment of this application. Figure 3 This is a flowchart illustrating an optional intelligent scheduling method for a container orchestration system provided in an embodiment of this application. Figure 4 This is a flowchart illustrating an optional intelligent scheduling method for a container orchestration system provided in an embodiment of this application. Figure 5 This is a flowchart illustrating an optional intelligent scheduling method for a container orchestration system provided in an embodiment of this application. Figure 6 This is a schematic diagram of the structure of an intelligent scheduling system provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0014] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0015] It should also be understood that, in the description of this application, unless otherwise stated, the " / " used in the specification and appended claims indicates that the related objects are in an "or" relationship. For example, A / B can mean A or B. The "and / or" in this application is merely a description of the relationship between the related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0016] Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, but are only used for distinguishing descriptions, and the terms "first" and "second" do not necessarily imply that they are different, nor should they be construed as indicating or implying relative importance.
[0017] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0018] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0019] With the rapid adoption of cloud computing, microservice architecture, and multi-tenant deployment models, container orchestration systems have become core infrastructure supporting the efficient operation of various businesses. Scheduling capabilities directly determine resource utilization efficiency, business operational stability, and service quality. However, existing container orchestration and scheduling solutions still suffer from numerous technical bottlenecks, making it difficult to meet the scheduling needs of complex scenarios. First, traditional scheduling schemes often rely on static rule configurations, executing scheduling based solely on Pod's basic resource requests (such as minimum CPU and memory requirements) and simple node selection conditions. This lacks the ability to perceive and adapt to dynamic load changes in the cluster. They cannot effectively handle dynamic scenarios such as load fluctuations and peak business traffic at different times, and they ignore the differences in characteristics between Pods (such as the different needs of latency-sensitive and compute-intensive Pods, and resource contention among multi-tenant Pods). This leads to some nodes being overcrowded while others are idle, resulting in consistently low resource utilization and a high risk of load imbalance.
[0020] Secondly, the data collection dimensions are scattered and lack systematic integration. Existing solutions often only collect basic node resource utilization data, failing to fully cover key information such as Pod running performance (such as startup time and restart frequency), application performance indicators (such as response latency and throughput), node historical running records, and load change trends. The one-sidedness of data support leads to a lack of scientific scheduling decisions and makes it difficult to accurately predict the running effect of Pods after deployment.
[0021] Furthermore, the candidate node selection mechanism is too simplistic, only satisfying the hard resource requests and basic affinity constraints of Pods, without fully considering key factors such as real-time load threshold constraints of nodes (e.g., excessively high resource utilization may cause performance fluctuations after the deployment of new Pods) and historical scheduling anomaly records (e.g., some nodes frequently cause Pods to crash due to hardware defects). As a result, the selected candidate nodes have stability risks, and Pods are prone to running abnormalities and service interruptions after deployment.
[0022] Finally, the existing scheduling scheme lacks a complete closed-loop optimization mechanism. After the scheduling decision is executed, there is no effective feedback channel for operational results and strategy iteration. The decision model and scheduling rules cannot be dynamically adjusted according to actual operational data, making it difficult to adapt to the resource optimization needs of a single cluster architecture and the cross-cluster collaborative scheduling needs of a multi-cluster architecture. Ultimately, this leads to low scheduling efficiency, significant fluctuations in service quality, serious resource waste, and an inability to cope with the dynamic adaptation challenges brought about by business expansion and changes in load characteristics.
[0023] To address the aforementioned issues, this application provides an example of an intelligent scheduling method for a container orchestration system. Please refer to [example provided]. Figure 1 As shown, Figure 1 A schematic flowchart illustrating an intelligent scheduling method for a container orchestration system provided in this application is shown. This is an example and not a limitation; the method can be applied to or run in electronic devices. The method includes: S101, obtain the cluster operation data of member clusters in a single-cluster architecture or multi-cluster architecture deployed by the container orchestration system.
[0024] The member cluster consists of multiple nodes, which are servers or computing units that host the operation of container groups (Pods). Cluster operation data includes node resource utilization, Pod performance, application performance metrics, and load change data.
[0025] S102, in response to receiving a scheduling request from a Pod to be scheduled, selects multiple candidate nodes from the member cluster that meet the Pod's running constraints.
[0026] S103 determines the target node from multiple candidate nodes based on cluster operation data, the feature information of the Pod to be scheduled, and the preset scheduling strategy of calling the AI inference engine.
[0027] S104 binds the Pod to be scheduled to the target node and monitors the running status data of the Pod after it is bound to the target node.
[0028] S105 optimizes the decision model and preset scheduling strategy of the AI inference engine based on the running status data, forming a scheduling closed loop.
[0029] The intelligent scheduling method for container orchestration systems provided in this application is applicable to mainstream container orchestration platforms such as Kubernetes. It can flexibly adapt to single-cluster architecture or multi-cluster architecture (multi-cluster architecture includes at least two mutually cooperating member clusters). Each member cluster consists of several nodes, which are physical servers, virtual computing units, or cloud server instances capable of supporting the operation of container groups (Pods) and are the underlying hardware resource carriers for Pod operation. Before the scheduling process starts, the container orchestration system continuously acquires cluster operation data through the data collection module deployed in the member cluster. It should be understood that this data collection module can collect and store various types of operation data in real time through the container orchestration system's native Metrics API, monitoring agents deployed at the node level (such as Prometheus Node Exporter), and application performance monitoring tools.
[0030] Specifically, node resource utilization includes core hardware resource usage data such as node CPU utilization, memory usage, disk I / O throughput, and network bandwidth usage; Pod performance includes runtime parameters such as Pod startup time, runtime, actual resource consumption, and number of restarts; application performance metrics include business-level performance data such as response latency, throughput, and error rate of applications hosted by Pods; and load change data includes information reflecting dynamic load changes, such as the overall cluster load and the load fluctuation trend of each node over different time periods, and the frequency of Pod creation / destruction. All of the above cluster operation data are stored in a time-series database to provide data support for subsequent scheduling decisions and model optimization. When the container orchestration system receives a scheduling request for a Pod to be scheduled (this request is submitted by the user as a Pod creation command or automatically triggered by the system to generate a Pod through expansion), the scheduling system first performs candidate node screening. During the screening process, the scheduling system verifies all nodes in the member cluster one by one according to the running constraints of the Pod to be scheduled. These running constraints include the resource requests declared by the Pod (such as minimum CPU and memory requirements), node selection constraints (such as Pod-node affinity / anti-affinity rules, node label matching requirements), taint tolerance configuration, and whether the currently available resources of the node meet the hard requirements for the Pod to run.
[0031] Through the above verification process, nodes that do not meet the constraints are excluded from all nodes, resulting in multiple candidate nodes that meet the basic requirements for Pod operation. If no candidate node meets the conditions, the Pod to be scheduled enters the scheduling waiting queue until a node meets the constraints and the filtering is triggered again.
[0032] After the candidate nodes are determined, the scheduling system first extracts the latest cluster operation data corresponding to the selected candidate nodes from the time-series database, and at the same time obtains the feature information of the Pod to be scheduled. This feature information includes the application identifier to which the Pod belongs, the tenant information, the resource requirement specifications, the business priority, and performance-sensitive characteristics (such as whether it is a latency-sensitive application). Subsequently, the scheduling system calls the pre-deployed AI inference engine, which integrates preset scheduling strategies (such as multi-tenant resource fair allocation strategy, online / offline load balancing strategy, resource utilization optimization strategy, etc.). The AI inference engine takes the cluster operation data and the feature information of the Pod to be scheduled as inputs and substitutes them into the trained decision model (such as a deep neural network model or reinforcement learning model) for inference calculation.
[0033] In some embodiments, the preset scheduling strategy includes tenant priority weights, resource quota rules, task scheduling constraints, and service quality assurance parameters; the load prediction data includes load change trends for future periods, resource supply and demand prediction results, and peak service traffic forecasts. The load change trend corresponds to the load changes of member clusters in a single-cluster architecture or each member cluster in a multi-cluster architecture. The load change of a member cluster is formed by summing the load changes of all its nodes.
[0034] Furthermore, the decision model, in conjunction with the objectives of the preset scheduling strategy, quantitatively evaluates the suitability of each candidate node for hosting the Pod to be scheduled. This evaluation dimension includes whether the remaining node resources can support the long-term stable operation of the Pod, whether the node load is balanced after the Pod is deployed, whether the performance requirements of the Pod are met, and whether it complies with the tenant resource allocation rules. Finally, it outputs the suitability score of each candidate node, and the scheduling system selects the candidate node with the highest score as the target node.
[0035] After identifying the target node, the scheduling system binds the Pod to be scheduled to the target node through the container orchestration system's API interface. For example, the target node's identifier is written into the Pod's `.spec.nodeName` field, completing the execution of the scheduling decision. After the binding operation is triggered, the kubelet component, the node agent on the target node, receives a notification from the container orchestration system and initiates the Pod creation and container instance startup process. Simultaneously, the data collection module continuously monitors the Pod's runtime status data after it is bound to the target node. This runtime status data includes the Pod's actual startup time, resource consumption fluctuations during operation, application performance, and whether resource contention or performance anomalies occur. This data is transmitted back to a time-series database in real time for correlation and storage. The correlation dimensions include scheduling decision records (such as the target node, scheduling time, and the cluster runtime data at that time) and the corresponding Pod runtime status data. Finally, the scheduling system initiates an optimization process based on the collected Pod runtime status data, forming a scheduling closed loop. This enables the AI inference engine's decision model and preset scheduling strategies to continuously adapt to changes in the cluster's runtime status, achieving self-evolution of scheduling decisions and ensuring the accuracy and optimization effectiveness of subsequent scheduling decisions. For example, the scheduling system periodically compares and analyzes the operational status data with the expected effects of scheduling decisions. If a deviation is found between the actual operational status and the expectations (such as resource shortages in Pods, performance falling short of expectations, or low node resource utilization leading to resource waste), this deviation data is input as feedback information into the model optimization module of the AI inference engine. Subsequently, the model optimization module uses the newly added operational status data to incrementally train or adjust the parameters of the decision model, enabling it to learn scheduling patterns that better fit the actual operational scenario. Simultaneously, based on the feedback results from the operational status data, the system dynamically optimizes the preset scheduling strategies in the AI inference engine, such as adjusting multi-tenant priority weights, updating resource allocation thresholds, and optimizing resource isolation parameters for mixed loads.
[0036] In some embodiments, such as Figure 2 As shown, based on cluster runtime data, the characteristic information of the Pods to be scheduled, and the preset scheduling strategy of calling the AI inference engine, the target node is determined from multiple candidate nodes, including: S201, based on cluster operation data, feature information of Pods to be scheduled, and the preset scheduling strategy of calling the AI inference engine, scores the suitability of each candidate node and obtains the score results of each candidate node.
[0037] S202, Based on the scoring results of each candidate node, determine the target node among multiple candidate nodes.
[0038] First, the scheduling system extracts the latest cluster operation data corresponding to the selected candidate nodes from the time-series database, and simultaneously obtains the feature information of the Pod to be scheduled. The feature information includes the application identifier of the Pod, the tenant information, the resource requirement specifications, the business priority, and performance-sensitive characteristics (such as whether it is a latency-sensitive application). Then, the scheduling system calls the pre-deployed AI inference engine, which integrates preset scheduling strategies (such as multi-tenant resource fair allocation strategy, online / offline load balancing strategy, resource utilization optimization strategy, etc.). The AI inference engine takes the cluster operation data and the feature information of the Pod to be scheduled as input, combines the goal orientation of the preset scheduling strategy, and substitutes it into the trained decision model (such as a deep neural network model or reinforcement learning model) to perform inference calculations, and quantifies the suitability of each candidate node to carry the Pod to be scheduled.
[0039] During the scoring process, the decision-making model assigns corresponding weights to different evaluation dimensions based on the priority of the preset scheduling strategy. The evaluation dimensions specifically include whether the remaining node resources can support the long-term stable operation of the Pod, whether the node load is balanced after the Pod is deployed, whether the performance requirements of the Pod are met (such as the requirements of latency-sensitive applications for node network latency), whether it complies with the tenant resource allocation rules, and whether it adapts to the resource isolation requirements of mixed online / offline loads. Finally, it outputs a specific score result for each candidate node, which is a quantitative value of 0-100 points. The higher the score, the higher the compatibility between the candidate node and the Pod to be scheduled. On the other hand, the scheduling system receives the score results of each candidate node returned by the AI inference engine, sorts the candidate nodes in descending order of score, and selects the first-ranked candidate node, i.e., the candidate node with the highest score, as the target node.
[0040] If multiple candidate nodes have the same score and are all the highest, then the final target node is determined by further combining auxiliary rules such as the node's historical scheduling success rate and hardware configuration priority (e.g., high-performance computing nodes are preferentially allocated to compute-intensive Pods).
[0041] In some embodiments, such as Figure 3 As shown, based on cluster operation data, the feature information of the Pods to be scheduled, and the preset scheduling strategy of calling the AI inference engine, the suitability score of each candidate node is calculated, and the score results of each candidate node are obtained, including: S301 obtains the real-time running status of each candidate node through the status awareness interface.
[0042] The real-time operating status includes CPU load, memory usage, I / O performance, and hardware topology information.
[0043] S302, extract the feature information of the Pod to be scheduled, including its tenant, resource requirements, performance sensitivity and business priority.
[0044] S303 inputs real-time running status and feature information into the AI inference engine, so that the decision model of the AI inference engine can predict the running effect of the Pod to be scheduled on each candidate node based on the preset scheduling strategy.
[0045] S304, based on the running performance of the Pod to be scheduled on each candidate node, outputs the adaptability score of each candidate node as the scoring result.
[0046] The fit score is used to quantify the expected running efficiency and resource fit of a Pod on the corresponding candidate node.
[0047] In some embodiments, the scheduling system can obtain the real-time running status of each candidate node through a preset state-aware interface. The state-aware interface collects dynamic running data of the candidate nodes in real time through communication with the node monitoring agent and hardware management module. The real-time running status specifically includes hardware-level data that directly affects the running efficiency of Pods, such as CPU load (e.g., current CPU utilization, run queue length), memory usage (e.g., percentage of used memory, cache memory size), I / O performance (e.g., disk read / write throughput, storage I / O latency, network I / O response speed), and hardware topology information (e.g., number of CPU cores and affinity configuration, memory slot distribution, storage device type and interface specifications, network adapter bandwidth).
[0048] Afterwards, the scheduling system extracts the feature information of unscheduled Pods to obtain the core factors affecting scheduling decisions. These include the tenant (identifying the user entity corresponding to the Pod, used to adapt to multi-tenant resource allocation rules), resource requirements (clarifying the specific requirements of the Pod for CPU, memory, storage, and network), performance sensitivity (such as whether it is a latency-sensitive, throughput-sensitive, or compute-intensive application), and business priority (such as core business Pods having higher priority than non-core business Pods).
[0049] Then, the scheduling system formats the collected real-time running status of candidate nodes and the extracted feature information of Pods to be scheduled, and sends it as input parameters to the AI inference engine. The decision model of the AI inference engine (such as a deep neural network model or a reinforcement learning model) combines the goal orientation of the preset scheduling strategy to perform feature fusion and inference calculation on the input data, focusing on predicting the running effect of the Pods to be scheduled on each candidate node. The running effect includes key indicators such as expected application response latency, throughput, resource utilization, running stability (such as whether resource contention is likely to occur), and probability of meeting the business SLA.
[0050] Finally, based on the predicted performance, the AI inference engine assigns a quantitative score to each candidate node. The score reflects the expected performance and resource suitability of the Pod on the corresponding candidate node. The score is a quantitative value from 0 to 100. The higher the expected performance (e.g., lower response latency, higher throughput) and the better the resource suitability (e.g., the higher the match between the node's remaining resources and the Pod's needs, the better the hardware topology matches the Pod's computing needs), the higher the corresponding score. The final score is the score for each candidate node. During the scoring process, the decision model assigns corresponding weights to different evaluation dimensions based on the priority of the preset scheduling strategy. For example, under the multi-tenant fairness strategy, the evaluation weight of tenant resource quota utilization rate will be increased; under the resource utilization optimization strategy, the weight of the utilization efficiency of the remaining node resources will be increased. The evaluation dimensions can specifically include whether the remaining node resources can support the long-term stable operation of the Pod, whether the node load is balanced after the Pod is deployed, whether the performance requirements of the Pod are met, whether it complies with the tenant resource allocation rules, and whether it adapts to the resource isolation requirements of mixed online / offline loads, etc., to ensure that the scoring results are consistent with the scheduling strategy objectives.
[0051] In addition, the scheduling system receives the scoring results of each candidate node returned by the AI inference engine, sorts the candidate nodes in descending order of score, and selects the first-ranked candidate node with the highest score as the target node. If multiple candidate nodes have the same score and are all the highest scores, the final target node is determined by further combining auxiliary rules such as the node's historical scheduling success rate and hardware configuration priority (e.g., high-performance computing nodes are preferentially allocated to compute-intensive Pods). In some embodiments, such as Figure 4 As shown, based on cluster runtime data, the characteristic information of the Pods to be scheduled, and the preset scheduling strategy of calling the AI inference engine, the target node is determined from multiple candidate nodes, including: S401 standardizes the cluster operation data and the feature information of the Pods to be scheduled, and extracts key feature vectors.
[0052] S402 calculates the resource adaptation score, performance prediction score, and stability score of each candidate node through the decision model of the AI inference engine.
[0053] S403: Based on the weights of the preset scheduling strategy, the scores of each candidate node are weighted and summed to obtain the final fit score of each candidate node.
[0054] S404: Sort multiple candidate nodes from high to low according to their final fitness scores, and select the node with the highest score as the target node.
[0055] In some embodiments, the scheduling system can perform standardization processing on the extracted Pod feature information to be scheduled and the cluster operation data obtained from the time-series database to eliminate the differences in the units of measurement of different data types (e.g., normalizing proportional data such as CPU utilization and memory usage to the [0,1] range, and standardizing numerical data such as network bandwidth and disk throughput according to a preset threshold). Redundant information is then eliminated through feature selection algorithms (e.g., feature filtering based on mutual information, variance thresholding) to extract key feature vectors that reflect resource adaptability, operational performance, and stability. It should be understood that these key feature vectors cover core dimensions such as the remaining proportion of node resources, the matching degree between Pod requirements and node hardware, the node load fluctuation coefficient, and parameters corresponding to application performance-sensitive features.
[0056] In some embodiments, the scheduling system inputs the standardized key feature vectors into the AI inference engine. The decision model of the AI inference engine (such as a deep neural network model or a reinforcement learning model) combines the preset scheduling strategy to calculate three core scores for each candidate node: resource adaptation score, which quantifies the degree of matching between the node's remaining resources and the Pod's resource requirements and resource utilization efficiency. The calculation is based on the degree of fit between the remaining amount of node CPU, memory, storage and other resources and the Pod's requirements, and the rationality of resource allocation (such as avoiding excessive resource reservation or insufficient allocation); performance prediction score, which predicts whether the application response latency, throughput, processing speed and other performance indicators meet the standards after the Pod is deployed, based on the Pod's performance sensitivity characteristics and the node's real-time running status; and stability score, which assesses the stability of the Pod's long-term operation on the node. The calculation is based on the node's historical load fluctuation range, the probability of resource contention, hardware failure records, and the risk of Pod restart.
[0057] In some embodiments, the AI inference engine assigns corresponding weights to resource adaptation score, performance prediction score, and stability score based on the priority configuration of preset scheduling strategies. For example, under the resource utilization optimization strategy, the weight of resource adaptation score is set to 0.5, the weight of performance prediction score is set to 0.3, and the weight of stability score is set to 0.2; under the latency-sensitive service scheduling strategy, the weight of performance prediction score is set to 0.6, the weight of stability score is set to 0.3, and the weight of resource adaptation score is set to 0.1. Then, the three scores of each candidate node are weighted and summed, that is, the final adaptation score = resource adaptation score × resource adaptation weight + performance prediction score × performance weight + stability score × stability weight, to obtain the final adaptation score of each candidate node (the score range is 0-100). If there are multiple preset scheduling strategies superimposed, the AI inference engine will first normalize the weights corresponding to each strategy according to the strategy priority, and then perform weighted summation calculation to ensure that the scoring results are consistent with the comprehensive scheduling objectives. Finally, the scheduling system receives the final fit scores of each candidate node returned by the AI inference engine, sorts the candidate nodes in descending order of score, and selects the node with the highest score as the target node. If multiple candidate nodes have the same score and are all the highest, the final target node is determined by further considering auxiliary rules such as the node's historical scheduling success rate, hardware configuration priority (e.g., high-performance computing nodes are prioritized for computationally intensive Pods), and the remaining tenant resource quota, ensuring the uniqueness and rationality of the scheduling decision.
[0058] In some embodiments, such as Figure 5 As shown, the decision-making model and preset scheduling strategy of the AI inference engine are optimized based on the running status data to form a scheduling closed loop, including: S501 compares the operating status data with the prediction results during scheduling decisions to obtain the deviation value.
[0059] The operational status data includes startup time, operational stability, resource consumption deviation, and service quality indicators.
[0060] S502, if the deviation value exceeds the preset threshold, adjust the model parameters and candidate node score weights of the AI inference engine's decision model.
[0061] S503 accumulates operational status data to form a historical dataset, which includes historical operational data of member clusters in a single-cluster architecture or member clusters in a multi-cluster architecture.
[0062] The historical operational data of each member cluster is formed by summing the historical operational data of all the nodes it contains.
[0063] S504 uses historical datasets to periodically train and update the decision model of the AI inference engine.
[0064] In some embodiments, the scheduling system extracts the Pod's running status data and corresponding scheduling decision prediction results from the time-series database. The running status data specifically includes the Pod's actual startup time, running stability (such as the number of restarts within the runtime, whether there are crashes or performance fluctuations), resource consumption deviation (such as the difference between actual CPU / memory consumption and the consumption predicted at the time of scheduling decision), and service quality indicators (such as actual response latency, throughput, error rate, etc.). The scheduling decision prediction results are the resource adaptation expectations, performance achievement expectations, and stability expectations output by the AI inference engine during scheduling.
[0065] First, the scheduling system compares the running status data with the corresponding prediction results item by item, and calculates the deviation value of each indicator through the preset deviation calculation algorithm (such as absolute error, relative error, mean square error, etc.). For example, the startup time deviation value = |actual startup time - predicted startup time|, and the resource consumption deviation value = |actual resource consumption - predicted resource consumption| / predicted resource consumption × 100%.
[0066] Afterwards, the scheduling system compares the calculated deviation values with preset thresholds (which can be configured by the user according to business needs, such as setting the startup time deviation threshold to 5 seconds and the resource consumption deviation threshold to 20%). If any one or more deviation values exceed the corresponding preset threshold, it indicates that the prediction result of the current scheduling decision deviates significantly from the actual operating effect. The scheduling system immediately triggers an adjustment mechanism. On the one hand, it adjusts the model parameters of the decision model in the AI inference engine (such as adjusting the weight coefficients of the neural network and the reward function parameters of the reinforcement learning model) to optimize the prediction accuracy of the model. On the other hand, it adjusts the weight configuration of the candidate node scores (if the performance prediction deviation is too large, the weight ratio of the performance prediction score is appropriately adjusted) to make the subsequent scoring results more consistent with the actual operating scenario.
[0067] Then, the scheduling system continuously accumulates the running status data of all Pods, and combines it with historical scheduling decision records to form a unified historical dataset. For a single cluster architecture, this historical dataset contains all the historical running data of that cluster; for a multi-cluster architecture, this historical dataset contains the historical running data of each member cluster, and the historical running data of each member cluster is formed by summarizing the historical running data of all its nodes (such as node resource usage history, Pod scheduling history, running status history, etc.), ensuring the comprehensiveness and representativeness of the historical dataset.
[0068] The scheduling system sets up regular training cycles (such as during the early morning hours when the cluster load is low). At the start of each training cycle, it calls the model training module of the AI inference engine, divides the historical dataset into a training set and a validation set according to a preset ratio, and uses the training set to perform incremental or full training on the decision model. During the training process, the model performance (such as prediction accuracy, bias control effect, etc.) is monitored in real time through the validation set. When the model performance reaches the preset standard, the training stops, and the new model is used to replace the old model in the AI inference engine, thus completing the update of the decision model. Through the aforementioned optimization process, the AI inference engine's decision model can continuously learn patterns from real-world operational scenarios, and the weight configuration of the preset scheduling strategy can dynamically adapt to changes in the cluster's operational status, achieving self-evolution of scheduling decisions. For example, if the performance prediction deviation of a certain type of latency-sensitive Pod repeatedly exceeds a threshold, the scheduling system optimizes the performance prediction logic by adjusting model parameters and increasing the weight of the performance prediction score. This ensures more accurate matching of high-performance nodes during subsequent scheduling, thereby continuously improving the accuracy and optimization effect of scheduling decisions, forming a complete closed loop of "scheduling decision-operation monitoring-deviation analysis-parameter adjustment-model update-optimized scheduling".
[0069] In some embodiments, before filtering multiple candidate nodes from the member cluster that meet the Pod's runtime constraints in response to receiving a scheduling request for the Pod to be scheduled, the method further includes: The custom controller receives custom resource objects submitted by external systems or users through its declarative interface. These custom resource objects include scheduling policy configuration objects and load prediction objects. Receive preset scheduling strategies and load prediction data through custom resource objects; By monitoring changes to custom resource objects through a custom controller, the updated preset scheduling strategy is synchronized to the AI inference engine and scheduling process, and load prediction data is used as input features for the AI inference engine. Adjust the Pod running constraints and adaptability score weights based on load forecast data to ensure that the preset scheduling strategy dynamically links with load changes.
[0070] In some embodiments, before the container orchestration system is ready to respond to the scheduling request of the Pod to be scheduled, it also needs to perform a preprocessing process of scheduling policy and load prediction, which specifically includes the following steps: First, the scheduling system receives custom resource objects submitted by external systems (such as automated operation and maintenance platforms, traffic prediction services) or users through the declarative interface provided by the custom controller. The custom controller is an extension component deployed on the control plane of the container orchestration system. The declarative interface follows the native API specification of the container orchestration system and supports users or external systems to submit configurations in a declarative manner. The custom resource object is a resource instance defined based on the container orchestration system's CRD (Custom Resource Definition) mechanism, specifically including a scheduling policy configuration object and a load prediction object. The scheduling policy configuration object is used to carry high-level scheduling rules, and the load prediction object is used to transmit load trend data for future periods.
[0071] The scheduling system derives preset scheduling policies through parsing the scheduling policy configuration object. These preset policies include multi-tenant resource fair allocation strategies, online / offline mixed load deployment strategies, latency-sensitive service priority strategies, and resource utilization optimization strategies, which can be flexibly defined by users according to business needs. Load prediction data is obtained through parsing the load prediction object. This data includes information such as the overall cluster load trend within a preset future time period (e.g., 1 hour, 3 hours), traffic growth predictions for specific applications or tenants, and resource supply-demand gap predictions. This data is generated by external intelligent analysis systems or by users based on business experience. Thirdly, the custom controller continuously monitors the change status of custom resource objects (including creation, update, and deletion operations). When an update to the scheduling policy configuration object is detected, the custom controller immediately synchronizes the updated preset scheduling policy to the policy configuration module of the AI inference engine and the policy execution stage of the entire scheduling process, ensuring that subsequent scheduling decisions adopt the latest policy.
[0072] When an update to the load forecast object is detected, the custom controller pushes the updated load forecast data to the data collection module. This data is then integrated with the cluster's operational data and used as input features for the AI inference engine, providing the decision model with a reference for future load trends. Fourthly, the scheduling system dynamically adjusts scheduling-related configurations based on the load forecast data: on the one hand, it adjusts Pod running constraints. For example, if load forecast data shows that cluster CPU resources will be strained at a certain time in the future, the CPU resource request constraints for Pods can be temporarily tightened to prevent low-priority Pods from excessively consuming resources. If it predicts that a node will experience resource shortages due to peak load, the filtering restrictions for that node can be temporarily added to the constraints to reduce the probability of Pods being scheduled to that node. On the other hand, it adjusts the adaptability score weights. For example, if load forecast data shows that the offline task load will increase significantly in the future, the weight of the "offline task resource compatibility" dimension in the resource adaptability score can be increased, or the multi-tenant score weights can be adjusted to ensure tenant resource needs. This achieves dynamic linkage between the preset scheduling strategy and load changes, ensuring that scheduling decisions adapt to future load trends in advance.
[0073] In some embodiments, the method further includes a multi-cluster scheduling step: The federated scheduling controller maintains a global resource view across clusters. The global resource view covers the resource sufficiency, network latency, and operating cost information of all member clusters in the multi-cluster architecture. The resource sufficiency of each member cluster is calculated by summing the available resources of all nodes contained in the member cluster. In response to receiving a cross-cluster scheduling request, candidate clusters that meet the cross-cluster scheduling requirements are filtered based on the global resource view. The candidate clusters belong to the member clusters in the multi-cluster architecture, and each candidate cluster contains multiple nodes. The global AI decision-making model is invoked to evaluate the suitability of each candidate cluster based on preset evaluation factors. These evaluation factors include at least one of the following: resource matching degree, task execution efficiency, operating cost, and disaster recovery priority. Assign the Pod to be scheduled to the target cluster with the highest suitability, and coordinate the target cluster to perform candidate node filtering and Pod binding operations. The target cluster is one of the candidate clusters, and the candidate node filtering operation is to select nodes that meet the Pod running constraints from multiple nodes contained in the target cluster. When any member cluster becomes overloaded or fails, the Pods bound to that member cluster will be migrated from the nodes contained in that member cluster to the nodes contained in other normal member clusters.
[0074] In multi-cluster architecture scenarios, this application embodiment also provides a dedicated multi-cluster scheduling step, which includes, but is not limited to, the following method steps: The first step is for the scheduling system to maintain a global resource view across clusters through the deployed federated scheduling controller. This federated scheduling controller is the core coordination component for cross-cluster scheduling. It establishes communication connections with the API servers of each member cluster, collects and summarizes key information of all member clusters in real time, and forms a global resource view.
[0075] In some embodiments, the information in the global resource view specifically includes the resource sufficiency, network latency, and operating cost information of each member cluster. Resource sufficiency refers to the overall resource availability of a member cluster, calculated by summing the available resources of all nodes within that cluster (such as the remaining CPU cores, free memory capacity, available storage space, and remaining network bandwidth). For example, the CPU resource sufficiency of a member cluster = (sum of remaining CPU cores of all nodes / total CPU cores of all nodes) × 100%, and the memory resource sufficiency is calculated similarly. Network latency includes cross-cluster network communication latency between member clusters (such as average and peak latency of data transmission between clusters) and network latency between member clusters and service terminals. Operating cost information includes quantifiable cost indicators such as hardware deployment costs, energy consumption costs, and maintenance costs for each member cluster. The federated scheduling controller periodically updates the global resource view to ensure that the view information is consistent with the actual status of each member cluster. The update cycle can be configured according to business needs (e.g., every 10 seconds or every minute). The second step involves the container orchestration system receiving a cross-cluster scheduling request (generated by user-specified cross-cluster Pod deployment, system-triggered cross-cluster scaling, or disaster recovery requirements). The federated scheduling controller then performs a candidate cluster filtering operation based on the global resource view. During this filtering process, the federated scheduling controller verifies all member clusters according to the cross-cluster scheduling requirements. These requirements include the total resource requirements of the Pods to be scheduled, the upper limit of network latency for the business (e.g., cross-cluster communication latency not exceeding 50ms), disaster recovery requirements (e.g., whether deployment in different regions is required), and the operating cost budget threshold. Member clusters that do not meet the requirements are excluded. For example, if a member cluster's resource availability cannot meet the total resource requirements of the Pods, or the cross-cluster network latency exceeds business requirements, or the operating cost exceeds the budget, that cluster is excluded, ultimately resulting in multiple candidate clusters that meet the cross-cluster scheduling requirements. The third step involves the federated scheduling controller, after the candidate clusters are determined, invoking a pre-deployed global AI decision-making model to evaluate the suitability of each candidate cluster. This global AI decision-making model collaborates with the AI inference engine within a single cluster. The evaluation process is based on preset evaluation factors, including at least one of the following: resource matching degree, task execution efficiency, operating cost, and disaster recovery priority. Resource matching degree assesses the degree to which the resource sufficiency of the candidate cluster matches the resource requirements of the Pod to be scheduled, including resource type matching (such as whether the GPU resources required by the Pod are provided), resource quantity matching, and future resource availability (combined with load prediction data). Task execution efficiency assesses the data processing speed and response latency compliance probability after the candidate cluster hosts the Pod, calculated comprehensively based on the cluster's hardware performance, network bandwidth, load status, and cross-cluster communication efficiency. Operating cost assesses the long-term operating cost of the Pod after deployment in the candidate cluster, combined with indicators such as the cluster's hardware cost, energy consumption, and maintenance costs. Disaster recovery priority, for core business Pods, assesses the independence of the candidate cluster's location, data center environment, and other clusters; the higher the independence, the higher the disaster recovery priority. The global AI decision-making model quantifies each evaluation factor into a specific score, and then performs a weighted sum based on preset weights (such as disaster recovery priority weight 0.4, task execution efficiency weight 0.3, resource matching weight 0.2, and operating cost weight 0.1 in core business scenarios; and operating cost weight 0.4, resource matching weight 0.3, task execution efficiency weight 0.2, and disaster recovery priority weight 0.1 in cost-sensitive scenarios) to obtain the total suitability score for each candidate cluster. The fourth step involves the federated scheduling controller selecting the candidate cluster with the highest overall fit score as the target cluster and allocating the Pods to be scheduled to this target cluster. This allocation is achieved through cross-cluster scheduling commands. The federated scheduling controller sends scheduling commands to the target cluster's scheduling system, containing data such as the Pod's characteristics, runtime constraints, and preset scheduling policies. Upon receiving the scheduling commands, the target cluster's scheduling system initiates a candidate node selection process within the single cluster (i.e., selecting nodes that meet the requirements based on the Pod's runtime constraints as described earlier) and subsequent Pod binding operations, completing the final deployment of the Pod within the target cluster. The federated scheduling controller monitors this process in real time to ensure successful scheduling execution. The fifth step involves the federated scheduling controller continuously monitoring the operational status of all member clusters. When any member cluster is detected to be overloaded or experiencing a failure (overload refers to cluster resource utilization exceeding a preset threshold, such as CPU utilization exceeding 90% for 5 consecutive minutes; failure includes cluster API server unavailability, widespread node downtime, network interruption, etc.), the Pod migration process is immediately initiated. First, the federated scheduling controller identifies all bound Pods within the faulty or overloaded cluster, distinguishing between core business Pods and non-core business Pods. Then, based on the global resource view, it selects other normally operating member clusters (i.e., clusters with sufficient resources and stable status) as migration target clusters. Next, the federated scheduling controller coordinates with the target cluster to reserve resources and sends Pod eviction instructions to the faulty / overloaded cluster, while simultaneously sending Pod creation instructions to the target cluster, migrating the Pods from the nodes of the original cluster to the nodes of the target cluster. During the migration process, service continuity for core business Pods is ensured (e.g., using a rolling migration method), while non-core business Pods can be migrated sequentially according to priority. After migration is complete, the global resource view and related scheduling records are updated.
[0076] In some embodiments, coordinating the target cluster to perform candidate node filtering and Pod binding operations includes: The federated scheduling controller synchronizes the cross-cluster scheduling instructions of Pods to the scheduling plugins of the target cluster. The scheduling plugin of the target cluster completes the node allocation of Pods in the target cluster according to the candidate node screening process, AI scoring process and binding process; By synchronizing the running status of Pods in the target cluster in real time through the federated scheduling controller, a full-process monitoring of cross-cluster scheduling is formed.
[0077] In some embodiments, after the federated scheduling controller selects the candidate cluster with the highest overall adaptation score as the target cluster, the federated scheduling controller accurately synchronizes the cross-cluster scheduling instructions of the Pod to the pre-deployed scheduling plugin of the target cluster through a secure communication link established with the API server of the target cluster. This scheduling plugin is an extension component in the target cluster that adapts to cross-cluster scheduling scenarios. It works in deep collaboration with the cluster's native scheduling system and is specifically responsible for receiving and parsing cross-cluster scheduling related requests.
[0078] The cross-cluster scheduling instructions include complete characteristic information of the Pod to be scheduled (including tenant identifier, resource requirement specifications, performance-sensitive characteristics, business priority, etc.), pre-processed Pod running constraints, preset scheduling strategies (including tenant priority weights, resource quota rules, task scheduling constraints and service quality assurance parameters), load prediction data, and adaptation evaluation reference information of the global AI decision model. This ensures that the scheduling plugin of the target cluster can obtain all the decision-making basis required to complete node allocation, and the instruction transmission process follows the security specifications of container orchestration systems, encrypting sensitive configuration information to ensure data security.
[0079] After the scheduling plugin in the target cluster successfully receives the cross-cluster scheduling instruction, it immediately initiates the entire node allocation process within the cluster, executing step by step according to the candidate node screening process, AI scoring process, and binding process defined above: In the candidate node screening stage, the scheduling plugin, based on the Pod running constraints in the instruction, sequentially completes hard condition filtering, real-time load status exclusion, and historical scheduling effect reference elimination to form a set of candidate nodes that meet the basic operating requirements; In the AI scoring process, the scheduling plugin calls the AI inference engine deployed locally in the target cluster, integrates the real-time running status of candidate nodes, Pod feature information, cluster running data, and load prediction data, and after standardization processing and feature extraction, calculates the resource adaptation score, performance prediction score, and stability score of each candidate node, and performs a weighted summation based on the weight configuration of the preset scheduling strategy to obtain the final adaptation score of each candidate node; In the binding process, the scheduling plugin selects target nodes in descending order of adaptation score, binds the Pod to the target node through the native API interface of the container orchestration system, triggers the kubelet component on the target node to perform Pod creation and container instance start-up operations, and finally completes the node allocation of the Pod in the target cluster.
[0080] Throughout the entire node allocation process described above, the federated scheduling controller continuously synchronizes Pod runtime status data through real-time bidirectional communication with the target cluster scheduling plugin. This data includes the progress and results of candidate node selection, intermediate and final scores of AI scoring, Pod binding status with target nodes, Pod startup progress, container instance runtime status, and any anomalies such as resource contention or performance abnormalities. This status data is fed back to the scheduling management module of the federated scheduling controller in real time and simultaneously stored in a time-series database and archived in association with cross-cluster scheduling records. This forms a complete monitoring link for cross-cluster scheduling, from instruction issuance, node selection, AI scoring, node binding to Pod runtime, ensuring that the scheduling execution process is traceable and auditable. If anomalies are detected during monitoring (such as candidate node selection failure, binding operation timeout, Pod startup anomalies, etc.), the target cluster scheduling plugin immediately reports anomaly alarms and detailed fault information to the federated scheduling controller. The federated scheduling controller can trigger corresponding fault tolerance mechanisms based on the anomaly type, including re-executing the node selection process within the target cluster, adjusting the target cluster's scheduling parameters, or re-evaluating candidate clusters and switching to a suboptimally suitable target cluster, ensuring the reliability of cross-cluster scheduling and the success rate of Pod deployment.
[0081] In some embodiments, the decision model of the AI inference engine includes a deep neural network model and a reinforcement learning model, and the decision model is trained in the following manner: Based on historical cluster operation data and scheduling effect data, a model training dataset is constructed. The historical cluster operation data includes the historical data of member clusters in a single cluster architecture or the historical data of each member cluster in a multi-cluster architecture. The historical data of each member cluster is formed by summarizing the historical data of all nodes it contains. The training dataset includes node status, Pod feature information and scheduling result labels under different load scenarios. A supervised learning approach is used to train a deep neural network model based on a model training dataset, learning the mapping relationship between node states and Pod performance. A reinforcement learning approach is adopted to train a reinforcement learning model based on the model training dataset, with the reward objectives of maximizing resource utilization and minimizing task latency, thereby optimizing the scheduling decision strategy. Regularly update the model parameters of deep neural network models and reinforcement learning models using newly added historical cluster operation data and scheduling effect data to enhance the adaptability of decision models to load changes in member clusters in a single cluster architecture or in member clusters in a multi-cluster architecture.
[0082] In some embodiments, the decision-making model of the AI inference engine includes a deep neural network model and a reinforcement learning model. The first step in model training is to build a model training dataset, which is generated by integrating historical cluster operation data and scheduling effect data. The historical cluster operation data covers the historical operation records of member clusters in a single cluster architecture or all member clusters in a multi-cluster architecture. The historical data of each member cluster is calculated by aggregating the historical data of all its nodes (including node resource usage history, load fluctuation data, hardware status change records, etc.) to ensure data integrity and hierarchical consistency. The scheduling effect data records the actual operating performance of the Pods corresponding to each scheduling decision, including quantitative indicators such as resource consumption, response latency, operational stability, and service quality compliance status.
[0083] In this embodiment, the training dataset is constructed according to load scenarios, comprehensively covering different operating scenarios such as low load, high load, fluctuating load, and peak load. Each scenario includes corresponding node status data (such as real-time resource utilization of nodes, hardware topology information, load fluctuation coefficient, etc.), Pod feature information (such as tenant, resource requirement specifications, performance sensitivity, business priority, etc.), and scheduling result labels. The label is a quantitative value of the actual running effect after the Pod is scheduled to the corresponding node, ensuring that the training dataset has complete scenario coverage and accurate labels, providing reliable data support for model training.
[0084] For deep neural network models, supervised learning is employed for training. For example, node state data and Pod feature information from the model training dataset are used as input features of the model. After standardization and feature dimension optimization, these are input into the deep neural network model. The scheduling result label is used as the model output target. The backpropagation algorithm is used to minimize the error between the model's predicted value and the actual label value. The mean squared error loss function is used to measure the prediction bias. The model's weight coefficients and bias parameters are continuously iterated and adjusted so that the model gradually learns the nonlinear mapping relationship between node state, Pod features, and Pod performance. Ultimately, this enables the model to accurately predict the core performance indicators of Pods on target nodes, such as resource consumption, response latency, and stability, based on the input features.
[0085] For reinforcement learning models, reinforcement learning is used for training. For example, the scheduling decision process is modeled as a Markov decision process, where the state space is defined as a set of node state data, Pod feature information, and load scenario information of the current cluster; the action space is defined as a configuration scheme for candidate node selection and suitability score weights; and the reward objective is set as a comprehensive optimization objective of maximizing resource utilization and minimizing task latency. The reward function is implemented through weighted summation (i.e., reward value = α × resource utilization - β × task latency, where α and β are weight coefficients dynamically adjusted according to a preset scheduling strategy to balance the priorities of the two optimization objectives). Based on the constructed model training dataset, the reinforcement learning model explores scheduling actions in the state space of different load scenarios. By calculating the immediate reward (resource utilization and task latency performance after a single scheduling) and the long-term cumulative reward (the comprehensive effect of multiple schedulings over a period of time) for each action, the model's action selection strategy is continuously optimized, enabling the model to output the optimal scheduling decision while meeting Pod running constraints, tenant priority rules, and service quality requirements.
[0086] After the initial training of the model is completed, a continuous update mechanism is initiated to enhance the model's dynamic adaptability: the scheduling system periodically (e.g., during the early morning hours when cluster load is low) collects newly added historical cluster operation data and scheduling effect data to expand the original model training dataset. The expanded dataset maintains consistency in scene classification attributes and data format. Incremental training is used to update the model parameters of the deep neural network model and the reinforcement learning model. Specifically, the deep neural network model fine-tunes the network weights and bias parameters by adding new data to avoid overfitting and improve its adaptability to new load features. The reinforcement learning model expands its state space based on the new scene data and dynamically adjusts the weight coefficients α and β of the reward function to optimize the balance between exploration and utilization of action selection strategies. The update process is monitored in real time using a validation set. Model performance evaluation metrics for deep neural network models include prediction accuracy, operational performance metrics such as prediction deviation, and performance evaluation metrics for reinforcement learning models include the improvement in resource utilization of scheduling decisions, the effect of task latency reduction, and the service quality compliance rate. When the model performance reaches the preset standards (e.g., prediction accuracy ≥ 95%, resource utilization improvement ≥ 10%, task latency reduction ≥ 15%), training is stopped, and the updated model replaces the old model currently in use in the AI inference engine. This ensures that the decision model can continuously adapt to the load changes of member clusters in a single-cluster architecture or in each member cluster in a multi-cluster architecture (e.g., load characteristic changes brought about by new services, resource distribution adjustments after cluster expansion, load coordination changes in cross-cluster scheduling scenarios, etc.), and always maintain the optimization effect and accuracy of scheduling decisions.
[0087] In some embodiments, the constraints for Pod operation include: Based on the Pod's resource request volume, node affinity constraints, taint tolerance rules, and hardware requirements, nodes that do not meet the hard conditions are filtered out.
[0088] Based on real-time node load status, nodes whose resource utilization exceeds a preset threshold are excluded.
[0089] Based on historical scheduling performance data, nodes that previously caused Pods to malfunction were removed, resulting in a number of candidate nodes.
[0090] In some embodiments, during the candidate node screening phase, the scheduling system first performs a first-level filtering based on the Pod's hard configuration requirements. This involves extracting the resource request amount declared by the Pod to be scheduled (including the minimum quota of CPU cores, memory capacity, storage space, and network bandwidth), node affinity constraints (such as the Pod specifying that it needs to be scheduled to a node with a specific label, or a node in its availability zone or rack location), taint tolerance rules (i.e., whether the Pod can tolerate different levels of taints such as "No Schedule" or "No Execute" set by the node), and hardware requirements (such as whether dedicated computing hardware such as GPUs or FPGAs is required, or specific requirements for CPU architecture, memory frequency, and storage media type). The scheduling system verifies all nodes in the cluster one by one according to the above configuration requirements. Any node that does not meet any of the hard conditions is directly filtered out. For example, if the node's available resources are lower than the Pod's request amount, the node label does not match the affinity constraints, the Pod cannot tolerate the node's taints, or the node does not have dedicated hardware support, the node is included in the exclusion list. This initially filters out a set of nodes that meet the basic operating requirements.
[0091] Based on the first-level screening results, the scheduling system initiates the second-level real-time load status verification. It obtains the core resource utilization data of each node in real time through the status awareness interface, including CPU utilization, memory usage, disk I / O utilization, and network bandwidth usage. The preset thresholds for the above resource utilization can be flexibly configured according to the cluster operation and maintenance strategy and business needs (e.g., the CPU utilization threshold is set to 85%, and the memory usage threshold is set to 80%). The scheduling system compares the real-time resource utilization of the nodes with the corresponding preset thresholds. If the utilization of any core resource exceeds the threshold, it indicates that the current load of the node is high and the remaining resources are insufficient to support the stable operation of the new Pod. In order to avoid the node overload caused by the deployment of the new Pod or to affect the performance and stability of the Pods already running on the node, the scheduling system excludes the node from the candidate set and only retains the nodes whose real-time load is in the safe operating range.
[0092] Finally, the scheduling system performs a third-level historical scheduling effect reference verification. Historical scheduling effect data is uniformly stored in a time-series database, detailing the operational status of each node after scheduling Pods, including whether Pods experienced restarts, crashes, resource contention, performance fluctuations, or other anomalies, as well as the frequency, duration, and severity of these anomalies. The scheduling system queries the historical data of the remaining nodes after the first two layers of filtering, setting an anomaly rate threshold (e.g., the anomaly rate of Pods scheduled to this node exceeding 10% in the past 30 days) as a verification standard. If a node has repeatedly caused Pod operational anomalies in the past, or has experienced serious anomalies that led to the interruption of core business Pod services, it indicates potential risks in hardware stability and resource allocation rationality, and the scheduling system removes it from the candidate set.
[0093] Through the above three-level progressive screening, multiple candidate nodes that simultaneously meet the hard configuration requirements, are safe for real-time load, and have a reliable history of operation are finally obtained. If no node meets the conditions after screening, the Pod to be scheduled enters the scheduling waiting queue until a node in the cluster meets the constraints and the screening process is triggered again.
[0094] In some embodiments, the method further includes: Based on the tenant priority and resource quota rules in the preset scheduling policy, scheduling weights are assigned to Pods of different tenants. Pods run on nodes of member clusters in a single-cluster architecture or on nodes of member clusters in a multi-cluster architecture.
[0095] During the evaluation process of candidate nodes, a tenant isolation factor is introduced to avoid the same node hosting too many Pods from different tenants.
[0096] Monitor the resource usage of each tenant's Pod and dynamically adjust tenant resource quotas and scheduling priorities.
[0097] In some embodiments, during the suitability scoring stage of target node decision-making, a dedicated scheduling weight is first assigned to the Pods of different tenants based on the tenant priority and resource quota rules explicitly defined in the preset scheduling strategy. This weight is suitable for the scheduling scenarios of member cluster nodes in a single-cluster architecture and the node scheduling scenarios of each member cluster in a multi-cluster architecture. The scheduling weight is calculated using a two-factor weighting method. The tenant priority factor is preset according to the tenant level (e.g., 1.2 for core business tenants, 1.0 for ordinary tenants, and 0.8 for test tenants), and the resource quota usage factor is dynamically calculated based on the tenant's current resource utilization rate (e.g., 1.1 when the tenant's used resources account for less than 50% of the quota, and 0.9 when it is higher than 80%). The final scheduling weight = tenant priority factor × resource quota usage factor serves as the basis for subsequent suitability scoring correction, ensuring that high-priority tenants and tenants with reasonable resource usage receive better scheduling opportunities.
[0098] In the AI scoring process for candidate nodes, a tenant isolation factor is introduced to achieve resource isolation and security protection between tenants, avoiding resource contention or data security risks caused by too many different tenants hosting Pods on the same node. The calculation of the tenant isolation factor is based on the tenant distribution of the Pods currently hosted by the candidate node. First, the number of different tenants on the node is counted, and the tenant diversity coefficient is calculated (e.g., the coefficient is 0.4 when the node hosts 2 different tenants, 0.2 when 3-4 tenants, and 0.1 when 5 or more tenants). Then, the final factor value is obtained through the formula "Tenant Isolation Factor = 1 - Tenant Diversity Coefficient".
[0099] This factor is multiplied by the candidate node’s basic fit score (a weighted sum of resource fit score, performance prediction score, and stability score) and the Pod scheduling weight to obtain the corrected final fit score, i.e., “final score = basic fit score × Pod scheduling weight × tenant isolation factor”. This allows nodes that carry fewer different tenants to get higher scores and guides Pods to be scheduled to nodes with a more balanced tenant distribution.
[0100] Meanwhile, the scheduling system continuously monitors the resource usage of each tenant's Pods, establishing a dynamic adjustment mechanism for tenant resource usage to ensure the rationality of resource quotas and scheduling priorities. Monitoring dimensions include real-time utilization rates, peak usage durations, resource idle periods, and the number of quota shortage alarms for each tenant's core resources such as CPU, memory, storage, and network. The scheduling system sets adjustment threshold trigger conditions: if a tenant's resource utilization rate is consistently below 30% of its quota for a long period (e.g., 7 consecutive days), it indicates a mismatch between its resource needs and the current quota, and the scheduling system dynamically lowers its resource quota limit (e.g., by 20%) and simultaneously reduces its tenant priority factor (e.g., from 1.0 to 0.9); if a tenant frequently experiences resource quota shortages (e.g., Pod scheduling failures due to resource shortages 5 times in the past 3 days), and its continuous resource utilization rate is above 80% of its quota, provided the overall cluster resources are sufficient, its resource quota limit is dynamically increased (e.g., by 30%), and its priority factor is maintained or increased; for resource stress scenarios caused by sudden traffic surges in core tenants, their priority factor can be temporarily increased to 1.5 to ensure the scheduling priority of core services. All adjustments are synchronized to the preset scheduling strategy storage module and AI inference engine to ensure that subsequent scheduling decisions are adapted to the latest tenant configurations in a timely manner.
[0101] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0102] Corresponding to the intelligent scheduling method of the container orchestration system in the above embodiment, Figure 6This is a schematic diagram of the structure of an intelligent scheduling system provided in an embodiment of this application. (Refer to...) Figure 6 The intelligent scheduling system includes: Processor 601 is used to obtain cluster operation data of member clusters in a single-cluster architecture or multi-cluster architecture of a container orchestration system deployment.
[0103] The member cluster consists of multiple nodes, which are servers or computing units that host the operation of container groups (Pods). Cluster operation data includes node resource utilization, Pod performance, application performance metrics, and load change data.
[0104] The scheduling plugin 602 is used to respond to a scheduling request received from a Pod to be scheduled, to filter multiple candidate nodes from the member cluster that meet the Pod's running constraints; to determine the target node among multiple candidate nodes based on cluster running data, the feature information of the Pod to be scheduled, and the preset scheduling strategy of the AI inference engine; to bind the Pod to be scheduled to the target node and monitor the running status data of the Pod after it is bound to the target node; and to optimize the decision model of the AI inference engine and the preset scheduling strategy based on the running status data to form a scheduling closed loop.
[0105] The intelligent scheduling system provided in this application embodiment is applicable to mainstream container orchestration platforms such as Kubernetes, and can flexibly adapt to single-cluster or multi-cluster architectures. It is particularly suitable for executing the intelligent scheduling method of the aforementioned container orchestration system. The intelligent scheduling system mainly includes two major components: a processor and a scheduling plugin. These components work together to complete the entire intelligent scheduling process from data collection, scheduling decision-making, Pod deployment to policy optimization, ensuring the resource utilization and operational stability of the container orchestration system.
[0106] The processor is deployed in the control plane of the container orchestration system, and distributed collection nodes are set up in each member cluster of a single-cluster architecture or in each member cluster of a multi-cluster architecture. The member cluster consists of multiple nodes, which are physical servers, virtual computing units, and other computing units that host the running of container Pods. The processor establishes a secure communication link with the native Metrics API of the container orchestration system, the monitoring agent deployed at the node level (such as Prometheus Node Exporter), and the hardware management module to collect cluster operation data in real time. The collection scope comprehensively covers four categories of data: node resource utilization, Pod running performance, application performance indicators, and load change data.
[0107] The data includes: node resource utilization (such as CPU utilization, memory usage, disk I / O throughput, and network bandwidth usage); Pod performance (such as startup time, runtime, actual resource consumption, and number of restarts); application performance metrics (such as response latency, throughput, and error rate of applications hosted by Pods); and load change data (such as load fluctuation trends of the cluster as a whole and individual nodes, and Pod creation / destruction frequency). The processor standardizes and converts the collected data before synchronizing it to a time-series database in real time. Simultaneously, a data association index is established to ensure that cluster operation data is traceably linked to subsequent scheduling decision records and Pod operation status data, providing comprehensive and reliable data support for the scheduling plugin's decision-making process and model optimization.
[0108] The scheduling plugin is also deployed in the control plane of the container orchestration system, establishing communication connections with the processor, AI inference engine, time-series database, and API servers of each member cluster. Its core functions are implemented step by step in the following process: First, the scheduling plugin listens for scheduling requests from the container orchestration system in real time. When it receives a scheduling request for a Pod to be scheduled (this request is either submitted by the user as a Pod creation command or automatically triggered by the system to expand the Pod), it immediately starts the candidate node screening process. During the screening process, the scheduling plugin obtains the latest cluster operation data of the member clusters from the time-series database, and combines it with the operation constraints of the Pod to be scheduled. It then uses a three-layer progressive verification to screen multiple candidate nodes that meet the requirements: First, it filters out nodes that do not meet the hard conditions based on the Pod's resource request volume, node affinity constraints, taint tolerance rules, and hardware requirements; then, it excludes nodes whose resource utilization exceeds a preset threshold based on the real-time node load status; finally, it removes nodes that have caused Pod operation abnormalities in the past by referring to historical scheduling effect data, forming a candidate node set.
[0109] After identifying candidate nodes, the scheduling plugin, based on cluster operation data provided by the time-series database and the characteristic information of the Pods to be scheduled (including core attributes such as the application identifier, tenant information, resource requirement specifications, business priority, and performance sensitivity of the Pod), calls the pre-deployed AI inference engine. Combining this with the preset scheduling strategies integrated into the AI inference engine (including tenant priority weights, resource quota rules, task scheduling constraints, and service quality assurance parameters), it initiates the target node decision-making process. The scheduling plugin first obtains the real-time operating status of each candidate node, integrates it with Pod characteristic information, cluster operation data, and load prediction data, and then performs standardized processing and feature extraction. The AI inference engine calculates the resource suitability score, performance prediction score, and stability score of each candidate node. This score is then combined with the weight configuration of the preset scheduling strategy, Pod scheduling weights, and tenant isolation factors to obtain the final suitability score for each candidate node. The scheduling plugin selects the target node in descending order of score. If multiple nodes have the highest score, the final target node is determined by combining auxiliary rules such as the node's historical scheduling success rate and hardware configuration priority.
[0110] After identifying the target node, the scheduling plugin binds the Pod to be scheduled to the target node through the API interface of the container orchestration system. Specifically, it writes the identifier of the target node into the Pod's .spec.nodeName field, triggering the kubelet component on the target node to perform Pod creation and container instance startup operations. At the same time, the scheduling plugin continuously monitors the runtime status data of the Pod after it is bound to the target node, including the actual Pod startup time, runtime stability, resource consumption fluctuations, application performance, and whether resource contention or performance anomalies occur. This runtime status data is transmitted back to the time-series database in real time and stored in association with the corresponding scheduling decision records.
[0111] Finally, the scheduling plugin initiates the optimization process based on the runtime status data stored in the time-series database, forming a scheduling closed loop: the scheduling plugin compares the actual runtime status data of the Pod with the prediction results output by the AI inference engine during scheduling item by item, and obtains the deviation value of each indicator through a preset deviation calculation algorithm. If the deviation value exceeds the preset threshold, the model parameters of the AI inference engine decision model and the candidate node scoring weights are adjusted immediately. At the same time, the scheduling plugin continuously accumulates runtime status data and historical scheduling records to form a historical dataset covering single-cluster or multi-cluster scenarios. It periodically calls the model training module of the AI inference engine and uses the historical dataset to incrementally train and update the decision model (deep neural network model and reinforcement learning model), dynamically optimizing the preset scheduling strategy, ensuring that the AI inference engine's decision model and the preset scheduling strategy can continuously adapt to changes in cluster load, and continuously improve the accuracy and optimization effect of scheduling decisions.
[0112] It is understood that the intelligent scheduling system embodiments and any implementation methods correspond to the intelligent scheduling method embodiments and any implementation methods of the container orchestration system, respectively. The technical effects corresponding to the intelligent scheduling system embodiments and any implementation methods can be found in the aforementioned technical effects corresponding to the intelligent scheduling method embodiments and any implementation methods of the container orchestration system, and will not be repeated here.
[0113] It should be noted that the intelligent scheduling system provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0114] The functional units and modules in the above embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of the embodiments of this application.
[0115] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0116] This application also provides an electronic device, which includes one or more processors and a memory; The memory is coupled to one or more processors. The memory is used to store computer program code, which includes computer instructions. One or more processors invoke the computer instructions to cause the electronic device to execute the intelligent scheduling method of the container orchestration system described above.
[0117] Figure 7This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 700 can be a mobile phone, smart screen, tablet computer, wearable electronic device, in-vehicle electronic device, augmented reality (AR) device, virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), projector, or a communication device such as a server, storage device, or base station, or a smart car, etc. This application embodiment does not impose any limitations on the specific type of electronic device.
[0118] The memory 701 can be used to store computer software programs 702 and modules. The processor 703 executes various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory 701. The memory 701 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device (such as audio data, telephone directory, etc.). In addition, the memory 701 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0119] The processor 703 may include one or more processors such as a central processing unit (CPU), an application processor (AP), and a baseband processor. The processor can serve as the nerve center and command center of the wireless router. The processor 703 can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. The memory 701 can be used to store executable program code, including instructions. The processor 703 executes various functional applications and data processing of the network device by running the instructions stored in the memory. The memory 701 may include a program storage area and a data storage area, such as storing data for audio signals to be played. For example, the memory may be Double Data Rate Synchronous Dynamic Random Access Memory (DDR) or Flash memory.
[0120] This application also provides a computer-readable storage medium storing computer instructions; when the computer-readable storage medium is used on an electronic device, it causes the electronic device to execute the intelligent scheduling method of the container orchestration system described above.
[0121] The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or can include one or more data storage devices such as servers or data centers that can be integrated with media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media, or semiconductor media (e.g., solid-state disks (SSDs)).
[0122] This application also provides a computer program product containing computer instructions, which, when run on an electronic device, enables the electronic device to execute the intelligent scheduling method of the container orchestration system described above.
[0123] The computer storage medium and computer program product provided in the embodiments of this application are used to execute the methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects corresponding to the methods provided above, and will not be repeated here.
[0124] In the above embodiments, implementation can also be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, Digital Subscriber Line, DSL) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc., and the storage medium can also include combinations of the above types of memory.
[0125] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0126] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments claimed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0127] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0128] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0129] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. An intelligent scheduling method of a container orchestration system, characterized in that, The method comprises the following steps: acquiring cluster running data of a member cluster in a single-cluster architecture or a multi-cluster architecture deployed by a container orchestration system, the member cluster being composed of a plurality of nodes, the nodes being servers or computing units carrying a container group Pod, the cluster running data including node resource utilization, Pod running performance, application performance indicators, and load change data; in response to receiving a scheduling request of a to-be-scheduled Pod, screening a plurality of candidate nodes from the member cluster that meet Pod running constraint conditions; determining a target node from the plurality of candidate nodes based on the cluster running data, feature information of the to-be-scheduled Pod, and a preset scheduling strategy of an AI inference engine; binding the to-be-scheduled Pod to the target node and monitoring running state data of the Pod after being bound to the target node; optimizing a decision model of the AI inference engine and the preset scheduling strategy based on the running state data, forming a scheduling closed loop.
2. The method of claim 1, wherein, The method further comprises the following steps: based on the cluster running data, the feature information of the to-be-scheduled Pod, and the preset scheduling strategy of the AI inference engine, scoring the adaptability of each candidate node to obtain a scoring result of each candidate node; determining the target node from the plurality of candidate nodes according to the scoring result of each candidate node.
3. The method of claim 2, wherein, The method further comprises the following steps: obtaining real-time running states of each candidate node through a state-aware interface, wherein the real-time running states include CPU load, memory occupation, I / O performance, and hardware topology information; extracting feature information of the to-be-scheduled Pod, wherein the feature information includes a tenant to which the to-be-scheduled Pod belongs, resource demand, performance sensitivity, and business priority; inputting the real-time running states and the feature information into the AI inference engine to predict running effects of the to-be-scheduled Pod in each candidate node based on the preset scheduling strategy through a decision model of the AI inference engine; outputting an adaptability score of each candidate node as the scoring result based on the running effects of the to-be-scheduled Pod in each candidate node, wherein the adaptability score is used to quantify an expected running efficiency and resource adaptability degree of the Pod in the corresponding candidate node.
4. The method of claim 1, wherein, The method further comprises the following steps: standardizing the cluster running data and the feature information of the to-be-scheduled Pod to extract a key feature vector; calculating a resource adaptability score, a performance prediction score, and a stability score of each candidate node through a decision model of the AI inference engine; and weighting and summing the scores of each candidate node according to the weights of the preset scheduling strategy to obtain a final fitness score of each candidate node; ranking the candidate nodes according to the final fitness scores of the candidate nodes from high to low, and selecting a node at the top of the ranking as a target node.
5. The method of claim 1, wherein, The optimization of the decision model of the AI inference engine and the preset scheduling strategy based on the running state data forms a scheduling closed loop, which includes: comparing the running state data with the predicted results at the time of scheduling decision to obtain a deviation value, wherein the running state data includes start-up time, running stability, resource consumption deviation, and service quality indicators; if the deviation value exceeds a preset threshold, adjusting the model parameters of the decision model of the AI inference engine and the candidate node score weights; accumulating the running state data to form a historical data set, the historical data set containing historical running data of the member clusters of the single-cluster architecture or each member cluster in the multi-cluster architecture, wherein the historical running data of each member cluster is formed by the historical running data of all nodes contained therein; periodically training and updating the decision model of the AI inference engine using the historical data set.
6. The method according to any one of claims 1 to 5, characterized in that, Before screening a plurality of candidate nodes that meet the Pod running constraint conditions from the member clusters in response to receiving a scheduling request of a Pod to be scheduled, the method further includes: receiving a custom resource object submitted by an external system or a user through a declarative interface of a custom controller, the custom resource object including a scheduling strategy configuration object and a load prediction object; receiving the preset scheduling strategy and the load prediction data through the custom resource object; monitoring changes in the custom resource object through the custom controller, synchronizing the updated preset scheduling strategy to the AI inference engine and the scheduling process, and taking the load prediction data as input features of the AI inference engine; adjusting the Pod running constraint conditions and the fitness score weights according to the load prediction data, so that the preset scheduling strategy dynamically links with load changes.
7. The method of claim 6, wherein, The preset scheduling strategy includes tenant priority weights, resource quota rules, task scheduling constraints, and service quality guarantee parameters; the load prediction data includes load change trends, resource supply and demand prediction results, and business traffic peak forecasts for future periods, the load change trends corresponding to the load changes of the member clusters of the single-cluster architecture or each member cluster in the multi-cluster architecture, the load changes of the member clusters being formed by the load changes of all nodes contained therein.
8. The method according to any one of claims 1 to 5, characterized in that, The method further includes a multi-cluster scheduling step: maintaining a global resource view across clusters through a federal scheduling controller, the global resource view covering resource adequacy, network delay, and running cost information of all member clusters in the multi-cluster architecture, the resource adequacy of each member cluster being calculated by aggregating the available resources of all nodes contained in the member cluster; In response to receiving the cross-cluster scheduling request, candidate clusters satisfying the cross-cluster scheduling requirement are screened based on the global resource view, the candidate clusters belong to member clusters in the multi-cluster architecture, and each candidate cluster contains multiple nodes; A global AI decision model is called to evaluate the adaptation degree of each candidate cluster based on preset evaluation factors, wherein the evaluation factors include at least one of resource matching degree, task execution efficiency, operation cost, and disaster recovery priority; The to-be-scheduled Pod is allocated to a target cluster with the highest adaptation degree, and the target cluster is coordinated to perform candidate node screening operation and Pod binding operation, wherein the target cluster is one of the candidate clusters, and the candidate node screening operation is to screen nodes meeting the Pod running constraint condition from multiple nodes contained in the target cluster; When the running state of any of the member clusters is overloaded or fails, the Pod bound in any of the member clusters is migrated from the node contained in any of the member clusters to the node contained in other normal member cluster.
9. The method of claim 8, wherein, The coordination of the target cluster to perform candidate node screening operation and Pod binding operation includes: The cross-cluster scheduling instruction of the Pod is synchronized to the scheduling plug-in of the target cluster through the federal scheduling controller; The node allocation of the Pod in the target cluster is completed through the scheduling plug-in of the target cluster according to the candidate node screening process, the AI scoring process and the binding process; The running state of the Pod in the target cluster is synchronized in real time through the federal scheduling controller to form a whole-process monitoring of cross-cluster scheduling.
10. The method according to any one of claims 1 to 5, characterized in that, The decision model of the AI inference engine includes a deep neural network model and a reinforcement learning model, and the decision model is trained in the following manner: Based on historical cluster running data and scheduling effect data, a model training data set is constructed, the historical cluster running data includes historical data of the single-cluster architecture or each member cluster in the multi-cluster architecture, the historical data of each member cluster is formed by aggregating the historical data of all nodes contained therein, and the training data set contains node state, Pod feature information and scheduling result label under different load scenarios; The deep neural network model is trained based on the model training data set in a supervised learning manner to learn the mapping relationship between node state and Pod running effect; The reinforcement learning model is trained based on the model training data set in a reinforcement learning manner, with maximizing resource utilization and minimizing task delay as reward targets, to optimize the scheduling decision strategy; Periodically, the model parameters of the deep neural network model and the reinforcement learning model are updated using newly added historical cluster running data and scheduling effect data to enhance the adaptability of the decision model to the load changes of the member clusters in the single-cluster architecture or the multi-cluster architecture.
11. The method according to any one of claims 1 to 5, characterized in that, The Pod running constraint condition includes: According to the resource request amount of the Pod, the node affinity constraint, the stain tolerance rule and the hardware demand, the nodes that do not meet the hard conditions are filtered out; In combination with the real-time node load status, nodes with resource usage exceeding a preset threshold are excluded; Referring to historical scheduling effect data, nodes that have caused abnormal Pod operation in the past are eliminated to obtain a plurality of candidate nodes.
12. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Based on the tenant priority and resource quota rules in the preset scheduling strategy, scheduling weights are assigned to Pods of different tenants, which run on nodes of member clusters of the single-cluster architecture or nodes of member clusters in the multi-cluster architecture; In the scoring process of the candidate nodes, a tenant isolation factor is introduced to avoid the same node carrying too many Pods of different tenants; The resource usage of each tenant Pod is monitored, and the tenant resource quota and scheduling priority are dynamically adjusted.
13. An intelligent dispatch system characterized by, An intelligent scheduling method for a container orchestration system as claimed in any one of claims 1 to 12, comprising: a processor configured to obtain cluster operation data of a member cluster in a single-cluster architecture or a multi-cluster architecture deployed by a container orchestration system, the member cluster comprising a plurality of nodes, the nodes being servers or computing units carrying Pod operation of a container group, the cluster operation data including node resource utilization, Pod operation performance, application performance indicators, and load change data; a scheduling plug-in configured to, in response to receiving a scheduling request of a to-be-scheduled Pod, screen a plurality of candidate nodes satisfying a Pod operation constraint condition from the member cluster, determine a target node from the plurality of candidate nodes based on the cluster operation data, feature information of the to-be-scheduled Pod, and a preset scheduling strategy of an AI inference engine, bind the to-be-scheduled Pod to the target node, and monitor operation state data of the Pod after being bound to the target node, and optimize a decision model of the AI inference engine and the preset scheduling strategy based on the operation state data to form a scheduling closed loop.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor, when executing the computer program, causes the electronic device to implement the method as claimed in any one of claims 1 to 12.
15. A computer program product, characterised in that, The computer program, when executed, causes the method as claimed in any one of claims 1 to 12 to be performed.
16. A computer-readable storage medium, the computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by the processor, implements the method as claimed in any one of claims 1 to 12.
Citation Information
Cited By
Cloud resource scheduling method and device, equipment, storage medium and program product
CN122044889A