Kubernetes task priority scheduling method and device based on AI

Through the AI-based task priority scheduling method, task resource allocation and priority are dynamically adjusted, which solves the problems of unbalanced resource allocation and insufficient task importance management in the Kubernetes cluster, and achieves more efficient resource utilization and timely response to critical tasks.

CN120849036APending Publication Date: 2025-10-28DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510853400.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

The scheduling mechanism of traditional Kubernetes clusters cannot perceive the actual resource consumption during task execution, resulting in a large discrepancy between resource allocation and usage, which affects resource utilization efficiency and task execution stability. Furthermore, the lack of fine-grained management of the differences in task importance means that critical tasks cannot obtain priority execution opportunities.

Method used

The AI-based task priority scheduling method obtains historical information and real-time resource information of tasks, and uses the task priority prediction model to dynamically adjust the resource allocation and priority of tasks, achieving fine-grained scheduling and supporting timely response and stable operation of high-priority tasks.

Benefits of technology

It improves the scheduling flexibility and resource utilization efficiency of the Kubernetes cluster, ensures timely response and stable operation of high-priority tasks, avoids resource shortage or waste, and improves the overall service quality of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849036A_ABST
    Figure CN120849036A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of task scheduling, and discloses an AI-based Kubernetes task priority scheduling method and device, and the method comprises the steps: obtaining the historical task information of at least one task in a Kubernetes cluster, the historical task information comprises the historical task execution information of each task, business importance levels and business timeliness information of businesses to which the tasks belong; inputting historical task information of at least one task into a task priority prediction model to obtain a task priority score of each task; acquiring real-time resource information of the Kubernetes cluster and real-time task execution information of each task; and scheduling each task based on the real-time resource information, the real-time task execution information of each task and the task priority score of each task. And a service fine-grained and dynamic resource scheduling strategy is realized. The resource allocation of each task can be flexibly adjusted according to the real-time resource of the current Kubernetes cluster and the task execution information of each task, and the resource can be dynamically preempted according to the task priority score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of task scheduling technology, and in particular to an AI-based Kubernetes task priority scheduling method and apparatus. Background Technology

[0002] With the rapid development of cloud computing and containerization technologies, Kubernetes clusters have become the mainstream container orchestration system, playing a core role in task scheduling and resource management. However, traditional Kubernetes cluster scheduling algorithms have revealed several limitations in practical applications. First, their scheduling mechanism is mainly based on static resource allocation strategies. This strategy cannot perceive the actual resource consumption during task execution, leading to a significant discrepancy between resource allocation and actual resource usage. On the one hand, this can easily result in excessive resource reservation and waste; on the other hand, in scenarios of sudden load or resource contention, resource shortages may occur, leading to task execution failures or delays, affecting the overall stability of the system, the overall resource utilization of the Kubernetes cluster, and task execution efficiency. Second, in terms of priority management, there is a lack of fine-grained differentiation of the importance of tasks at the business level. As a result, critical tasks often fail to obtain priority execution opportunities in scenarios with intense resource competition, thus affecting the service quality of the system. Therefore, there is an urgent need for a scheduling mechanism that can adapt to dynamic changes in resource demand and support fine-grained priority scheduling to improve the scheduling flexibility and resource utilization efficiency of Kubernetes clusters, and ensure the timely response and stable operation of high-priority tasks. Summary of the Invention

[0003] The purpose of this invention is to provide at least one AI-based Kubernetes task priority scheduling method and apparatus, which can at least solve the technical problems of poor scheduling flexibility and low resource utilization efficiency of Kubernetes clusters, and at least achieve the technical effect of improving the scheduling flexibility and resource utilization efficiency of Kubernetes clusters.

[0004] To address the aforementioned technical problems, at least one embodiment of this application provides an AI-based Kubernetes task priority scheduling method, comprising: acquiring historical task information of at least one task in a Kubernetes cluster, the historical task information including historical task execution information of each task, the business importance level and business timeliness information of the business to which each task belongs; inputting the historical task information of at least one task into a task priority prediction model to obtain a task priority score for each task; acquiring real-time resource information of the Kubernetes cluster and real-time task execution information of each task; and scheduling each task based on the real-time resource information, the real-time task execution information of each task, and the task priority score of each task.

[0005] This solution accurately obtains task priority scores that indicate the differences in importance among tasks at the business level, based on the business importance level, business timeliness information, and task execution information of each task. Tasks are scheduled based on their priority scores, real-time resource information, and real-time task execution information, enabling a fine-grained and dynamic resource scheduling strategy. It not only flexibly adjusts resource allocation for each task based on the current real-time resources of the Kubernetes cluster and the task execution information, but also supports dynamic resource preemption based on task priority scores. This avoids significant discrepancies between allocated resources and the actual resources required by the tasks, preventing resource shortages or waste, and improving the scheduling flexibility and resource utilization efficiency of the Kubernetes cluster. It also ensures timely response and stable operation of high-priority tasks.

[0006] In some examples, tasks are scheduled based on real-time resource information, real-time task execution information, and task priority scores. This includes: adjusting the task priority score for each task based on real-time resource information, real-time task execution information, and task priority scores to obtain a new task priority score; and scheduling tasks based on real-time resource information, real-time task execution information, and the new task priority scores.

[0007] In some examples, real-time task execution information includes task resource requirements and task running nodes. Real-time resource information includes the remaining resources of each node. Based on the real-time resource information, the real-time task execution information, and the task priority score, the task priority score is adjusted to obtain a new task priority score. This includes: for critical tasks whose task priority score is higher than a preset priority score, obtaining the remaining resources of the target node of the critical task from the remaining resources of each node; if the task resource requirements are higher than a preset task requirement threshold and higher than the remaining resources of the target node, then determining the increase in the task priority score of the critical task based on the difference between the task resource requirements and the remaining resources of the target node; and using the sum of the increase and the task priority score as the new task priority score of the critical task. Here, the preset task requirement threshold corresponds one-to-one with the critical tasks.

[0008] In some examples, real-time task execution information also includes a task delay impact value. The task priority score is adjusted based on real-time resource information, real-time task execution information, and task priority score to obtain a new task priority score. This includes: determining an increase in the task priority score based on the task delay impact value, where the task delay impact value represents the total impact of delayed execution of the task on tasks with execution priority scores higher than a preset high priority threshold; and summing the increase value with the task priority score as the new task priority score.

[0009] In some examples, real-time resource information includes the remaining resources of the Kubernetes cluster. The task priority score is adjusted based on the real-time resource information, the real-time task execution information, and the task priority score to obtain a new task priority score. For example, if the remaining cluster resources are less than the cluster resource threshold and the task priority score is lower than the preset low priority threshold, the task priority score is set to 0.

[0010] In some examples, tasks are scheduled based on real-time resource information, real-time task execution information, and new task priority scores. This includes: sorting the task priority scores of each task from highest to lowest to obtain a task priority order; executing tasks sequentially based on the task priority order; determining the node scheduling score of each node in the Kubernetes cluster based on the new task priority score, real-time task execution information, and real-time resource information; selecting the node with the highest scheduling score to schedule the task; and updating the real-time resource information.

[0011] In some examples, the method further includes: obtaining sample task information for at least one task in the Kubernetes cluster, including sample task execution information for each task, the importance level of the business to which each task belongs, and the timeliness information of the business; performing data preprocessing on the sample task information of at least one task to obtain target sample information for each task, including data cleaning and data standardization; performing feature engineering on the target sample information for each task to obtain feature information and label information for each task; and training an initial task priority prediction model based on the target sample information, feature information, and label information for each task to obtain a task priority prediction model that meets preset requirements.

[0012] At least one embodiment of this application also provides an AI-based Kubernetes task priority scheduling device, comprising: an acquisition unit, configured to acquire historical task information of at least one task in a Kubernetes cluster, the historical task information including historical task execution information of each task, the business importance level and business timeliness information of the business to which each task belongs; an input unit, configured to input the historical task information of at least one task into a task priority prediction model to obtain a task priority score for each task; the acquisition unit is further configured to acquire real-time resource information of the Kubernetes cluster and real-time task execution information of each task; and a scheduling unit, configured to schedule each task based on the real-time resource information, the real-time task execution information of each task, and the task priority score of each task.

[0013] At least one embodiment of this application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the above-described AI-based Kubernetes task priority scheduling method.

[0014] At least one embodiment of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described AI-based Kubernetes task priority scheduling method. Attached Figure Description

[0015] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.

[0016] Figure 1 This is a flowchart illustrating an AI-based Kubernetes task priority scheduling method provided in one embodiment of this application;

[0017] Figure 2 This is a schematic diagram of the structure of an AI-based Kubernetes task priority scheduling system provided in one embodiment of this application;

[0018] Figure 3 This is a schematic diagram of the structure of a task priority prediction model provided in one embodiment of this application;

[0019] Figure 4 This is a schematic diagram of a task scheduling process provided in one embodiment of this application;

[0020] Figure 5This is a schematic diagram of an AI-based Kubernetes task priority scheduling device provided in another embodiment of this application;

[0021] Figure 6 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the various embodiments of this application to help readers better understand this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for the convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.

[0023] It should be noted that the acquisition or use of data in the embodiments of this application requires the user's consent. The relevant data can only be obtained after the user's authorization, and the acquisition or use of the data complies with the provisions of relevant laws and regulations.

[0024] To facilitate understanding of the embodiments of this application, we will first introduce the relevant content of the AI-based Kubernetes task priority scheduling method.

[0025] Kubernetes: An open-source container orchestration platform that enables automated deployment, scaling, and management of containerized applications. It optimizes resource utilization by allocating Pods (container groups) to cluster nodes through a scheduler. In this solution, it serves as the foundational platform for task scheduling.

[0026] Pod: The smallest deployable and manageable unit of computing in Kubernetes. It is a collection of closely related containers that share storage and network resources. In this scheme, tasks run in the form of Pods.

[0027] Scheduler: A core component of Kubernetes, responsible for assigning Pods to appropriate nodes.

[0028] Mean Squared Error Loss (MSE): Used to measure the average error between the model's predicted values ​​and the actual values.

[0029] Cross-entropy loss function: commonly used in classification tasks to measure the difference in probability distribution between the predicted result and the true label.

[0030] With the rapid development of cloud computing and containerization technologies, Kubernetes clusters have become the mainstream container orchestration system, playing a core role in task scheduling and resource management. However, traditional Kubernetes cluster scheduling algorithms have revealed several limitations in practical applications. First, their scheduling mechanism is mainly based on static resource allocation strategies. This strategy cannot perceive the actual resource consumption during task execution, leading to a significant discrepancy between resource allocation and actual resource usage. On the one hand, this can easily result in excessive resource reservation and waste; on the other hand, in scenarios of sudden load or resource contention, resource shortages may occur, leading to task execution failures or delays, affecting the overall stability of the system, the overall resource utilization of the Kubernetes cluster, and task execution efficiency. Second, in terms of priority management, there is a lack of fine-grained differentiation of the importance of tasks at the business level. As a result, critical tasks often fail to obtain priority execution opportunities in scenarios with intense resource competition, thus affecting the service quality of the system. Therefore, there is an urgent need for a scheduling mechanism that can adapt to dynamic changes in resource demand and support fine-grained priority scheduling to improve the scheduling flexibility and resource utilization efficiency of Kubernetes clusters, and ensure the timely response and stable operation of high-priority tasks.

[0031] To address the technical problems of poor scheduling flexibility and low resource utilization efficiency in Kubernetes clusters, this invention proposes an AI-based Kubernetes task priority scheduling method. The implementation details of the AI-based Kubernetes task priority scheduling method in this embodiment are described below. The following content is only for the convenience of understanding and is not necessary for implementing this solution.

[0032] Example 1:

[0033] The AI-based Kubernetes task priority scheduling method in this embodiment can be applied to electronic devices with communication, computing, and data storage capabilities, specifically to Kubernetes clusters. Its specific process can be as follows: Figure 1 Shown, including:

[0034] Step 110: Obtain historical task information for at least one task in the Kubernetes cluster. The historical task information includes the historical task execution information of each task, the business importance level of the business to which each task belongs, and the business timeliness information.

[0035] like Figure 2 As shown, the data collection and preprocessing module collects data from the Kubernetes cluster to obtain historical task information for at least one task in the Kubernetes cluster.

[0036] Specifically, a task is the basic unit for realizing business, and business refers to the core services or functions provided by an enterprise or organization.

[0037] Specifically, historical task execution information refers to information related to task execution generated at historical moments. This information includes the execution time, resource usage information, and task dependencies. Specifically, the execution time includes the task start time and completion time. Resource usage information includes the actual usage of various resources during task execution, including at least CPU, memory, disk I / O, and network bandwidth. Task dependencies include the execution order of tasks and the dependencies between them.

[0038] For example, the lifecycle events of a task can be monitored through the Kubernetes API, that is, the task start timestamp is recorded when the task starts and the task end timestamp is recorded when the task is completed, so as to obtain the historical task execution time of each task based on the task start timestamp and task end timestamp.

[0039] For example, the Kubernetes resource monitoring interface can be used to periodically collect and save the actual usage of resources such as CPU, memory, disk I / O, and network bandwidth during task execution. This data can serve as a source of historical task resource usage information. For instance, the actual usage of various resources for each Pod on a node can be obtained through the api / v1 / nodes / {nodeName} / metrics interface. Here, each pod corresponds one-to-one with a task, and the pod is the execution vehicle for the task.

[0040] For example, the execution order of each task and the task execution dependencies of each task can be extracted by parsing the task configuration file (e.g., parsing a YAML-formatted task configuration file).

[0041] Specifically, the business importance level is used to indicate the degree of importance of a business; the higher the business importance level, the greater the importance of the corresponding business.

[0042] Specifically, business timeliness information indicates the urgency of executing a business task in terms of time. If the business timeliness information indicates real-time processing, it means that the business needs to be processed in real time, indicating a high degree of urgency in terms of time. If the business timeliness information indicates batch processing, it means that the business needs to be executed in terms of time.

[0043] Step 120: Input the historical task information of at least one task into the task priority prediction model to obtain the task priority score for each task.

[0044] like Figure 2As shown, historical task information for at least one task is input into the task priority prediction model to obtain the task priority score for each task. The task priority prediction model is constructed based on the prediction model construction module.

[0045] Specifically, a task priority prediction model is a model used in task scheduling scenarios to predict the task priority score. The task priority score represents the priority of a task in being scheduled or in acquiring resources; a higher score indicates a greater priority in scheduling and resource acquisition. The output of the task priority prediction model can be numerical or categorical data. If numerical, the output is the task priority score; if categorical, the output is the task priority level. In this solution, the preferred output is the task priority score.

[0046] Specifically, the task priority score obtained by the task priority prediction model is predicted based on the historical data of each task. The task priority score can reflect the importance and urgency of each task under normal circumstances. In other words, it can indicate the priority of a task in acquiring resources under normal circumstances.

[0047] In some examples, such as Figure 3 As shown, the process of building a task priority prediction model by the prediction model building module is as follows: Based on the data collection and preprocessing module, sample task information of at least one task in the Kubernetes cluster is obtained. The sample task information includes the sample task execution information of each task, the importance level of the sample business to which each task belongs, and the timeliness information of the sample business. Based on the data collection and preprocessing module, the sample task information of at least one task is preprocessed to obtain the target sample information of each task. The data preprocessing includes data cleaning and data standardization. Based on the data collection and preprocessing module, the target sample information of each task is feature-engineered to obtain the feature information and label information of each task. The prediction model building module trains the initial task priority prediction model based on the target sample information, feature information, and label information of each task to obtain a task priority prediction model that meets the preset requirements.

[0048] In some cases, if no task priority prediction model exists after obtaining the historical task information of at least one of the aforementioned tasks, then the historical task information of the aforementioned at least one task can be used as sample task information for at least one task. Furthermore, if a task priority prediction model exists after obtaining the historical task information of the aforementioned at least one task, then the task priority prediction model can be directly used to predict the historical task information of at least one task to obtain the task priority score for each task.

[0049] Therefore, the data collection and preprocessing module breaks through the traditional single data acquisition mode, enabling collaborative collection of multi-source data such as Kubernetes cluster task execution time, resource usage, dependencies, and business metrics. Furthermore, the module cleanses and standardizes the collected data, providing high-quality, structured data for the task priority prediction model construction module, thus addressing the shortcomings of traditional data processing methods, such as incompleteness and insufficient feature mining. This improves the performance of the constructed task priority prediction model.

[0050] Specifically, sample task execution information refers to the information generated during task execution related to the task. This information includes the task execution time, resource usage information, and task dependencies. Specifically, the task execution time includes the task start time and task completion time. Resource usage information includes the actual usage of various resources during task execution, including at least CPU, memory, disk I / O, and network bandwidth. Task dependencies include the execution order of tasks and the task execution dependencies between them.

[0051] Specifically, the sample business importance level is used to indicate the degree of importance of a business. The higher the sample business importance level, the higher the importance of the corresponding business.

[0052] Specifically, the timeliness information of sample business operations is used to indicate the urgency of executing business operations in the time dimension. If the timeliness information is real-time processing, it means that the business needs to be processed in real time, indicating a high urgency in the time dimension. If the timeliness information is batch processing, it means that the urgency in the time dimension is low.

[0053] In some examples, data cleaning of sample task information for at least one task includes: using statistical methods to identify outliers in the sample task information for at least one task and removing outliers; for the sample task information of each task, obtaining the missing values ​​in the sample task information; if the missing value is numerical data, then using the mean imputation method, the median imputation method, or a machine learning model to impute the missing value; if the missing value is categorical data, then using the default value as the missing value.

[0054] Specifically, numerical data refers to data that can be processed mathematically, representing quantitative information. Categorical data represents qualitative information, such as the time-sensitive information represented through real-time and batch processing mentioned above.

[0055] For example, for multiple CPU utilization rates in the sample task information of each task, the mean and standard deviation of multiple CPU utilization rates are calculated, and data exceeding the mean plus 3 times the standard deviation are regarded as outliers and removed; for multiple sample task execution times in the sample task information of each task, the execution duration of each execution task is obtained based on the execution time of each sample task, resulting in multiple execution durations, and the sample task execution times corresponding to the longer or shorter execution durations among the multiple execution durations are corrected or deleted manually or by algorithms.

[0056] For example, if the missing value is memory usage and there are few missing values, the missing values ​​can be filled using the mean or median. If the missing value is memory usage and there are many missing values, machine learning models such as random forest regression can be used to predict and fill the missing values. If the missing value is time-sensitive information, it can be filled with the default value of "unknown".

[0057] In some examples, data standardization processing of sample task information for at least one task includes: converting time-type data in the sample task information of each task into target time-type data of a fixed length sequence; converting categorical data in the sample task information of each task into target categorical data in binary vector form; normalizing numerical data in the sample task information of each task to obtain target numerical data; and using the target time-type data, target categorical data, and target numerical data of each task as the target sample information of each task.

[0058] Specifically, time-based data refers to data used to represent time or moments in time.

[0059] For example, for time-based data, if the length of the time-based data is greater than the fixed-length sequence, it is truncated to make its length equal to the fixed-length sequence; if the length of the time-based data is less than the fixed-length sequence, it is padded with a preset value to make its length equal to the fixed-length sequence. The preset value can be 0 or other numerical values. Furthermore, to facilitate model processing, the fixed-length time-based data can be further converted into relative time interval data for model processing. For the specific implementation process of converting fixed-length time-based data into relative time interval data for model processing, please refer to existing technologies, which will not be elaborated further here.

[0060] For example, for categorical data, taking the importance level of sample business as an example, one-hot encoding can be used to convert the importance level of sample business into a binary vector form, which is convenient for the initial task priority prediction model to process. For details on one-hot encoding, please refer to existing technologies; further explanation is not provided here.

[0061] For example, for numerical data, taking CPU utilization as an example, a normalization formula is used to map it to the [0,1] interval to avoid interference from differences in units and improve the model training effect. For details regarding the normalization formula and the specific implementation process of data normalization, please refer to existing technologies; further details will not be elaborated here.

[0062] In some examples, feature engineering is performed on the target sample information of each task to obtain the feature information and label information of each task. This includes: calculating derived features on the target sample information for each task to obtain target derived information; constructing a task dependency graph based on the target sample information and extracting graph features from the task dependency graph; performing weighted fusion processing on each data in the target sample information to generate a target business priority score, and using the target business priority score as the label of the target sample information; and using the target derived information and graph features as the feature information of the target sample information.

[0063] Specifically, the target-derived information includes features such as task resource utilization rate, task resource utilization growth rate, and task execution time volatility. Task resource utilization rate refers to the average amount of system resources a task occupies during its execution; task resource utilization growth rate refers to the rate at which resource usage changes over time during task execution; and task execution time volatility refers to the stability of the execution time when a task is executed multiple times. For example, taking CPU utilization rate within task resource utilization as an example, the CPU utilization rate is obtained using the formula: CPU utilization rate = CPU usage / Total CPU usage.

[0064] Specifically, a task dependency graph describes the dependencies between tasks and their execution order. Graph features include topological characteristics such as node in-degree, out-degree, and shortest path length, reflecting the importance of the task corresponding to a node in the task dependency graph—that is, the degree to which the execution of the task corresponding to the node affects other tasks. In-degree refers to the number of edges pointing to the node, and out-degree refers to the number of edges originating from the node and pointing to other nodes. The shortest path length is the shortest edge between the task corresponding to the node and the starting node of the task dependency graph.

[0065] Specifically, the target service priority score indicates the order in which its corresponding service acquires resources, that is, the order in which the service is scheduled. The higher the target service priority score, the earlier its resource acquisition order. For example, the target service priority score is obtained by weighted summing of the sample service importance level, sample task execution information, and sample service timeliness information in the target sample information.

[0066] In some examples, such as Figure 3The initial task priority prediction model shown can be any of RNN, LSTM, or Transformer. Specifically, the choice of model depends on the data characteristics and application scenario. RNN is suitable for processing task data with time-series characteristics, capturing the dependencies between data sequences through hidden layer neuron connections. LSTM introduces gating mechanisms (including input gate, forget gate, and output gate) on top of RNN, effectively solving the problem of long-sequence dependencies and is suitable for long-term trend prediction. Transformer, based on a self-attention mechanism, can process data in parallel, offering high efficiency and accuracy when handling long sequences and complex features, and is particularly suitable for task data involving the fusion of multiple feature types.

[0067] In some examples, an initial task priority prediction model is trained based on the target sample information, feature information, and label information of each task to obtain a task priority prediction model that meets preset requirements. This includes: using the target sample information, feature information, and label information of each task as a training dataset; dividing the training dataset into a training set, a validation set, and a test set according to a preset ratio; training the initial task priority prediction model using the data in the training set to obtain an intermediate priority prediction model; in training the initial task priority prediction model based on the data in the training set, using the validation set to verify whether the trained initial task priority prediction model is an intermediate priority prediction model. If the model performance of the initial task priority prediction model no longer improves or overfitting occurs, it is considered an intermediate priority prediction model; and using the test set to test whether the intermediate priority prediction model meets the prediction requirements. If the verification meets the preset requirements, the intermediate priority prediction model is used as the task priority prediction model.

[0068] For example, such as Figure 3 As shown, the training dataset can be divided into a training set, a validation set, and a test set in a 7:1:2 ratio. The training set is used to train the initial task priority prediction model, allowing it to learn its parameters. The validation set is used to evaluate the model's performance and adjust hyperparameters during training, including the learning rate, the number of network layers, and the number of neurons. The test set is used to objectively evaluate whether the validated priority prediction model meets the prediction requirements, thus determining the task priority prediction model.

[0069] In some examples, the training process for training an initial task priority prediction model using training set data to obtain an intermediate priority prediction model is as follows: Training data from the training set is input into the initial task priority prediction model in batches. For each training data point, forward propagation is used to calculate the prediction result, and a loss function is used to calculate the loss between the prediction result and the label of that training data. Backpropagation is used to update the model parameters based on the loss value, resulting in the trained initial task priority prediction model. When the latest initial task priority prediction model is obtained based on a preset number of training data points, a validation set is used to determine if the latest initial task priority prediction model meets preset requirements. If it does, the latest initial task priority prediction model is used as the intermediate priority prediction model. If it does not meet the preset requirements, the process of re-exercising the steps of calculating the prediction result using forward propagation for each training data point, calculating the loss between the prediction result and the label of that training data using a loss function, and updating the model parameters based on the loss value using backpropagation is repeated until an intermediate priority prediction model is obtained.

[0070] Specifically, if the initial task priority prediction model outputs numerical data, its accuracy and mean squared error can be obtained, and its compliance with preset requirements can be determined based on these metrics. Specifically, if the accuracy is not lower than the preset accuracy threshold and the mean squared error is not higher than the preset mean squared error threshold, the initial task priority prediction model is considered to meet the preset requirements. If the output set contains categorical data, its accuracy and recall can be obtained, and its compliance with the prediction requirements can be determined based on these metrics. Specifically, if the accuracy is not lower than the preset accuracy threshold and the recall is not higher than the preset recall threshold, the intermediate priority prediction model is considered to meet the preset requirements. If the difference between the accuracy, recall, or mean squared error and the preset requirements exceeds a preset gap, the initial task priority prediction model can be adjusted by increasing or decreasing the number of network layers, adjusting the number of neurons, optimizing the learning rate, increasing the regularization strength, or using the Dropout method, or by any one or more of these methods.

[0071] Specifically, the loss function can be selected based on the data type of the prediction model's output based on the initial task priority. If the output is numerical data, the Mean Squared Error (MSE) loss function is used; if the output is categorical data, the Cross-Entropy loss function is used. Details regarding the MSE and Cross-Entropy loss functions (e.g., formulas) can be found in existing technologies and will not be elaborated upon here.

[0072] Specifically, updating model parameters based on loss values ​​using backpropagation includes: combining backpropagation with the Adam algorithm and updating model parameters based on loss values. The Adam (Adaptive Moment Estimation) algorithm combines the advantages of Momentum and RMSProp (Root Mean Square Propagation), automatically adjusting the learning rate of each parameter during backpropagation, making model training more efficient, stable, and with fast convergence and good stability.

[0073] For example, testing whether an intermediate priority prediction model meets prediction requirements using a test set includes: using data from the test set as input to the intermediate priority prediction model and obtaining the model's output; forming an output set based on the model's output obtained from the test set; if the data in the output set is categorical data, calculating precision and recall based on the corresponding output set, and determining whether the intermediate priority prediction model meets prediction requirements based on precision and recall; if the data in the output set is numerical data, calculating precision and mean squared error based on the output set, and determining whether the intermediate priority prediction model meets preset requirements based on precision and mean squared error. Specifically, when the precision is not lower than the precision threshold in the preset requirements, and the recall is not higher than the recall threshold in the preset requirements, the intermediate priority prediction model is considered to meet the preset requirements. When the precision is not lower than the precision threshold in the preset requirements, and the mean squared error is not higher than the mean squared error threshold in the preset requirements, the intermediate priority prediction model is considered to meet the preset requirements.

[0074] In some examples, the method also includes: after obtaining a task priority prediction model that meets preset requirements, saving the task priority prediction model in a specific format. This specific format can be HDF5 (Hierarchical DataFormat Version 5, a universal binary file format) or SavedModel. SavedModel is a complete model saving format.

[0075] Step 130: Obtain real-time resource information of the Kubernetes cluster and real-time task execution information of each task.

[0076] like Figure 2 As shown, real-time resource information of the Kubernetes cluster and real-time task execution information of each task can be obtained through the real-time monitoring system.

[0077] Specifically, real-time resource information describes the real-time resource status of the Kubernetes cluster and its nodes. This includes the remaining resources of each node (e.g., remaining CPU and memory), the remaining resources of the Kubernetes cluster (e.g., remaining CPU and memory), and the resource utilization of each node (e.g., CPU utilization, disk space utilization, and network bandwidth utilization). Furthermore, real-time resource information also describes the node usage status, which may include the health status of each node (e.g., healthy, faulty, or requiring maintenance), the number of historical node failures, resource usage fluctuations, and node network connectivity status (e.g., inter-node network latency and packet loss rate).

[0078] Specifically, real-time task execution information describes the task execution status, including: task resource requirements, task running nodes, task latency impact value, real-time task execution progress (e.g., percentage of completed steps), and real-time task execution status (e.g., whether the task encountered any anomalies, such as Pod crashes, container restarts, or task freezes). The task latency impact value represents the total impact of the task's delayed execution on tasks with a priority score higher than a preset high-priority threshold. A higher latency impact value indicates that the task should be scheduled and executed with greater priority.

[0079] Step 140: Schedule each task based on real-time resource information, real-time task execution information of each task, and task priority score of each task.

[0080] In some examples, such as Figure 4 As shown, in step 140 above, tasks are scheduled based on real-time resource information, real-time task execution information of each task, and task priority scores of each task, including steps 1401-1402:

[0081] Step 1401: For each task, adjust the task priority score based on real-time resource information, real-time task execution information, and task priority score to obtain a new task priority score.

[0082] like Figure 2 As shown, the priority dynamic adjustment module adjusts the task priority score based on real-time resource information, real-time task execution information, and task priority score to obtain a new task priority score.

[0083] Specifically, regarding the aforementioned step 1401, the specific implementation method of adjusting the task priority score based on real-time resource information, real-time task execution information of the task, and task priority score of the task to obtain a new task priority score can be found in optional Exemplary 1, optional Exemplary 2, and optional Exemplary 3.

[0084] Optionally, Example 1: During the actual execution of a task, there may be a surge in critical task demands and insufficient node resources. Scheduling tasks solely based on task priority scores obtained under normal circumstances may lead to untimely scheduling of critical tasks, resulting in the unavailability of critical services corresponding to those tasks. To avoid this, this solution proposes a task priority score adjustment method, specifically: Real-time task execution information includes task resource demand information and task running nodes. Real-time resource information includes the remaining resources of each node. Based on the real-time resource information, the task's real-time task execution information, and the task's task priority score, the task priority score is adjusted to obtain a new task priority score. This includes: for critical tasks with a task priority score higher than a preset priority score, obtaining the remaining resources of the target node of the critical task's running node from the remaining resources of each node; if the task resource demand information is higher than a preset task demand threshold and higher than the remaining resources of the target node, then determining the increase in the critical task's task priority score based on the difference between the task resource demand information and the remaining resources of the target node; and using the sum of the increase and the task priority score as the new task priority score for the critical task. Here, the preset task demand threshold corresponds one-to-one with each critical task.

[0085] Specifically, a critical task refers to a task that provides critical services, that is, a task with a higher priority score compared to the preset priority score. The preset priority score is the criterion for judging whether a task is a critical task, and it can be set according to actual needs, which will not be elaborated on here.

[0086] Specifically, node remaining resources refer to the amount of remaining available resources of a node.

[0087] Specifically, a task execution node refers to a node that executes a task.

[0088] Specifically, the preset task requirement threshold refers to the amount of resources required by a critical task under normal circumstances.

[0089] In some examples, the increment for the task priority score of a critical task is determined based on the difference between the task resource requirement information and the remaining resources of the target node. This includes multiplying the difference between the task resource requirement information and the remaining resources of the target node by a preset ratio, and using this product as the increment for the task priority score of the critical task. The preset ratio can be set according to actual needs, and can be set to 0.2, 0.5, 1, 2, or other values.

[0090] For example, when the task priority score of task 1 is higher than the preset priority score, task 1 is considered critical task 1. When the memory resource demand of critical task 1 surges, exceeding the preset memory demand threshold, and the available resources of the node are much less than the memory resource demand of critical task 1, the task priority score of the critical task is increased to improve the priority of the critical task in acquiring resources. This allows the critical task to acquire resources earlier than it would normally, ensuring that the critical task can acquire sufficient resources when node resources are insufficient and task resource demand surges, thus avoiding the problem of critical service unavailability due to failure to compete for resources.

[0091] Optionally, in Example 2, it should be understood that during the actual execution of tasks, there are interdependent relationships between tasks. Critical tasks may depend on non-critical tasks, and if non-critical tasks are not executed in a timely manner, it will severely affect the execution of critical tasks that depend on them. Therefore, scheduling tasks solely based on task priority scores determined by simple task execution dependencies can easily lead to critical tasks that depend on non-critical tasks, despite having priority access to resources, being unable to execute because the non-critical tasks they depend on have not yet been executed. This results in the unavailability of critical services corresponding to critical tasks and waste of resources due to the critical tasks continuously preempting them. To address this issue, this solution proposes a method to solve the above problems. Specifically, real-time task execution information also includes a task delay impact value. Based on real-time resource information, real-time task execution information, and the task priority score, the task priority score is adjusted to obtain a new task priority score. This includes: determining an increase in the task priority score based on the task delay impact value, where the task delay impact value represents the total impact of delayed task execution on tasks with execution priority scores higher than a preset high-priority threshold; and using the sum of the increase and the task priority score as the new task priority score. Therefore, by increasing the task priority score of each task through the task delay impact value, a stronger correlation is formed between the task priority score and the critical tasks that depend on it. This prevents critical tasks from being delayed because their dependent non-critical tasks have not yet been executed, thus ensuring that critical tasks can be executed in a timely and smooth manner.

[0092] Specifically, the task delay impact value is determined based on the critical tasks that depend on it. Tasks with a priority score higher than a preset high priority threshold are considered critical tasks. The task delay impact value can be obtained by weighted summation based on the importance level of the critical tasks that depend on it.

[0093] In some examples, the increase in task priority score is determined based on the impact value of task delay, including: using the impact value of task delay as the increase in task priority score; or using the product of the impact value of task delay and the increase ratio as the increase in task priority score. The increase ratio can be set according to actual needs and can be 0.1, 0.2, 0.3, 1, 2, or other values.

[0094] Optionally, as an example three, when cluster resources are scarce, to ensure the smooth execution of critical tasks, this solution proposes a method: real-time resource information includes the remaining resources of the Kubernetes cluster. Based on the real-time resource information, the real-time task execution information, and the task priority score, the task priority score is adjusted to obtain a new task priority score. This includes: if the remaining cluster resources are less than the cluster resource threshold, and the task priority score is lower than a preset low priority threshold, then the task priority score is set to 0. Therefore, when cluster resources are scarce, setting the task priority score to 0 prevents non-critical tasks from preempting resources, helping to ensure that the remaining cluster resources prioritize the smooth execution of critical tasks.

[0095] Specifically, the cluster resource threshold is used to determine whether cluster resources are strained. If the remaining resources of the cluster are lower than the cluster resource threshold, it indicates that the cluster resources are strained.

[0096] Specifically, tasks with a priority score below a preset low-priority threshold are classified as non-critical tasks. Non-critical tasks refer to tasks with low timeliness requirements and whose delayed execution does not affect critical services. For example, a background log analysis task.

[0097] When adjusting the task priority scores of each task, one or a combination of the adjustment methods proposed in Optional Exemplary 1, Optional Exemplary 2 and Optional Exemplary 3 can be used. Specifically, any two adjustment methods can be used in combination, for example, the two adjustment methods proposed in Optional Exemplary 1 and Optional Exemplary 2 can be used in combination; or three adjustment methods can be used in combination.

[0098] Step 1402: Schedule each task based on real-time resource information, real-time task execution information of each task, and new task priority scores of each task.

[0099] In some examples, such as Figure 2As shown, the scheduler can schedule tasks based on real-time resource information, real-time task execution information, and new task priority scores, thereby enabling the scheduling of tasks in the Kubernetes cluster.

[0100] In some examples, such as Figure 2 As shown, in step 1402 above, before scheduling each task based on real-time resource information, real-time task execution information of each task, and the new task priority score of each task, the method further includes: inputting the new task priority score of each task into the scheduling strategy integration module to obtain a task priority score in a standard format; using the task priority score in the standard format as the new task priority score; and sending the new task priority score to the scheduler. The standard format can be JSON format.

[0101] Among these features, new task priority scores can be passed to the scheduler through the Kubernetes extended scheduler interface.

[0102] Optionally, in step 1402 above, scheduling each task based on real-time resource information, real-time task execution information of each task, and new task priority scores of each task includes: sorting the task priority scores of each task in descending order to obtain a task priority order; and performing the following operations based on the task priority order: determining the node scheduling score of each node in the Kubernetes cluster based on the new task priority score of the task, the real-time task execution information of the task, and real-time resource information, selecting the node with the highest node scheduling score to schedule the task, and updating the real-time resource information.

[0103] In some examples, the method determines the node scheduling score of each node in the Kubernetes cluster based on the task's new task priority score, real-time task execution information, and real-time resource information. Before selecting the node with the highest scheduling score to schedule the task, the method further includes: determining the initial node whose remaining resources are greater than the task resource requirement information in the real-time task execution information from the real-time resource information; using the initial node as the node in the Kubernetes cluster to obtain the node scheduling score of each node in the Kubernetes cluster. Therefore, by excluding nodes that do not meet the task's resource requirements, the method avoids calculating the node scheduling score for nodes that do not meet the task's resource requirements, thus avoiding wasted computing resources. This improves the computational efficiency of obtaining the node scheduling score of each node.

[0104] In some examples, the node scheduling score of each node in the Kubernetes cluster is determined based on the new task priority score, real-time task execution information, and real-time resource information. This includes: obtaining the task weight corresponding to the task priority score for each node; obtaining the remaining resources of the node from the real-time task execution information; obtaining the task resource requirement information from the real-time task execution information; calculating the resource ratio based on the total resource requirement and the remaining resources of the node; determining the resource weight of the resource ratio; determining the task status value and task status weight based on the node usage status in the real-time task execution information; and determining the node scheduling score based on the task weight, task priority score, resource ratio, resource weight, task status value, and task status weight. The task status value describes the stability of task execution; the higher the task status value, the worse the stability of task execution.

[0105] Specifically, the task status value and task status weight are determined based on the node usage status in the real-time task execution information. The node scheduling score is then determined based on the task weight, task priority score, resource ratio, resource weight, task status value, and task status weight, as shown in the following formula:

[0106]

[0107] Among them, w1 is the resource weight, w2 is the task status weight, and w3 is the task weight. Each weight can be dynamically adjusted according to business needs and cluster characteristics. The ratio of NodeRemainingResources to TaskResourceRequest is the resource ratio. NodeStability is the task status value, and TsaskPriority is the task priority score.

[0108] Therefore, by prioritizing tasks with high priority scores and selecting nodes with high scheduling scores for scheduling, we can achieve a balance between resources and service quality. This is achieved by reserving nodes with low resource utilization and high stability for tasks with high priority scores, and allocating nodes with high resource utilization to tasks with low priority scores.

[0109] In some examples, the method further includes, if a new task joins the Kubernetes cluster, obtaining the task priority score of the new task and re-executing the scheduling of tasks based on real-time resource information, real-time task execution information of each task, and the task priority score of each task. Specifically, when Kubernetes cluster resources are insufficient, the new task preempts resources from tasks with lower task priority scores. When Kubernetes cluster resources are sufficient, the new task is scheduled according to its task priority score.

[0110] In summary, this solution obtains historical task information for at least one task in the Kubernetes cluster. This historical task information includes the historical execution information of each task, the business importance level of the business to which each task belongs, and the business timeliness information. The historical task information of at least one task is input into a task priority prediction model to obtain the task priority score for each task. Real-time resource information of the Kubernetes cluster and real-time task execution information of each task are obtained. Tasks are scheduled based on this real-time resource information, real-time task execution information, and task priority scores. Based on the business importance level, business timeliness information, and task execution information of each task, accurate task priority scores that indicate the differences in importance among tasks at the business level are obtained. Scheduling tasks based on their priority scores, real-time resource information, and real-time task execution information enables a fine-grained and dynamic resource allocation strategy. This not only allows for flexible adjustment of resource allocation based on the current real-time resources of the Kubernetes cluster and the task execution information of each task, but also supports dynamic resource preemption based on task priority scores. This avoids significant deviations between allocated resources and the actual resources required by tasks, leading to resource shortages or waste, and improves the scheduling flexibility and resource utilization efficiency of the Kubernetes cluster. It also ensures timely response and stable operation of high-priority scoring tasks.

[0111] Specifically, a task priority prediction model is obtained by training deep learning models such as RNN, LSTM, or Transformer. This model is then used to predict the historical task information of each task, resulting in a task priority score. Unlike traditional methods that rely on static rules or simple algorithms to obtain task priorities, this approach fully leverages the powerful feature learning capabilities of neural network models to automatically uncover potential patterns in task data, achieving accurate prediction of task priority scores and filling the technological gap in intelligent prediction during task scheduling in Kubernetes clusters. Tasks are scheduled by combining real-time resource information, real-time task execution information, and task priority scores. A dynamic priority adjustment mechanism is implemented, which can adaptively adjust scheduling based on dynamic changes in tasks and the cluster. This allows for rapid adaptive adjustments to various unforeseen circumstances, improving scheduling efficiency and resource utilization. For example, in the event of resource shortages or task delays, the task priorities and scheduling strategies of each task can be quickly and accurately adjusted, ensuring that the scheduling strategy adapts to various unforeseen situations in real time.

[0112] This solution forms a complete AI-based Kubernetes task priority scheduling system through data collection and preprocessing modules, a prediction model building module, a priority dynamic adjustment module, and a scheduler. Through the collaborative work of these modules, it achieves end-to-end intelligent scheduling, from data acquisition and priority prediction to dynamic adjustment and intelligent scheduling. This provides a brand-new systematic scheduling solution for Kubernetes task scheduling, significantly improving scheduling flexibility and resource utilization efficiency.

[0113] Example 2:

[0114] Another embodiment of this application relates to an AI-based Kubernetes task priority scheduling device. The implementation details of this AI-based Kubernetes task priority scheduling device are described below. The following details are provided for ease of understanding and are not essential for implementing this solution. A schematic diagram of the AI-based Kubernetes task priority scheduling device 500 in this embodiment can be seen as follows: Figure 5 As shown, it includes an acquisition unit 51, an input unit 52, and a scheduling unit 53.

[0115] The acquisition unit 51 is used to acquire historical task information of at least one task in the Kubernetes cluster. The historical task information includes the historical task execution information of each task, the business importance level of the business to which each task belongs, and the business timeliness information.

[0116] Input unit 52 is used to input historical task information of at least one task into the task priority prediction model to obtain the task priority score of each task.

[0117] The acquisition unit 51 is also used to acquire real-time resource information of the Kubernetes cluster and real-time task execution information of each task.

[0118] The scheduling unit 53 is used to schedule each task based on real-time resource information, real-time task execution information of each task, and task priority score of each task.

[0119] In some examples, when scheduling unit 53 is used to schedule tasks based on real-time resource information, real-time task execution information of each task, and task priority score of each task, it specifically performs the following: for each task, it adjusts the task priority score based on real-time resource information, real-time task execution information of each task, and task priority score of each task to obtain a new task priority score; and schedules each task based on real-time resource information, real-time task execution information of each task, and new task priority score of each task.

[0120] In some examples, real-time task execution information includes task resource requirement information and task running nodes. Real-time resource information includes the remaining resources of each node. When scheduling unit 53 adjusts the task priority score based on real-time resource information, real-time task execution information, and task priority score to obtain a new task priority score, it specifically performs the following: For critical tasks whose task priority score is higher than a preset priority score, it obtains the remaining resources of the target node of the critical task's running node from the remaining resources of each node; if the task resource requirement information is higher than a preset task requirement threshold and higher than the remaining resources of the target node, it determines the increase in the task priority score of the critical task based on the difference between the task resource requirement information and the remaining resources of the target node; the sum of the increase and the task priority score is taken as the new task priority score of the critical task; wherein, the preset requirement threshold of the task corresponds one-to-one with the critical task.

[0121] In some examples, the real-time task execution information also includes a task delay impact value. When the scheduling unit 53 adjusts the task priority score based on real-time resource information, the real-time task execution information of the task, and the task priority score of the task to obtain a new task priority score, it specifically performs the following: determining the increase value of the task priority score based on the task delay impact value, wherein the task delay impact value is used to represent the total impact value of the delayed execution of the task on the task whose execution priority score is higher than a preset high priority threshold; and using the sum of the increase value and the task priority score as the new task priority score.

[0122] In some examples, real-time resource information includes the remaining resources of the Kubernetes cluster. When scheduling unit 53 adjusts the task priority score based on real-time resource information, real-time task execution information of the task, and task priority score of the task to obtain a new task priority score, it specifically does the following: if the remaining cluster resources are less than the cluster resource threshold and the task priority score is lower than the preset low priority threshold, then the task priority score is set to 0.

[0123] In some examples, when scheduling unit 53 schedules tasks based on real-time resource information, real-time task execution information of each task, and new task priority scores of each task, it specifically performs the following: sorting the task priority scores of each task in descending order to obtain the task priority order; executing tasks sequentially based on the task priority order; determining the node scheduling score of each node in the Kubernetes cluster based on the new task priority score of the task, the real-time task execution information of the task, and the real-time resource information; selecting the node with the highest node scheduling score to schedule the task; and updating the real-time resource information.

[0124] In some examples, the acquisition unit 51 is also used to: acquire sample task information of at least one task in the Kubernetes cluster, the sample task information including sample task execution information of each task, the sample business importance level of the business to which each task belongs, and the sample business timeliness information; perform data preprocessing on the sample task information of at least one task to obtain target sample information of each task, the data preprocessing including data cleaning and data standardization; perform feature engineering on the target sample information of each task to obtain feature information and label information of each task; and train an initial task priority prediction model based on the target sample information, feature information, and label information of each task to obtain a task priority prediction model that meets preset requirements.

[0125] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units are absent in this embodiment.

[0126] Example 3:

[0127] Another embodiment of this application relates to an electronic device, such as... Figure 6As shown, it includes: at least one processor 901; and a memory 902 communicatively connected to at least one processor 901; wherein the memory 902 stores instructions executable by at least one processor 901, the instructions being executed by at least one processor 901 to enable at least one processor 901 to execute the AI-based Kubernetes task priority scheduling method in the above embodiments.

[0128] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.

[0129] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.

[0130] Example 4:

[0131] Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.

[0132] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0133] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of this application.

Claims

1. An AI-based Kubernetes task priority scheduling method, characterized in that, include: Obtain historical task information for at least one task in the Kubernetes cluster. The historical task information includes historical task execution information for each task, as well as the business importance level and business timeliness information of the business to which each task belongs. The historical task information of at least one of the tasks is input into the task priority prediction model to obtain the task priority score of each task; Obtain real-time resource information of the Kubernetes cluster and real-time task execution information of each task; The tasks are scheduled based on the real-time resource information, the real-time task execution information of each task, and the task priority score of each task.

2. The AI-based Kubernetes task priority scheduling method according to claim 1, characterized in that, Scheduling of each task based on the real-time resource information, the real-time task execution information of each task, and the task priority score of each task includes: For each task, the task priority score is adjusted based on the real-time resource information, the real-time task execution information of the task, and the task priority score of the task to obtain a new task priority score for the task. The tasks are scheduled based on the real-time resource information, the real-time task execution information of each task, and the new task priority score of each task.

3. The AI-based Kubernetes task priority scheduling method according to claim 2, characterized in that, The real-time task execution information includes task resource requirement information and task running nodes. The real-time resource information includes the remaining resources of each node. The adjustment of the task priority score based on the real-time resource information, the task's real-time task execution information, and the task's task priority score to obtain a new task priority score includes: For critical tasks whose priority scores are higher than preset priority scores, the remaining target node resources of the task running node of the critical task are obtained from the remaining node resources of each node. If the task resource requirement information is higher than the preset task requirement threshold and higher than the remaining resources of the target node, then the increase in the task priority score of the key task is determined based on the difference between the task resource requirement information and the remaining resources of the target node. The sum of the increased value and the task priority score is taken as the new task priority score for the critical task; The preset requirement thresholds for the tasks correspond one-to-one with the key tasks.

4. The AI-based Kubernetes task priority scheduling method according to claim 2, characterized in that, The real-time task execution information also includes a task latency impact value. The adjustment of the task priority score based on the real-time resource information, the task's real-time task execution information, and the task's task priority score to obtain a new task priority score includes: The increase in the task priority score is determined based on the task delay impact value, wherein the task delay impact value is used to represent the total impact of the delayed execution of the task on tasks whose execution task priority score is higher than a preset high priority threshold; The sum of the increased value and the task priority score is taken as the new task priority score for the task.

5. The AI-based Kubernetes task priority scheduling method according to claim 2, characterized in that, The real-time resource information includes the remaining resources of the Kubernetes cluster. The adjustment of the task priority score based on the real-time resource information, the real-time task execution information of the task, and the task priority score of the task to obtain a new task priority score includes: If the remaining resources of the cluster are less than the cluster resource threshold, and the task priority score is lower than the preset low priority threshold, then the task priority score is set to 0.

6. The AI-based Kubernetes task priority scheduling method according to claim 2, characterized in that, The scheduling of each task based on the real-time resource information, the real-time task execution information of each task, and the new task priority score of each task includes: The task priority scores of each task are sorted from highest to lowest to obtain the task priority order. Based on the task priority order, the following steps are performed sequentially: Based on the new task priority score of the task, the real-time task execution information of the task, and the real-time resource information, the node scheduling score of each node in the Kubernetes cluster is determined, the node with the highest node scheduling score is selected to schedule the task, and the real-time resource information is updated.

7. The AI-based Kubernetes task priority scheduling method according to claim 1, characterized in that, The method further includes: Obtain sample task information for at least one task in a Kubernetes cluster. The sample task information includes sample task execution information for each task, the sample business importance level of the business to which each task belongs, and the sample business timeliness information. Data preprocessing is performed on sample task information of at least one of the tasks to obtain target sample information for each task. The data preprocessing includes data cleaning and data standardization. Feature engineering is performed on the target sample information of each task to obtain the feature information and label information of each task. The initial task priority prediction model is trained based on the target sample information, feature information, and label information of each task to obtain a task priority prediction model that meets the preset requirements.

8. An AI-based Kubernetes task priority scheduling device, characterized in that, include: The acquisition unit is used to acquire historical task information of at least one task in the Kubernetes cluster. The historical task information includes historical task execution information of each task, business importance level and business timeliness information of the business to which each task belongs. An input unit is used to input the historical task information of at least one of the tasks into a task priority prediction model to obtain a task priority score for each task. The acquisition unit is also used to acquire real-time resource information of the Kubernetes cluster and real-time task execution information of each task; The scheduling unit is used to schedule each task based on the real-time resource information, the real-time task execution information of each task, and the task priority score of each task.

9. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the AI-based Kubernetes task priority scheduling method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the AI-based Kubernetes task priority scheduling method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Flash lamp control method and device, medium and program product

    CN121586131A

  • Flash control method, device, medium and program product

    CN121586131B