Multi-target collaborative scheduling optimization method and system for AI training tasks on cloud

By analyzing historical load data and dynamically adjusting resource allocation strategies, the problem of inaccurate resource demand prediction for cloud servers in multi-tenant scenarios has been solved, achieving efficient, energy-saving, and fair resource allocation and optimizing multi-objective scheduling.

CN121833295AInactive Publication Date: 2026-04-10JIANGSU AOGONG INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-16
Publication Date
2026-04-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing cloud servers struggle to accurately predict each tenant's resource needs in multi-tenant scenarios based on historical data, leading to extended task waiting times, unreasonable resource allocation, and an inability to balance task completion time, energy costs, and resource fairness.

Method used

By analyzing historical load data, resource demand is predicted, and a method of fusion of safety control coefficients and historical data is used to dynamically adjust resource allocation strategies. In combination with the current environmental conditions, multiple objectives are prioritized to achieve efficient, energy-saving and fair allocation of resources.

Benefits of technology

It improved the accuracy of resource allocation, reduced operating costs, ensured resource supply during peak business periods, avoided resource waste and shortages, and optimized multi-objective resource scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833295A_ABST
    Figure CN121833295A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-target collaborative scheduling optimization method and system for AI training tasks on cloud, and the method comprises the steps: analyzing historical load data, and predicting the resource demand of each tenant at a next detection time point; judging whether the deviation degree between the predicted resource demand quantity of the same tenant and the actual resource demand quantity is within a set safety regulation coefficient range or not, and if not, extracting historical load data of the tenant under the same event; calculating the predicted resource demand quantity at the next detection time point after fusion; and based on the current environment state and the predicted resource demand quantity of each tenant, determining the priority of multiple targets, and carrying out dynamic allocation regulation and control on resource data of the cloud platform so as to balance multi-tenant multi-target resource allocation. According to the invention, real-time interaction is carried out through the predicted resource demand and the environment state, so that resource distribution is continuously adjusted, multi-target and multi-user resource distribution requirements are met, and efficient, energy-saving and fair resource scheduling is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of cloud service technology and relates to a multi-objective collaborative scheduling optimization method and system for AI training tasks in the cloud. Background Technology

[0002] With the rapid development and widespread adoption of cloud computing technology, cloud servers have become a core infrastructure supporting various enterprise and personal applications. In a multi-tenant cloud environment, the performance of resource scheduling strategies directly determines the operational efficiency, cost control, and service quality of the data center. However, current cloud servers, when facing complex and ever-changing multi-tenant scenarios, cannot predict the resource demands of each tenant based on historical data. This results in an inability to rationally allocate resources based on each tenant's resource demands and the current environmental conditions, considering factors such as task completion time, energy costs, and resource fairness.

[0003] In multi-tenant scenarios, tasks submitted by different tenants often exhibit significant heterogeneity, with varying resource requirements (such as computational, memory, or communication needs). Existing traditional resource scheduling algorithms lack a global awareness of each tenant's resource requirements, often making it difficult to accurately predict these requirements. They typically employ a passive, reactive scheduling strategy, which not only causes tasks to queue for extended periods after submission, significantly prolonging overall task completion time and reducing cluster throughput, but also fails to effectively balance resource scheduling conflicts among multiple tenants. Summary of the Invention

[0004] The purpose of this invention is to provide a multi-objective collaborative scheduling optimization method and system for cloud-based AI training tasks, which solves the above-mentioned technical problems.

[0005] The objective of this invention can be achieved through the following technical solutions: A multi-objective collaborative scheduling optimization method for cloud-based AI training tasks includes: Extract historical load data for each tenant within a detection cycle prior to the current detection time, analyze the historical load data within the detection cycle, and predict the resource requirements of each tenant at the next detection time. Obtain the actual resource requirements of each tenant, and determine whether the deviation between the predicted resource requirements and the actual resource requirements of the same tenant is within the set safety control coefficient range. The set safety control coefficient range is [0, u], where u is greater than 0, so that the predicted resource requirements are greater than the actual resource requirements. If a resource demand is outside the set safety control coefficient range, then extract the historical load data corresponding to that tenant under the same event. The resource demand predicted by the tenant at the current detection time point within the current detection cycle and the historical load data corresponding to the same event are merged to obtain the resource demand predicted at the next detection time point after fusion, so that the fused resource demand is close to the actual resource demand of the tenant. The system obtains the current environmental status of the cloud platform at the current detection time. Based on the current environmental status and the predicted resource requirements of each tenant, it determines the priority of multiple objectives and dynamically allocates and regulates the cloud platform's resource data according to the priority of multiple objectives to balance the resource allocation of multiple tenants and multiple objectives.

[0006] Furthermore, historical load data for each detection time point within a detection cycle prior to the previous detection time point is extracted. By analyzing the historical load data of each tenant, the burst coefficient of resource demand for each tenant at two adjacent detection time points is determined. Based on the set security control coefficient, the resource demand required by each tenant at the next detection time point is predicted. The resource demand includes CPU demand, memory demand, and bandwidth demand.

[0007] Furthermore, based on historical load data, the method for determining the CPU requirements for each tenant at the next detection time point includes: Extract the CPU utilization rate of each tenant at each detection time point within the detection period; Based on the CPU utilization rate corresponding to each detection time point, the CPU utilization rate threshold corresponding to the time ratio threshold r is selected, and the time ratio threshold r takes the value range [0,1]. Based on the CPU utilization corresponding to two adjacent detection time points, analyze the burst coefficient corresponding to the sudden change in CPU utilization, and use the set security control coefficient to analyze the secure CPU utilization of each tenant at the next detection time point. Based on the secure CPU utilization of each tenant, predict the CPU demand of each tenant at the next detection time point.

[0008] Furthermore, based on the predicted security CPU utilization of each tenant, the CPU core count and CPU utilization at the previous detection time point are used to predict the CPU demand at the next detection time point, so as to obtain the CPU demand of each tenant at the next detection time point.

[0009] Furthermore, the method for calculating the burst coefficient corresponding to the sudden change in CPU utilization includes: obtaining the maximum CPU utilization of the same tenant in historical load data within the current detection period; taking the ratio between the maximum CPU utilization of the tenant and the CPU utilization threshold, and subtracting the value 1 to obtain the burst coefficient of CPU utilization, wherein the CPU utilization threshold is a set CPU utilization to control the utilization of the allocated CPU.

[0010] Furthermore, the actual resource demand of each tenant at each detection time point is compared with the predicted resource demand at that detection time point to determine the degree of deviation between the predicted and actual resource demand of the same tenant. The determination of the degree of deviation is based on the ratio of the difference between the predicted and actual resource demand. The ratio of the difference is equal to the difference between the predicted and actual demand of the same tenant under the same resource conditions, and the ratio of the difference to the actual demand.

[0011] Furthermore, if there is a tenant whose predicted resource demand at the next detection time point is not within the set safety control coefficient range, the predicted resource demand will be dynamically adjusted.

[0012] Furthermore, the methods used to dynamically adjust the predicted resource demand include: Extract the actual resource demand at the testing time point corresponding to the current testing time point of the current testing period within the previous testing period, and analyze the year-on-year growth rate coefficient; Based on the year-on-year growth rate and the resource demand G at the next testing time point in the previous testing cycle, we analyze the resource demand at the next testing time point in the current testing cycle, so as to predict the resource demand at the next testing time point in the current testing cycle based on the resource demand in the previous testing cycle. Based on the resource demand predicted for the next detection time point in the current detection period from the previous detection period, and the resource demand predicted for the next detection time point in the current detection period, a weighted fusion is performed to obtain the predicted resource demand for the next detection time point in the current detection period.

[0013] Furthermore, based on the determined priorities of the multiple objectives, the initial weight coefficients corresponding to each objective are selected, and the weight coefficients of each objective are adjusted using the comprehensive allocation evaluation coefficients determined by the multi-objective evaluation function to meet the resource allocation requirements of multi-tenant multi-objectives.

[0014] Based on the same inventive concept, this invention discloses a multi-objective collaborative scheduling optimization system for cloud AI training tasks, applied to any of the aforementioned multi-objective collaborative scheduling optimization methods for cloud AI training tasks, comprising: The resource demand forecasting module is used to extract the historical load data of each tenant within a detection cycle before the current detection time, analyze the historical load data of the detection cycle, and predict the resource demand of each tenant at the next detection time. The prediction and judgment module is used to obtain the actual resource demand of each tenant and determine whether the deviation between the predicted resource demand and the actual resource demand of the same tenant is within the set safety control coefficient range. The set safety control coefficient range is [0, u], where u is greater than 0, so that the predicted resource demand is greater than the actual resource demand. If a resource demand is outside the set safety control coefficient range, then extract the historical load data corresponding to that tenant under the same event. The resource data fusion module is used to merge the resource demand predicted by the tenant at the current detection time point within the current detection cycle with the historical load data corresponding to the same event, so as to obtain the resource demand predicted at the next detection time point after fusion, so that the fused resource demand is close to the actual resource demand of the tenant. The multi-objective control module is used to obtain the environmental status of the cloud platform at the current detection time. Based on the current environmental status and the predicted resource demand of each tenant, it determines the priority of multiple objectives and dynamically allocates and controls the resource data of the cloud platform according to the priority of multiple objectives to balance the resource allocation of multiple tenants and multiple objectives.

[0015] The beneficial effects of this invention are: This invention analyzes the historical load data of each tenant within a detection cycle prior to the current detection time point to filter the CPU utilization threshold, memory usage threshold, and I / O throughput threshold corresponding to the time ratio threshold. Combined with the burst coefficient corresponding to each resource parameter within the current detection cycle, it predicts the resource demand required by each tenant at the next detection time point, so that the predicted resource demand is greater than the actual resource demand, thus satisfying the resource needs of each tenant.

[0016] This invention analyzes the deviation between the predicted resource demand of each tenant and the actual resource demand to determine whether the deviation is within the set safety control coefficient range. If not, it filters out the same events and uses the actual resource demand at the same detection time point under the same event to predict the resource demand of each tenant at the current detection time point. By using the same events in the past data, the accuracy of the predicted resource demand of each tenant is improved.

[0017] This invention uses a weighted fusion of the resource demand of each tenant predicted by the same event and the resource data volume of each tenant predicted by the historical load data of the current detection period. It can adjust the resource demand of the next detection time point predicted by the current detection time point in the current detection period by using the historical data of the previous detection period of the same event. This effectively solves the lag problem in prediction by a single data source. While ensuring the supply of resources during peak business periods, it suppresses the inflated allocation of resources caused by occasional fluctuations, reduces operating costs, and achieves highly accurate calculation of resource demand before cloud resource allocation.

[0018] The steps employed in this invention can calculate the resource requirements of each tenant at the next detection time point, thereby enabling real-time prediction of the resource requirements of each tenant at each detection time point. Furthermore, based on the environmental state at the current detection time point, the priority of the multi-objectives corresponding to each tenant during the resource allocation process is determined, and resource data is dynamically allocated according to the priority of the multi-objectives, so that the allocation of the predicted resource requirements of each tenant meets the current environmental state, thereby balancing the resource allocation needs of multiple tenants and multiple objectives.

[0019] This invention continuously adjusts resource allocation by interacting in real time with predicted resource demand and environmental conditions to meet the resource allocation needs of multiple objectives and users, thereby achieving efficient, energy-saving, and fair resource collaborative scheduling optimization. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart of the multi-objective cooperative scheduling optimization method in this invention; Figure 2 This is a flowchart of the method for dynamically adjusting resource demand in this invention; Figure 3 This is a schematic diagram of the multi-objective collaborative scheduling optimization system in this invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] Current cloud servers struggle to balance task completion time, energy costs, and resource fairness in multi-tenant scenarios, thus failing to provide efficient, energy-saving, and fair intelligent scheduling strategies.

[0024] In a multi-tenant environment, tasks submitted by different tenants often have different resource requirements. Existing traditional scheduling algorithms are often too simple and lack the ability to globally perceive task load. Especially for complex AI training tasks, the execution time depends not only on the amount of allocated computing resources, but also on the communication and data waiting time between nodes. Existing technologies often struggle to accurately predict the resource requirements and load peaks and troughs of tasks, thus failing to accurately predict the resource requirements of each tenant. Consequently, they cannot balance the multi-tenant, multi-objective problem based on the predicted resource requirements of each tenant.

[0025] To avoid sudden task peaks, resources are over-configured, causing all servers to operate at high power. Conversely, during task troughs, a lack of accurate prediction of each tenant's resource needs leads to servers entering low-power sleep mode, increasing task completion time. Using a single objective for control results in unreasonable resource allocation, making it difficult to balance resource allocation across multiple objectives and hindering collaborative optimization. To address these issues, such as... Figure 1 As shown, this application discloses a multi-objective cooperative scheduling optimization method for cloud-based AI training tasks, including: Step W1: Extract the historical load data of each tenant within a detection cycle before the current detection time point, analyze the historical load data of the detection cycle, and predict the resource requirements of each tenant at the next detection time point. The historical load data includes CPU utilization, memory usage, average queuing time of a single task, and average execution time of a single task during the detection period.

[0026] Historical load data for each detection time point within a detection cycle prior to the previous detection time point is extracted. By analyzing the historical load data of each tenant, the burst coefficient of resource demand for each tenant at two adjacent detection time points is determined. Based on the set security control coefficient, the resource demand required by each tenant at the next detection time point is predicted. The resource demand includes CPU demand, memory demand, and bandwidth demand.

[0027] Specifically, the CPU utilization rate of each tenant in the historical load data within the detection period is obtained. By analyzing the CPU utilization rate in the historical load data of each tenant, the CPU requirement of each tenant at the next detection time point is determined. The methods used include: Step 11: Extract the CPU utilization rate of each tenant at each detection time point; The time interval between two adjacent detection time points is the same, which constitutes a detection time period. Using the same detection time period is to obtain the CPU utilization corresponding to each detection time point and reduce the impact of changes in sampling interval on the accuracy of CPU utilization. The detection time period can be 30 minutes, 1 hour, or 1 day, etc., and the length of the detection time period is less than the length of a detection cycle.

[0028] The duration within a detection cycle is greater than the duration corresponding to two adjacent detection time points, thereby ensuring that a data model of CPU utilization is fitted based on the CPU utilization corresponding to each detection time point, and obtaining the trend of CPU utilization of each tenant over time.

[0029] Step 12: Based on the CPU utilization corresponding to each detection time point, filter the CPU utilization threshold corresponding to the time ratio threshold r. The time ratio threshold is a ratio coefficient with a value range of [0,1]. The value of the time ratio threshold r is determined by the R&D department. The selection of the time ratio threshold r directly affects the selection of the CPU utilization corresponding to the time ratio threshold. Here, the time ratio threshold is 0.95, but other values ​​such as 0.92 can also be selected.

[0030] When the time ratio threshold r is set to 0.95, it means that the CPU utilization rate corresponding to 95% of the detection time points within a detection cycle is lower than a certain CPU utilization rate, i.e., t. low / T=r, where T represents the total duration of one detection cycle, t low This represents the duration when CPU utilization is less than the CPU utilization threshold. The CPU utilization values ​​at each detection time point are used for judgment. Based on the above time ratio threshold r, the CPU utilization corresponding to the time ratio threshold r can be determined, and the CPU utilization corresponding to the time ratio threshold r is used as the CPU utilization threshold.

[0031] Step 13: Based on the CPU utilization corresponding to two adjacent detection time points, analyze the burst coefficient corresponding to the sudden change in CPU utilization, and use the set security control coefficient to analyze the secure CPU utilization of each tenant at the next detection time point. The secure CPU utilization rate for each tenant is calculated as follows: CPU utilization threshold * (1 + burst coefficient) * security control coefficient. The burst coefficient is related to historical load data, and the security control coefficient ranges from 1.1 to 1.3. By setting the security control coefficient, the future CPU utilization rate is controlled with a safety margin based on historical data to cope with sudden demand.

[0032] The burst coefficient of CPU utilization can be obtained from the CPU utilization in historical load data within a detection period. The detection period can be 1 day, 1 week, 1 month, 1 quarter, 1 year, or special holidays, etc. In this embodiment, a method for calculating the burst coefficient is disclosed, including: obtaining the maximum value of CPU utilization of the same tenant in historical load data within the current detection period; taking the ratio between the maximum value of the tenant's CPU utilization and the CPU utilization threshold, and subtracting the value 1 to obtain the burst coefficient of CPU utilization. The CPU utilization threshold is a set CPU utilization to control the utilization rate of the allocated CPU.

[0033] The burst coefficient determined by the ratio between the maximum CPU utilization value and the CPU utilization threshold in the above historical load data is the maximum burst coefficient, providing reliable numerical support for the reserved CPU utilization data.

[0034] Step 14: Based on the security CPU utilization of each tenant, predict the CPU demand of each tenant at the next detection time point.

[0035] Based on the predicted security CPU utilization of each tenant, the CPU core count and CPU utilization at the previous detection time point are used to predict the CPU demand at the next detection time point, so as to obtain the CPU demand of each tenant at the next detection time point.

[0036] Specifically, the CPU requirement at the next detection time point = the number of CPU cores at the previous detection time point * the CPU utilization rate at the previous detection time point / the safe CPU utilization rate at the next detection time point.

[0037] By analyzing the CPU utilization in historical load data within a detection period, the safe CPU utilization of each tenant is dynamically analyzed. Based on the safe CPU utilization and the number of CPU cores at the previous detection time point, the CPU demand required by each tenant at the next detection time point is predicted. This allows for the prediction of the CPU demand required at the next detection time point based on historical data. Similarly, based on the memory usage and I / O throughput in the historical load data within the previous detection period, the memory demand and I / O throughput of each tenant are predicted.

[0038] Specifically, the method for predicting the memory requirement corresponding to the next detection time point includes: obtaining the memory usage of each tenant at each detection time point in the previous detection cycle, and filtering the memory usage threshold corresponding to the time ratio threshold. The time ratio threshold corresponding to the memory usage threshold can be the same as or different from the time ratio threshold corresponding to the CPU utilization threshold. Based on the memory usage corresponding to two adjacent detection time points, analyze the burst coefficient corresponding to the sudden change in memory usage, and use the set safety control coefficient to calculate the memory demand of each tenant at the next detection time point. Formula: Memory demand of each tenant at the next detection time point = memory usage threshold * (1 + burst coefficient) * security control coefficient. The burst coefficient is related to the memory usage in historical load data. The security control coefficient is between 1.1 and 1.3. The security control coefficient is used to make a safety margin for future memory usage based on historical data in order to cope with sudden memory demand.

[0039] Similarly, the mutation coefficient corresponding to a sudden increase in memory usage is obtained by analyzing the memory usage of the same tenant within the current detection period, calculating the rate of change corresponding to the change in memory usage between two adjacent detection time points, screening the maximum rate of change in memory usage, and using the maximum rate of change in memory usage as the mutation coefficient of memory usage.

[0040] Based on the I / O throughput corresponding to each detection time point in the historical load data, the I / O throughput corresponding to the time ratio threshold is filtered out, and the I / O throughput corresponding to the time ratio threshold is selected as the I / O throughput threshold. Based on the I / O throughput threshold corresponding to the time ratio threshold, combined with the security control coefficient, the bandwidth demand of each tenant at the next detection time point is predicted. The bandwidth requirement is calculated as I / O throughput threshold * security control coefficient, where the security control coefficient ranges from 1.1 to 1.3.

[0041] In this embodiment, the formulas for predicting secure CPU utilization, memory requirements, and bandwidth requirements all include a security control coefficient. The values ​​of the security control coefficients can be the same or different. The specific selection of the security control coefficient depends on the difference between the actual resource requirements and the predicted resource requirements, so that the predicted resource requirements are greater than the actual resource requirements. Thus, through the security control coefficient, the secure CPU utilization, memory requirements, and bandwidth requirements can be predicted based on the CPU utilization threshold, memory usage threshold, and I / O throughput threshold.

[0042] This embodiment also discloses a method for predicting required resource needs based on historical load data. The method extracts the CPU utilization, memory usage, and I / O throughput corresponding to each detection time point within the detection period, and fits the CPU utilization, memory usage, and I / O throughput corresponding to each time point to construct a historical data model. The historical data model includes a CPU utilization model, a memory usage model, and an I / O throughput model.

[0043] Historical data model: G i (t) = a i t 5 +b i t 4 +c i t 3 +d i t 2 +e i t+f i G i (t) represents the CPU utilization model, memory usage model, and I / O throughput model of the i-th resource, respectively, i=1,2,3, corresponding to CPU utilization, memory usage, and I / O throughput, and t represents the t-th detection time point. G i (t) coefficient a i b i c i d i e i and f i For the i-th resource, the coefficients of the corresponding historical data model are solved using the least squares method to obtain 'a' for each historical data model. i b i c i d i e i and f i The specific value of the coefficient.

[0044] Based on the historical data models of each resource, the CPU utilization, memory usage, and I / O throughput corresponding to the next detection time point can be predicted. Then, according to the above calculation formula, they can be converted into the required resource demand, that is, the CPU demand, memory demand, and bandwidth demand are obtained.

[0045] Step W2: Obtain the actual resource requirements of each tenant, and determine whether the deviation between the predicted resource requirements and the actual resource requirements of the same tenant is within the set safety control coefficient range. The set safety control coefficient range is [0, u], where u is greater than 0, so that the predicted resource requirements of each tenant are greater than the actual resource requirements, thereby meeting the usage needs of each tenant. By comparing the actual resource demand of each tenant at each detection time point with the predicted resource demand at that detection time point, the degree of deviation between the predicted resource demand and the actual resource demand of the same tenant can be determined.

[0046] The basis for judging the degree of deviation is the ratio of the difference between the tenant's predicted resource demand and the actual resource demand at a certain detection time point. The difference ratio is equal to the difference between the predicted demand and the actual demand of the same tenant under the same resource. The ratio of the difference to the actual demand reflects whether the degree of deviation between the predicted resource demand and the actual resource demand is within the set safety control coefficient range.

[0047] In this embodiment, the upper limit value u of the set safety control coefficient range can be determined based on experience, and is generally preferred to be 0.15.

[0048] In this embodiment, the difference ratio between the actual resource demand and the predicted resource demand at the previous detection time point can be used to reflect the degree of deviation between the predicted and actual resource demand of the same tenant. Alternatively, the difference ratio between the actual and predicted resource demand at the previous n detection time points can be used and averaged to obtain the average difference ratio. The average difference ratio is used to reflect the degree of deviation between the predicted and actual resource demand.

[0049] The above steps, by using historical load data, allow for advance prediction of the resource requirements needed at each subsequent detection time point, facilitating the advance reservation of resources. They also enable dynamic adjustment of the predicted resource requirements based on a comparison between the actual resource requirements of each tenant and the predicted resource requirements, thus avoiding insufficient or excessive predicted resource requirements that could affect the resource allocation of each tenant.

[0050] Step W3: If a resource demand is outside the set safety control coefficient range, extract the historical load data corresponding to that tenant under the same event. The same event is divided according to the tenant's resource demand, period and / or date. In this embodiment, step W1 obtains the predicted resource requirements (CPU requirements, memory requirements, and bandwidth requirements) of each tenant at the next detection time point. It then determines whether the deviation between the predicted resource requirements and the actual resource requirements of each tenant is within the set safety control coefficient range. If they are not within the set safety control coefficient range, the predicted resource requirements of each tenant are dynamically adjusted according to the deviation between the predicted and actual resource requirements to balance the difference between the predicted and actual resource requirements. This ensures that at the next detection time point, the predicted resource requirements of each tenant meet the actual resource requirements, avoiding excessive predicted resource requirements that would lead to resource waste, and also avoiding insufficient resource reservations caused by the predicted resource requirements failing to meet the actual resource requirements.

[0051] If, among the deviations between the predicted and actual resource demand of the same tenant, one deviation is outside the set safety control coefficient range, the same event in the current detection cycle is obtained, where the resource demand is similar, so that data prediction can be performed based on historical data from the previous detection cycle of the same event.

[0052] Based on the predicted resource demand of each tenant, select the historical load data corresponding to the same event that is closest to the current detection cycle to exclude large differences between the same events due to the long time period.

[0053] Step W4: The resource demand predicted by the tenant at the current detection time point within the current detection cycle and the historical load data corresponding to the same event are fused together to obtain the resource demand predicted at the next detection time point after fusion. This makes the fused resource demand closer to the actual resource demand of the tenant, reducing the deviation between the predicted resource demand and the actual resource demand caused by only using the historical load data of the previous detection cycle within the current detection cycle, which does not meet the set safety control coefficient range. If a tenant's predicted resource demand at the next detection time point differs from the tenant's actual resource demand by a factor that is outside the set safety control coefficient range, then... Figure 2 As shown, the predicted resource demand is dynamically adjusted, and the adjustment method includes: Step 1: Extract the actual resource demand at the detection time point corresponding to the current detection time point of the current detection period within the previous detection period, and analyze the year-on-year growth rate coefficient; formula: ,in, This represents the year-on-year growth rate coefficient of the k-th resource at the t-th detection time point. This represents the actual resource demand for the k-th resource at the t-th detection time point within the previous detection period. This represents the actual resource requirement of the k-th resource at the t-th detection time point within the current detection period.

[0054] In this embodiment, the average of the year-on-year growth coefficients corresponding to the corresponding detection time points in the previous detection cycle and the current detection cycle can also be used for calculation.

[0055] By using historical data from the current testing time point within the current testing cycle and comparing it with historical data from the same testing time point in the previous testing cycle, the growth rate of resource demand at the current testing time point within the current testing cycle compared to the same testing time point in the previous testing cycle can be displayed. This allows for the accurate calculation of resource demand at the next testing time point based on the year-on-year growth rate coefficient at the current testing time point.

[0056] Step 2: Based on the year-on-year growth rate coefficient and the resource demand G at the next testing time point in the previous testing cycle, analyze the resource demand at the next testing time point in the current testing cycle, so as to predict the resource demand at the next testing time point in the current testing cycle based on the resource demand in the previous testing cycle. Based on the resource requirements at the next detection time point within the previous detection cycle, predict the resource requirements at the next detection time point within the current detection cycle: , and These represent the predicted resource demand for the k-th resource at the next detection time point within the current detection period, and the actual resource demand for the k-th resource at the next detection time point within the previous detection period, respectively.

[0057] It should be noted in this embodiment that the next detection time point in the previous detection cycle corresponds to the next detection time point in the current detection cycle.

[0058] Step 3: Based on the resource demand predicted in the previous detection cycle for the next detection time point in the current detection cycle, and the resource demand predicted based on the next detection time point in the current detection cycle, perform weighted fusion to obtain the predicted resource demand for the next detection time point in the current detection cycle.

[0059] formula: , This represents the predicted resource demand for the k-th resource at the next detection time point within the current detection cycle. Let λ+β=1, where λ represents the predicted resource demand for the k-th resource at the next detection time point within the current detection period after weighted fusion, and β represents the weighted coefficient corresponding to the predicted resource demand at the next detection time point within the current detection period based on the resource demand at the next detection time point within the previous detection period. The values ​​of λ and β range from [0,1].

[0060] Specifically, the values ​​of λ and β are chosen based on the predicted resource demand at the next detection time point and the degree of dependence on the previous and current detection cycles. When the predicted data dependent on the current detection cycle is greater than the predicted data based on the previous detection cycle, β > λ; when the predicted data dependent on the current detection cycle is less than the predicted data based on the previous detection cycle, β < λ; and when the degree of dependence on the predicted data of the current detection cycle is approximately the same as that based on the previous detection cycle, β = λ.

[0061] By weighted and fused with historical data from the previous testing cycle and the previous testing time period, the resource demand for the next testing time point predicted based on the current testing time point in the current testing cycle can be adjusted using the historical data from the previous testing cycle. This effectively solves the lag problem in prediction from a single data source, ensuring resource supply during peak business periods while suppressing inflated resource allocation caused by occasional fluctuations, reducing operating costs, and achieving highly accurate calculation of resource demand before cloud resource allocation.

[0062] Step A5: Obtain the current environmental status of the cloud platform at the current detection time. Based on the current environmental status and the predicted resource demand of each tenant, determine the priority of multiple objectives. Dynamically allocate and regulate the resource data of the cloud platform according to the priority of multiple objectives to balance the resource allocation of multiple tenants and multiple objectives.

[0063] The multiple objectives include task completion time, energy consumption cost, and resource fairness among multiple tenants. The environmental status includes current CPU remaining amount, current memory remaining amount, remaining task completion time, and current electricity price.

[0064] Based on the determined priorities of the multiple objectives, the initial weight coefficients corresponding to each objective are selected, and the weight coefficients of each objective are adjusted using the comprehensive allocation evaluation coefficients determined by the multi-objective evaluation function to meet the resource allocation requirements of multi-tenant multi-objectives.

[0065] Since different environmental conditions can affect the priority determination of multiple objectives, the method for determining the priority of multiple objectives includes: Step 51: Calculate the load based on the predicted resource demand of each tenant, and determine whether the calculated load exceeds the set load threshold. If it exceeds the set load threshold, the task completion time has the highest priority among the multiple objectives. Step 52: Conversely, determine the energy consumption cost at the current detection time point and whether the energy consumption cost exceeds the set energy consumption threshold. If it exceeds the set energy consumption threshold, the energy consumption cost has the highest priority among the multiple targets. Step 53: Conversely, extract the deadline and reserved time of each task corresponding to each tenant. Based on the deadline and reserved time of each task, determine the start time of each task corresponding to each tenant. Determine whether there is a task whose start time is earlier than the current detection time. If so, extract the tenant whose task start time is earlier than the current detection time and determine the highest priority of resource fairness among multiple tenants. Step 54: If none of the above conditions are met, then set the priority of the multiple objectives to the same level.

[0066] The load threshold and energy consumption threshold are designed based on experience to meet the priority judgment conditions. Since the importance of task completion time is greater than the importance of energy consumption cost, and the importance of energy consumption cost is greater than the importance of resource fairness among multi-tenants, task completion time, energy consumption cost and resource fairness among multi-tenants are judged first.

[0067] By determining the priority of multiple objectives, the weights corresponding to the multiple objectives are initially determined. The weights corresponding to the multiple objectives are time weights. Energy consumption weight And fairness weight Based on the priorities among the multiple objectives, the initial weight values ​​for each objective are initially set. , and The range is [0, 1]. When multiple objectives have the same priority level, the weight coefficient of each objective is 1 / 3.

[0068] Based on the weight values ​​of the above multi-objectives, the rationality of the current resource data allocation is comprehensively evaluated considering the current environmental state and the predicted resource demands of each tenant. This evaluation of the rationality of the current resource data allocation employs a multi-objective evaluation function:

[0069] in, This represents the comprehensive allocation evaluation coefficient, used to reflect the rationality of resource allocation among tenants under the current environmental conditions. , and The parameter is dimensionless and its value ranges from [-1, 1]. This represents the time efficiency parameter corresponding to the completion of all tasks for all tenants. , and These represent the weights corresponding to task completion time and task waiting time, respectively. When a task is urgent, ,on the contrary, ,and , and The value range is [0,1], and can be adjusted according to the urgency of the task. This indicates the number of tasks completed at the current detection time. This indicates the total number of tasks pending processing in the cluster. This represents the average time for tasks to wait in the cluster. This indicates the maximum allowed waiting time threshold; This represents the energy consumption cost parameter at the current detection time point. , and These represent the weights of resource allocation and economic cost, respectively, with values ​​ranging from [0,1]. , This represents the total power consumption of the cluster at the current detection time. This represents the maximum power consumption when the cluster is fully loaded. V represents the electricity price at the current detection time, and V represents the benchmark electricity price.

[0070] This represents the fairness parameter at the current detection time point. , This represents the fairness sensitivity coefficient, with a value range of [0,1], where M represents the total number of tenants. This represents the amount of resources that should be allocated to the i-th tenant, which is the predicted resource demand of the i-th tenant. This represents the actual amount of resources allocated to the i-th tenant.

[0071] Determine the comprehensive allocation evaluation coefficient If the value is greater than 0, it indicates that the resource allocation under the current environmental condition is reasonable. If it is less than 0, it indicates that the resource allocation under the current environmental condition is not reasonable. Therefore, the initial weight values ​​of each objective are adjusted according to the value of the comprehensive allocation evaluation coefficient, so that the comprehensive allocation evaluation coefficient corresponding to the weight value of each objective after adjustment is greater than 0. This ensures that the allocation of the predicted resource demand of each tenant meets the current environmental condition, and realizes the allocation and regulation of resource data according to the priority of each objective after adjustment, so as to balance the resource allocation needs of multiple tenants and multiple objectives.

[0072] Specifically, if the comprehensive allocation evaluation coefficient If the weight is greater than 0, the weight coefficient of the highest priority target is increased, and the weight coefficients of other targets are decreased; conversely, the weight coefficient of the highest priority target is decreased, and the weight coefficients of other targets are increased, and the sum of the adjusted weight coefficients of all targets equals 1.

[0073] Based on the same inventive concept, such as Figure 3 As shown, this invention also discloses a multi-objective collaborative scheduling and optimization system for cloud-based AI training tasks, comprising: The resource demand forecasting module is used to extract the historical load data of each tenant within a detection cycle before the current detection time, analyze the historical load data of the detection cycle, and predict the resource demand required by each tenant at the next detection time. The prediction and judgment module is used to obtain the actual resource demand of each tenant and determine whether the deviation between the predicted resource demand and the actual resource demand of the same tenant is within the set safety control coefficient range. The set safety control coefficient range is [0, u], where u is greater than 0, so that the predicted resource demand is greater than the actual resource demand. If a resource demand is outside the set safety control coefficient range, then extract the historical load data corresponding to that tenant under the same event. The resource data fusion module is used to merge the resource demand predicted by the tenant at the current detection time point within the current detection cycle with the historical load data corresponding to the same event, so as to obtain the resource demand predicted at the next detection time point after fusion, so that the fused resource demand is close to the actual resource demand of the tenant. The multi-objective control module is used to obtain the environmental status of the cloud platform at the current detection time. Based on the current environmental status and the predicted resource demand of each tenant, it determines the priority of multiple objectives and dynamically allocates and controls the resource data of the cloud platform according to the priority of multiple objectives to balance the resource allocation of multiple tenants and multiple objectives.

[0074] The specific methods involved in this system can be found in the multi-objective cooperative scheduling optimization method described above, and will not be elaborated further.

[0075] In practical applications, a computer-readable storage medium can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0076] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0077] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0078] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0079] In the description of this specification, references to terms such as "an embodiment," "example," and "specific example" indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0080] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A multi-objective collaborative scheduling optimization method for cloud-based AI training tasks, characterized in that, include: Extract historical load data for each tenant within a detection cycle prior to the current detection time, analyze the historical load data within the detection cycle, and predict the resource requirements of each tenant at the next detection time. Obtain the actual resource requirements of each tenant, and determine whether the deviation between the predicted resource requirements and the actual resource requirements of the same tenant is within the set safety control coefficient range. The set safety control coefficient range is [0, u], where u is greater than 0, so that the predicted resource requirements are greater than the actual resource requirements. If a resource demand is outside the set safety control coefficient range, then extract the historical load data corresponding to that tenant under the same event. The resource demand predicted by the tenant at the current detection time point within the current detection cycle and the historical load data corresponding to the same event are merged to obtain the resource demand predicted at the next detection time point after fusion, so that the fused resource demand is close to the actual resource demand of the tenant. The system obtains the current environmental status of the cloud platform at the current detection time. Based on the current environmental status and the predicted resource requirements of each tenant, it determines the priority of multiple objectives and dynamically allocates and regulates the cloud platform's resource data according to the priority of multiple objectives to balance the resource allocation of multiple tenants and multiple objectives.

2. The multi-objective collaborative scheduling optimization method for cloud-based AI training tasks according to claim 1, characterized in that, Historical load data for each detection time point within a detection cycle prior to the previous detection time point is extracted. By analyzing the historical load data of each tenant, the burst coefficient of resource demand for each tenant at two adjacent detection time points is determined. Based on the set security control coefficient, the resource demand required by each tenant at the next detection time point is predicted. The resource demand includes CPU demand, memory demand, and bandwidth demand.

3. The multi-objective collaborative scheduling optimization method for cloud-based AI training tasks according to claim 2, characterized in that, Based on historical load data, the method for determining the CPU requirements for each tenant at the next detection time point includes: Extract the CPU utilization rate of each tenant at each detection time point within the detection period; Based on the CPU utilization rate corresponding to each detection time point, the CPU utilization rate threshold corresponding to the time ratio threshold r is selected, and the time ratio threshold r takes the value range [0,1]. Based on the CPU utilization corresponding to two adjacent detection time points, analyze the burst coefficient corresponding to the sudden change in CPU utilization, and use the set security control coefficient to analyze the secure CPU utilization of each tenant at the next detection time point. Based on the secure CPU utilization of each tenant, predict the CPU demand of each tenant at the next detection time point.

4. The multi-objective collaborative scheduling optimization method for cloud-based AI training tasks according to claim 3, characterized in that, Based on the predicted security CPU utilization of each tenant, the CPU demand at the next detection time is predicted using the number of CPU cores and the CPU utilization at the previous detection time.

5. The multi-objective collaborative scheduling optimization method for cloud-based AI training tasks according to claim 3, characterized in that, The method for calculating the burst coefficient corresponding to the sudden change in CPU utilization includes: obtaining the maximum CPU utilization of the same tenant in historical load data within the current detection period; taking the ratio between the maximum CPU utilization of the tenant and the CPU utilization threshold, and subtracting the value 1 to obtain the burst coefficient of CPU utilization, wherein the CPU utilization threshold is a set CPU utilization to control the utilization of the allocated CPU.

6. The multi-objective collaborative scheduling optimization method for cloud-based AI training tasks according to claim 3, characterized in that, Each tenant's actual resource demand at each detection time point is compared with the predicted resource demand at that detection time point to determine the degree of deviation between the predicted and actual resource demand of the same tenant. The determination of the degree of deviation is based on the ratio of the difference between the predicted and actual resource demand. The ratio of the difference is equal to the difference between the predicted and actual demand of the same tenant under the same resource conditions, and the ratio of the difference to the actual demand.

7. The multi-objective collaborative scheduling optimization method for cloud-based AI training tasks according to claim 6, characterized in that, If a tenant's predicted resource demand at the next detection time point is not within the set safety control coefficient range, the predicted resource demand will be dynamically adjusted.

8. The multi-objective collaborative scheduling optimization method for cloud-based AI training tasks according to claim 7, characterized in that, The methods used to dynamically adjust the predicted resource demand include: Extract the actual resource demand at the testing time point corresponding to the current testing time point of the current testing period within the previous testing period, and analyze the year-on-year growth rate coefficient; Based on the year-on-year growth rate and the resource demand G at the next testing time point in the previous testing cycle, we analyze the resource demand at the next testing time point in the current testing cycle, so as to predict the resource demand at the next testing time point in the current testing cycle based on the resource demand in the previous testing cycle. Based on the resource demand predicted for the next detection time point in the current detection period from the previous detection period, and the resource demand predicted for the next detection time point in the current detection period, a weighted fusion is performed to obtain the predicted resource demand for the next detection time point in the current detection period.

9. A multi-objective collaborative scheduling optimization method for cloud-based AI training tasks according to claim 8, characterized in that, Based on the determined priorities of the multiple objectives, the initial weight coefficients corresponding to each objective are selected, and the weight coefficients of each objective are adjusted using the comprehensive allocation evaluation coefficients determined by the multi-objective evaluation function to meet the resource allocation requirements of multi-tenant multi-objectives.

10. A multi-objective collaborative scheduling and optimization system for cloud-based AI training tasks, characterized in that, A multi-objective collaborative scheduling optimization method for cloud-based AI training tasks, applied to any one of claims 1-9, includes: The resource demand forecasting module is used to extract the historical load data of each tenant within a detection cycle before the current detection time, analyze the historical load data of the detection cycle, and predict the resource demand of each tenant at the next detection time. The prediction and judgment module is used to obtain the actual resource demand of each tenant and determine whether the deviation between the predicted resource demand and the actual resource demand of the same tenant is within the set safety control coefficient range. The set safety control coefficient range is [0, u], where u is greater than 0, so that the predicted resource demand is greater than the actual resource demand. If a resource demand is outside the set safety control coefficient range, then extract the historical load data corresponding to that tenant under the same event. The resource data fusion module is used to merge the resource demand predicted by the tenant at the current detection time point within the current detection cycle with the historical load data corresponding to the same event, so as to obtain the resource demand predicted at the next detection time point after fusion, so that the fused resource demand is close to the actual resource demand of the tenant. The multi-objective control module is used to obtain the environmental status of the cloud platform at the current detection time. Based on the current environmental status and the predicted resource demand of each tenant, it determines the priority of multiple objectives and dynamically allocates and controls the resource data of the cloud platform according to the priority of multiple objectives to balance the resource allocation of multiple tenants and multiple objectives.