Computing power cloud resource scheduling method and system based on adaptive learning

Through the adaptive learning computing power cloud resource scheduling method, long and short-term memory networks are used to train prediction models and optimize resource allocation, the problem of idle or overload in traditional scheduling methods is solved, and the high concurrency response capability and cost-effectiveness of cloud services are improved.

CN120336026AActive Publication Date: 2025-07-18邓慕超
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510492108.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Traditional cloud resource scheduling methods are difficult to cope with dynamic changes and diversity requirements of task load, resulting in idle or overload of resources, increasing operating costs and limiting the responsiveness of cloud services in high concurrency scenarios.

Method used

The computing power cloud resource scheduling method based on adaptive learning is adopted, and the prediction model is trained through long-term memory networks, combined with the system resource usage status and task priority, the resource allocation plan is optimized to reduce resource idleness or overload.

Benefits of technology

It improves the responsiveness of cloud services in high concurrency scenarios, reduces the probability of idle or overloading of resources, and reduces operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336026A_ABST
    Figure CN120336026A_ABST
Patent Text Reader

Abstract

The invention provides a computing power cloud resource scheduling method and system based on adaptive learning, and the method comprises the steps: training a long-short-term memory network through a training data set to obtain a prediction model, predicting a prediction value of a short-term resource demand through the prediction model, and carrying out the prediction of the short-term resource demand in combination with a current resource use state of a system. According to the method, whether the trend of sudden increase or decrease of the resource demand exists is judged to obtain an adjusted predicted value, the accuracy of the predicted value is ensured, the state of the resources in the system is obtained, then the task priorities of the system are sequenced, the computing power allocation proportion required by each task is determined to obtain a preliminary resource allocation scheme, and the resource allocation efficiency is improved. According to the method, the initial resource allocation scheme is optimized in combination with the constraint condition of resource scheduling to obtain an optimized resource allocation scheme, so that the probability of idle or overload of resources is reduced, the operation cost is reduced, and the response capability of cloud services in a high-concurrency scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and particularly to a computing power cloud resource scheduling method and system based on adaptive learning. Background Art

[0002] Currently, as the core pillar in the field of information technology, cloud computing's resource scheduling ability directly determines the balance between system performance and cost-effectiveness. Especially in computing-intensive applications such as artificial intelligence and big data analysis, efficient resource utilization has become the key to industry competition. However, traditional cloud resource scheduling methods mostly rely on static configurations or simple rule-based strategies, making it difficult to cope with the dynamic changes in task loads and diverse requirements, resulting in widespread phenomena of resource idleness or overload. This inefficiency not only increases operating costs but also limits the response ability of cloud services in high-concurrency scenarios.

[0003] Therefore, it is necessary to provide a new computing power cloud resource scheduling method and system based on adaptive learning to solve the above problems existing in the prior art. Summary of the Invention

[0004] The purpose of the present invention is to provide a computing power cloud resource scheduling method and system based on adaptive learning, which reduces the probability of resource idleness or overload.

[0005] To achieve the above purpose, the computing power cloud resource scheduling method based on adaptive learning of the present invention includes: Obtain the original training data, and then use the sliding window method to divide the original training data by time to obtain a training data set. The original training data includes the historical load task volume and the historical real-time task request volume; Train a long short-term memory network with the training data set to obtain a prediction model; Predict the predicted value of short-term resource requirements through the prediction model, and then combine the current resource usage status of the system to determine whether there is a trend of sudden increase or decrease in resource requirements, so as to obtain an adjusted predicted value, and determine whether it is necessary to increase the resources in the system according to the adjusted predicted value; Obtain the status of resources in the system, and then sort the task priorities of the system to determine the computing power allocation ratio required for each task, so as to obtain a preliminary resource allocation plan; Optimize the preliminary resource allocation plan in combination with the constraints of resource scheduling to obtain an optimized resource allocation plan; Schedule resources according to the optimized resource allocation plan.

[0006] The beneficial effects of the computing power cloud resource scheduling method based on adaptive learning are as follows: The long short-term memory network is trained through the training data set to obtain a prediction model. The prediction value of the short-term resource demand is predicted through the prediction model. Combining the current resource usage status of the system, it is judged whether there is a trend of sudden increase or decrease in resource demand to obtain an adjusted prediction value, ensuring the accuracy of the prediction value. The status of resources in the system is obtained, and then the task priorities of the system are sorted to determine the computing power allocation ratio required for each task to obtain a preliminary resource allocation plan. Combining the constraints of resource scheduling, the preliminary resource allocation plan is optimized to obtain an optimized resource allocation plan, reducing the probability of resource idleness or overload, reducing the operating cost, and improving the response ability of cloud services in high-concurrency scenarios.

[0007] Optionally, after obtaining the original training data, it further includes: Performing data cleaning on the original training data.

[0008] Optionally, before performing the training of the long short-term memory network through the training data set, it further includes: Comparing the historical load task volume in each time period with a preset load task volume threshold; Dividing each time period into a peak period or a trough period according to the comparison result, and assigning a first weight for training the long short-term memory network to the historical load task volume in the peak period, and assigning a second weight for training the long short-term memory network to the historical load task volume in the trough period, where the first weight is greater than the second weight.

[0009] Optionally, when predicting the prediction value of the short-term resource demand through the prediction model, it further includes: When the historical real-time task request volume changes suddenly, comparing the prediction value of the short-term resource demand predicted by the prediction model with a preset sudden threshold; If the prediction value of the short-term resource demand predicted by the prediction model exceeds the preset sudden threshold, then adjust the weight of the historical real-time task request volume in the prediction model and re-predict the prediction value of the short-term resource demand.

[0010] Optionally, combining the current resource usage status of the system to judge whether there is a trend of sudden increase or decrease in resource demand to obtain an adjusted prediction value, including: The current resource usage status of the system includes the real-time load quantity. Perform weighted averaging on the real-time load quantity and the prediction value of the short-term resource demand to obtain a prediction value during adjustment; Compare the prediction value during adjustment with a preset adjustment threshold. If the prediction value during adjustment is greater than the preset adjustment threshold, then judge the resource demand trend; Judge whether the resource demand trend is a sudden increase or a decrease by the change in the size of the predicted value in multiple consecutive adjustments to obtain a preliminary adjustment direction. Combine the preliminary adjustment direction with the task type of the real-time load to update the predicted value during adjustment and obtain the adjusted predicted value.

[0011] Optionally, the resources in the system are provided by several processors, and the status of the resources in the system includes the computing power that several processors are not loaded with; obtain the status of the resources in the system, then sort the task priorities of the system, and determine the computing power allocation ratio required for each task to obtain a preliminary resource allocation plan, including: Obtain the computing power that several processors in the system are not loaded with. Sort the task priorities according to the computing power requirements of the tasks in the system, and determine the computing power allocation ratio required for each task. Allocate tasks to several processors in turn. If the computing power required by the current task exceeds the computing power of the corresponding processor for the load, then allocate the current task to the next processor. If the same processor is allocated multiple tasks, process the tasks according to the task priorities.

[0012] Optionally, the constraint conditions for resource scheduling include the load computing power limit of the processor; combine the constraint conditions for resource scheduling to optimize the preliminary resource allocation plan to obtain an optimized resource allocation plan, including: Judge that after the processor is allocated tasks, the loaded computing power of the processor exceeds the load computing power limit of the processor, then re-allocate the processor for the task with the lowest priority to obtain an optimized resource allocation plan.

[0013] Optionally, the computing power cloud resource scheduling method based on adaptive learning further includes: Judge whether the result of subtracting the average computing power from the loaded computing power of the processor is greater than the equilibrium threshold. If it is judged that the result of subtracting the average computing power from the loaded computing power of the processor is greater than the equilibrium threshold, then allocate the task with the smallest required computing power in the processor with the largest loaded computing power to the processor with the smallest loaded computing power.

[0014] Optionally, the resources in the system are provided by several processors, and the computing power cloud resource scheduling method based on adaptive learning further includes: Judge whether the difference between the adjusted predicted value and the real-time load is greater than the target threshold. If it is judged that the difference between the adjusted predicted value and the real-time load is greater than the target threshold, then update the original training data to update the prediction model.

[0015] Optionally, the computing power cloud resource scheduling method based on adaptive learning further includes: Generate a predicted value through the updated prediction model; Determine whether the difference between the predicted value generated by the updated prediction model and the actual load is greater than the target threshold; If it is determined that the difference between the predicted value generated by the updated prediction model and the actual load is greater than the target threshold, continue to update the original training data to update the prediction model.

[0016] Optionally, the resources in the system are provided by several processors, and the method for scheduling computing power cloud resources based on adaptive learning further includes: Retrieve the scheduling execution log of the resources from the system; Obtain the idle rate and overload rate of the resources through the scheduling execution log; Adjust the scheduling frequency of the resources according to the idle rate and the overload rate. The present invention also provides a computing power cloud resource scheduling system based on adaptive learning, including a data processing unit, a training unit, a model operation unit, and a resource allocation unit. The data processing unit is used to obtain the original training data, and then use the sliding window method to segment the original training data by time to obtain a training data set. The original training data includes the historical load task volume and the historical real-time task request volume. The training unit is used to train the long short-term memory network through the training data set to obtain a prediction model. The model operation unit is used to predict the predicted value of the short-term resource demand through the prediction model, and then combine the current resource usage status of the system to determine whether there is a sudden increase or decrease trend in the resource demand to obtain an adjusted predicted value, and determine whether it is necessary to increase the resources in the system according to the adjusted predicted value. The resource allocation unit is used to optimize the preliminary resource allocation plan in combination with the constraints of resource scheduling to obtain an optimized resource allocation plan, and schedule resources according to the optimized resource allocation plan. Description of the Drawings

[0017] Figure 1 It is a flowchart of the method for scheduling computing power cloud resources based on adaptive learning in some embodiments of the present invention. Detailed Embodiments

[0018] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art to which the present invention pertains. The words such as "including" used herein mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items.

[0019] In view of the problems existing in the prior art, an embodiment of the present invention provides a computing power cloud resource scheduling method based on adaptive learning. Referring to Figure 1 , the computing power cloud resource scheduling method based on adaptive learning includes the following steps:

[0020] S1: Obtain the original training data, and then use the sliding window method to divide the original training data by time to obtain a training data set. The original training data includes the historical load task volume and the historical real-time task request volume. For example, at 14:00 on a certain past day, the number of tasks being processed by the system is 50, and the number of tasks waiting to be processed is 80. Then the load task volume is the sum of the number of tasks being processed and the number of tasks waiting to be processed by the system, that is, 130. Relative to this moment, it is the historical load task volume. The request volume of the task at the current moment is the historical real-time task request volume. For example, if it is 12:00 at this time, the task request volume at 12:00 is the historical real-time task request volume. If it is 15:00 at this time, the task request volume at 15:00 is the historical real-time task request volume.

[0021] In some embodiments, after obtaining the original training data, it further includes: cleaning the original training data. Among them, the original training data may have missing values or abnormal points due to the failure of the acquisition device or network delay in the system.

[0022] When there are missing values, the filling method can be used to fill the missing values. For example, if the historical load task volume of the system is missing within a certain hour, it can be filled with the average value of the previous and next two hours, or filled with a percentage of the average value of the previous and next two hours, which can be 70%.

[0023] When abnormal points appear, the median can be used to replace the abnormal points. For example, the historical load task volume of the system in the first minute is 50, the historical load task volume of the system in the third minute is 60, and the historical load task volume of the system in the second minute is 600. It can be seen that the historical load task volume of the system in the second minute is much larger than that in the first minute and the third minute, exceeding the reasonable range. Then, the median 55 of the historical load task volume of the system in the first minute and the third minute can be used to replace the historical load task volume 600 of the system in the second minute to ensure the smoothness of the data.

[0024] In some embodiments, the process of dividing the original training data by time using the sliding window method can be expressed as follows: setting the sliding window size to 1 hour and the step size to 15 minutes, then 24 hours are divided into 96 time periods, and the average value of the original training data is taken within each time period. For example, the first time period includes the first time point, the second time point, the third time point, and the fourth time point. The historical load task volume at the first time point is 44, the historical load task volume at the second time point is 48, the historical load task volume at the third time point is 46, and the historical load task volume at the fourth time point is 52. Then, the average value 45 of 44, 48, 46, and 42 is taken as the historical load task volume of the first time period.

[0025] In some embodiments, the original training data obtained in step S1 is at least the original training data of any whole day in the past, and can also be the original training data of any 2 days, any 3 days, or even any 30 days in the past.

[0026] In some embodiments, before training the long short-term memory network with the training dataset, it further includes: comparing the historical load task volume in each time period with a preset load task volume threshold; dividing each time period into a peak period or a trough period according to the comparison result, and assigning a first weight for training the long short-term memory network to the historical load task volume in the peak period, and assigning a second weight for training the long short-term memory network to the historical load task volume in the trough period, where the first weight is greater than the second weight.

[0027] The preset load task volume thresholds include the peak load task volume threshold and the off-peak load task volume threshold. Taking 75 as an example for the peak load task volume threshold and 25 as an example for the off-peak load task volume threshold, here, both the peak load task volume threshold and the off-peak load task volume threshold can be adjusted according to the actual situation, and no specific limitations are imposed here. Taking the total historical load task volume of each time period of any day as an example, the total historical load task volume from 14:00 to 15:00 is 90, and the total historical load task volume from 22:00 to 24:00 is 20. Since the total historical load task volume of 90 from 14:00 to 15:00 is greater than the peak load task volume threshold of 75, it is determined that the period from 14:00 to 15:00 is the peak period, while the total historical load task volume of 20 from 22:00 to 24:00 is less than the off-peak load task volume threshold of 25, so it is determined that the period from 22:00 to 24:00 is the off-peak period.

[0028] In some embodiments, the first weight is 0.6 and the second weight is 0.4. When the time of the peak period is 1.5 times the time of the off-peak period, the first weight can be adjusted to 0.7 and the second weight can be adjusted to 0.3 to form a balanced training set.

[0029] S2: Train the long short-term memory network through the training data set to obtain a prediction model.

[0030] In some embodiments, in step S2, the long short-term memory network (Long Short-Term Memory, LSTM) in the prior art is selected to construct the prediction model. Since the specific training process of the long short-term memory network is the prior art, it will not be elaborated in detail here.

[0031] S3: Predict the predicted value of the short-term resource demand through the prediction model, and then combine the current resource usage status of the system to determine whether there is a sudden increase or decrease trend in the resource demand to obtain the adjusted predicted value.

[0032] In some embodiments, when predicting the predicted value of the short-term resource demand through the prediction model, it further includes: when the historical real-time task request volume changes suddenly, compare the predicted value of the short-term resource demand predicted by the prediction model with a preset sudden change threshold; if the predicted value of the short-term resource demand predicted by the prediction model exceeds the preset sudden change threshold, adjust the weight of the historical real-time task request volume in the prediction model and re-predict the predicted value of the short-term resource demand.

[0033] In some embodiments, when the real-time task request volume is more than 1.5 times the previous task request volume, it is determined that the current real-time task request volume has changed suddenly.

[0034] In some embodiments, the historical load task volume and the historical real-time task request volume jointly determine the prediction model. When there is a sudden change in the historical real-time task request volume, the predicted value of the short-term resource demand predicted by the prediction model may be affected by the historical real-time task request volume, resulting in an abnormally large predicted value of the short-term resource demand predicted by the prediction model. Therefore, when there is a sudden change in the historical real-time task request volume, the predicted value of the short-term resource demand is compared with the preset sudden change threshold. When the predicted value of the short-term resource demand is greater than the preset sudden change threshold, it means that the predicted value of the short-term resource demand predicted by the prediction model is abnormally large. By adjusting the weight of the historical real-time task request volume in the prediction model, the influence on the predicted value when there is a sudden change in the historical real-time task request volume can be reduced.

[0035] In some embodiments, in combination with the current resource usage status of the system, it is determined whether there is a trend of sudden increase or decrease in resource demand to obtain an adjusted predicted value, including: The current resource usage status of the system includes the real-time load quantity. The real-time load quantity and the predicted value of the short-term resource demand are weighted and averaged to obtain a predicted value during adjustment; The predicted value during adjustment is compared with a preset adjustment threshold. If the predicted value during adjustment is greater than the preset adjustment threshold, the resource demand trend is determined; Through the magnitude change of multiple consecutive predicted values during adjustment, it is determined whether the resource demand trend is a sudden increase or a decrease to obtain a preliminary adjustment direction; The preliminary adjustment direction is combined with the task type of the real-time load to update the predicted value during adjustment to obtain an adjusted predicted value. Among them, when the real-time load quantity becomes smaller, the predicted value during adjustment becomes smaller, and when the real-time load quantity becomes larger, the predicted value during adjustment also becomes larger accordingly. This method can smooth the influence of mutations in historical data through the integration of history and real-time.

[0036] In some embodiments, taking the predicted value of the short-term resource demand as 60 as an example, that is, the prediction model predicts that the task volume in the next stage is 60, and the real-time load quantity is 80, the weight of the real-time load quantity is 0.4, and the predicted value of the short-term resource demand is 0.6. Then, by weighting and averaging the real-time load quantity and the predicted value of the short-term resource demand, the predicted value during adjustment can be obtained as 68.

[0037] In some embodiments, taking the preset adjustment threshold as 70 as an example, the predicted value during adjustment 68 is less than the preset adjustment threshold 70, so the predicted value during adjustment is updated to 68, and the adjusted predicted value obtained is 68.

[0038] In some other embodiments, taking the preset adjustment threshold as 70 as an example, during the adjustment, if the predicted value 78 is greater than the preset adjustment threshold 70, and a number of predicted values obtained during the adjustment are retrieved. If the several predicted values obtained during the adjustment sequentially include 68, 70, and 75 in chronological order, it can be determined that the resource demand trend is a sudden increase, and the preliminary adjustment direction is to increase the predicted value; if the several predicted values obtained during the adjustment sequentially include 88, 85, and 79 in chronological order, it can be determined that the resource demand trend is a decrease, and the preliminary adjustment direction is to decrease the predicted value.

[0039] In some embodiments, the task types of the real-time load can be divided into normal operation tasks, difficult data operation tasks, and simple data operation tasks. Correspondingly, the difficult data operation tasks require a large amount of resources, while the simple data operation tasks require a small amount of resources. That is to say, the difficult data operation tasks occupy a larger amount of resources than the simple data operation tasks.

[0040] For example, one normal operation task needs to occupy 50 units of resources in the system, one difficult data operation task needs to occupy 80 units of resources in the system, and one simple data operation task needs to occupy 20 units of resources in the system.

[0041] When the task type of the real-time load is a difficult data operation task, it means that the task type in the next stage is also likely to be a difficult data operation task, and a task in the next stage requires more system resources. To avoid system overload, it is necessary to increase the estimated predicted value. If the preliminary adjustment direction is to increase the predicted value, then increase it by 20% based on the predicted value during the adjustment to obtain the adjusted predicted value; if the preliminary adjustment direction is to decrease the predicted value, then decrease it by 10% based on the predicted value during the adjustment to obtain the adjusted predicted value.

[0042] When the task type of the real-time load is a simple data operation task, it means that the task type in the next stage is also likely to be a simple data operation task, and a task in the next stage requires less system resources. To avoid system overload, it is necessary to decrease the estimated predicted value. If the preliminary adjustment direction is to increase the predicted value, then increase it by 10% based on the predicted value during the adjustment to obtain the adjusted predicted value; if the preliminary adjustment direction is to decrease the predicted value, then decrease it by 20% based on the predicted value during the adjustment.

[0043] In some embodiments, the adjusted predicted value is compared with the number of tasks with the highest load in the system. When the ratio of the adjusted predicted value to the number of tasks with the highest load in the system is greater than or equal to 80%, it is determined that the system needs to increase resources, for example, increasing from 3 servers to [number of servers not specified in the original text] servers.

[0044] S4: Obtain the status of resources in the system based on the adjusted predicted values, then sort the task priorities of the system, and determine the computing power allocation ratio required for each task to obtain a preliminary resource allocation plan.

[0045] In some embodiments, the resources in the system are provided by several processors, and the status of the resources in the system includes the computing power not loaded by several processors; obtaining the status of the resources in the system, then sorting the task priorities of the system, and determining the computing power allocation ratio required for each task to obtain a preliminary resource allocation plan includes: obtaining the computing power not loaded by several processors in the system; sorting the task priorities according to the computing power requirements of the tasks in the system, and determining the computing power allocation ratio required for each task; allocating tasks to several processors in turn. If the computing power required by the current task exceeds the computing power of the corresponding processor for loading, then allocate the current task to the next processor. If the same processor is allocated multiple tasks, process the tasks according to the task priorities.

[0046] Taking a system including a group of Graphics Processing Unit (GPU) clusters as an example, a group of GPU clusters includes 4 GPUs. The total computing power of each GPU is 100 units, and the total computing power of 4 GPUs is 400 units. 400 units are the resources in the system.

[0047] The 4 GPUs are the first GPU, the second GPU, the third GPU, and the fourth GPU respectively. Among them, the first GPU is fully loaded, the second GPU is fully loaded, the third GPU is idle, and the fourth GPU is idle. Then the status of the resources in the system is that the first GPU is fully loaded, the second GPU is fully loaded, the third GPU is idle, and the fourth GPU is idle.

[0048] Suppose there are 5 tasks, namely the first task, the second task, the third task, the fourth task, and the fifth task. The computing power required for the first task is 50 units, the computing power required for the second task is 30 units, the computing power required for the third task is 40 units, the computing power required for the fourth task is 20 units, and the computing power required for the fifth task is 10 units. Sort the task priorities of the system according to the computing power required by the tasks, that is, the greater the computing power required by the task, the higher the priority of the task. Then the priorities of the first task, the second task, the third task, the fourth task, and the fifth task are the first task, the third task, the second task, the fourth task, and the fifth task. Among them, the computing power allocation ratio required for the first task is 50 / 150, the computing power allocation ratio required for the second task is 30 / 150, the computing power allocation ratio required for the third task is 40 / 150, the computing power allocation ratio required for the fourth task is 20 / 150, and the computing power allocation ratio required for the fifth task is 10 / 150.

[0049] Allocate the first task to the first processor, the second task to the second processor, the third task to the third processor, the fourth task to the fourth processor, and the fifth task to the first processor. However, since the first processor and the second processor are fully loaded, allocate the first task and the fifth task to the third processor. At this time, the third processor processes a total of three tasks, namely the first task, the third task, and the fifth task. The total computing power required for the three tasks is 100 units. At this time, the computing power of the third processor can meet the computing power requirements of the three tasks. If the computing power of the third processor cannot meet the computing power requirements of the three tasks, allocate the tasks to the next processor, that is, the fourth processor; the third processor processes the first task, the third task, and the fifth task in turn according to the task priorities. Allocate the second task to the fourth processor. At this time, the fourth processor processes a total of two tasks, namely the second task and the fourth task. The total computing power required for the two tasks is 50 units (such as based on the number of CPU / GPU cores or floating-point operation capabilities). At this time, the computing power of the fourth processor can meet the computing power requirements of the three tasks; the fourth processor processes the first task and the third task in turn according to the task priorities.

[0050] S5: Combine the constraint conditions of resource scheduling to optimize the preliminary resource allocation plan to obtain an optimized resource allocation plan.

[0051] In some embodiments, the constraint conditions of resource scheduling include the load computing power limit of the processor; combining the constraint conditions of resource scheduling to optimize the preliminary resource allocation plan to obtain an optimized resource allocation plan, including: judging that after the processor is assigned tasks, the computing power already loaded by the processor exceeds the load computing power limit of the processor, then reallocate the processor for the task with the lowest priority to obtain an optimized resource allocation plan.

[0052] Taking the load computing power limit of the processor as 80% of the load as an example, that is, if the total computing power of the graphics processing unit is 100 units, then the graphics processing unit can process a maximum of 80 units of tasks.

[0053] Suppose the graphics processing unit processes two tasks, namely the first task and the second task. The priority of the first task is higher than that of the second task. The computing power requirement of the first task is 50 units, and the computing power requirement of the second task is 35 units. The total computing power requirement of the two tasks is 85 units. When the graphics processing unit processes the first task and the second task, the load is 85%, exceeding the load capacity limit of 80%. Then, the second task needs to be reallocated to the processor.

[0054] In some embodiments, the computing power cloud resource scheduling method based on adaptive learning includes: determining whether the result of subtracting the average computing power from the loaded computing power of the processor is greater than the balance threshold; if it is determined that the result of subtracting the average computing power from the loaded computing power of the processor is greater than the balance threshold, then allocate the task with the smallest required computing power in the processor with the largest loaded computing power to the processor with the smallest loaded computing power. Here, the average computing power refers to the average of the loaded computing powers of all processors.

[0055] Suppose there are 4 graphics processing units, namely the first graphics processing unit, the second graphics processing unit, the third graphics processing unit, and the fourth graphics processing unit. The first graphics processing unit processes the first task, the second graphics processing unit processes the second task, the third graphics processing unit processes the third task, and the fourth graphics processing unit processes the fourth task and the fifth task. The computing power required for the first task is 38 units, so the loaded computing power of the first graphics processing unit is 38 units. The computing power required for the second task is 35 units, so the loaded computing power of the second graphics processing unit is 35 units. The computing power required for the third task is 41 units, so the loaded computing power of the third graphics processing unit is 41 units. The computing power required for the fourth task is 40 units, and the computing power required for the fifth task is 6 units, so the loaded computing power of the fourth graphics processing unit is 46 units. The average computing power of the first graphics processing unit, the second graphics processing unit, the third graphics processing unit, and the fourth graphics processing unit is 40 units. If the balance threshold is set to 5 units, then the result of subtracting the average computing power from the loaded computing power of the fourth graphics processing unit is 6 units, which is greater than the balance threshold of 5 units. Then, allocate the fifth task with the smallest required computing power in the fourth graphics processing unit to the second graphics processing unit with the smallest loaded computing power.

[0056] S6: Schedule resources according to the optimized resource allocation plan.

[0057] In some embodiments, the resources in the system are provided by several processors, and the method for scheduling computing power cloud resources based on adaptive learning further includes: determining whether the difference between the adjusted predicted value and the real-time load is greater than a target threshold; if it is determined that the difference between the adjusted predicted value and the real-time load is greater than the target threshold, updating the original training data to update the prediction model.

[0058] In some embodiments, the method for scheduling computing power cloud resources based on adaptive learning further includes: generating a predicted value through the updated prediction model; determining whether the difference between the predicted value generated by the updated prediction model and the actual load is greater than the target threshold; if it is determined that the difference between the predicted value generated by the updated prediction model and the actual load is greater than the target threshold, continuing to update the original training data to update the prediction model.

[0059] After scheduling the computing power cloud resources by generating an adjusted predicted value through the prediction model, the system will generate new original training data. If the difference between the adjusted predicted value and the real-time load is greater than the target threshold, it indicates that there is a large deviation in the predicted value generated by the prediction model, and thus the prediction model needs to be corrected and updated. Updating the prediction model with the new original training data can reduce the deviation of the subsequently generated predicted values and make the generated predicted values meet the requirements.

[0060] In some embodiments, the method for scheduling computing power cloud resources based on adaptive learning further includes: retrieving the scheduling execution log of the resources from the system; obtaining the idle rate and overload rate of the resources through the scheduling execution log; and adjusting the scheduling frequency of the resources according to the idle rate and the overload rate.

[0061] Suppose there are 2 graphics processors, namely the first graphics processor and the second graphics processor. Within a preset time period, for example, 10 minutes, the idle rate of the first graphics processor is 20% and the overload rate is 1%. Since the idle rate of the first graphics processor is too high, the scheduling frequency of the first graphics processor needs to be increased. The idle rate of the second graphics processor is 1% and the overload rate is 25%. Since the overload rate of the second graphics processor is too high, the scheduling frequency of the second graphics processor needs to be decreased, that is, the tasks that should have been assigned to the second graphics processor can be preferentially assigned to the first graphics processor.

[0062] The present invention also provides a computing power cloud resource scheduling system based on adaptive learning for implementing the above-mentioned computing power cloud resource scheduling, which includes a data processing unit, a training unit, a model operation unit, and a resource allocation unit. The data processing unit is used to obtain original training data, and then divide the original training data by time using a sliding window method to obtain a training data set. The original training data includes historical load task volume and historical real-time task request volume. The training unit is used to train a long short-term memory network through the training data set to obtain a prediction model. The model operation unit is used to predict the predicted value of short-term resource demand through the prediction model, and then combine the current resource usage status of the system to determine whether there is a trend of sudden increase or decrease in resource demand to obtain an adjusted predicted value, and determine whether it is necessary to increase resources in the system according to the adjusted predicted value. The resource allocation unit is used to optimize the preliminary resource allocation plan in combination with the constraints of resource scheduling to obtain an optimized resource allocation plan, and schedule resources according to the optimized resource allocation plan.

[0063] Although the embodiments of the present invention have been described in detail above, it is obvious to those skilled in the art that various modifications and changes can be made to these embodiments. However, it should be understood that such modifications and changes are all within the scope and spirit of the present invention described in the claims. Moreover, the present invention described herein may have other embodiments and can be implemented or realized in various ways.

Claims

1. A computing power cloud resource scheduling method based on adaptive learning, characterized in that, Including: Obtain the original training data, and then use the sliding window method to divide the original training data by time to obtain a training data set. The original training data includes the historical load task volume and the historical real-time task request volume; Train the long short-term memory network through the training data set to obtain a prediction model; Predict the predicted value of the short-term resource demand through the prediction model, and then combine the current resource usage status of the system to determine whether there is a sudden increase or decrease trend in the resource demand, so as to obtain an adjusted predicted value, and judge whether it is necessary to increase the resources in the system according to the adjusted predicted value; Obtain the status of the resources in the system, then sort the task priorities of the system, and determine the computing power allocation ratio required for each task to obtain a preliminary resource allocation plan; Optimize the preliminary resource allocation plan in combination with the constraints of resource scheduling to obtain an optimized resource allocation plan; Schedule resources according to the optimized resource allocation plan.

2. The method for scheduling computing power cloud resources based on adaptive learning according to claim 1, wherein, Before performing the training of the long short-term memory network through the training data set, it also includes: Compare the historical load task volume in each time period with a preset load task volume threshold; Divide each time period into a peak period or a trough period according to the comparison result, and assign a first weight for training the long short-term memory network to the historical load task volume in the peak period, and assign a second weight for training the long short-term memory network to the historical load task volume in the trough period. The first weight is greater than the second weight.

3. The method for scheduling computing power cloud resources based on adaptive learning according to claim 1, wherein When predicting the predicted value of the short-term resource demand through the prediction model, it also includes: When the historical real-time task request volume changes suddenly, compare the predicted value of the short-term resource demand predicted by the prediction model with a preset sudden threshold; If the predicted value of the short-term resource demand predicted by the prediction model exceeds the preset sudden threshold, adjust the weight of the historical real-time task request volume in the prediction model and re-predict the predicted value of the short-term resource demand.

4. The method for scheduling computing power cloud resources based on adaptive learning according to claim 1, wherein, Combining the current resource usage status of the system to determine whether there is a sudden increase or decrease trend in the resource demand to obtain an adjusted predicted value, including: The current resource usage status of the system includes the real-time load quantity. Perform a weighted average on the real-time load quantity and the predicted value of the short-term resource demand to obtain a predicted value during adjustment; Compare the predicted value during adjustment with a preset adjustment threshold. If the predicted value during adjustment is greater than the preset adjustment threshold, judge the resource demand trend; Judge whether the resource demand trend is a sudden increase or decrease through the size change of multiple consecutive predicted values during adjustment to obtain a preliminary adjustment direction; Combine the preliminary adjustment direction with the task type of the real-time load to update the predicted value during adjustment to obtain an adjusted predicted value.

5. The method for scheduling computing power cloud resources based on adaptive learning according to claim 1, characterized in that The resources in the system are provided by several processors. The status of the resources in the system includes the computing power unloaded by several processors; obtain the status of the resources in the system, then sort the task priorities of the system, and determine the computing power allocation ratio required for each task to obtain a preliminary resource allocation plan, including: Obtain the computing power unloaded by several processors in the system; Sort the task priorities according to the computing power requirements of the system's tasks, and determine the computing power allocation ratio required for each task; Allocate tasks to several processors in sequence. If the computing power required by the current task exceeds the computing power of the corresponding processor's load, then allocate the current task to the next processor. If the same processor is allocated multiple tasks, process the tasks according to the task priorities.

6. The method for scheduling computing power cloud resources based on adaptive learning according to claim 5, characterized in that, The constraint conditions for resource scheduling include the load computing power limit of the processor; combining the constraint conditions for resource scheduling, optimize the preliminary resource allocation plan to obtain an optimized resource allocation plan, including: Judge that after the processor is allocated tasks, if the computing power already loaded by the processor exceeds the load computing power limit of the processor, then reallocate the processor for the task with the lowest priority to obtain an optimized resource allocation plan.

7. The method for scheduling computing power cloud resources based on adaptive learning according to claim 5 or 6, characterized in that It also includes: Judge whether the result of the computing power already loaded by the processor minus the average computing power is greater than the balance threshold; If it is judged that the result of the computing power already loaded by the processor minus the average computing power is greater than the balance threshold, then allocate the task with the smallest required computing power in the processor with the largest already loaded computing power to the processor with the smallest already loaded computing power.

8. The method for scheduling computing power cloud resources based on adaptive learning according to claim 1, wherein The resources in the system are provided by several processors, and the computing power cloud resource scheduling method based on adaptive learning further includes: Judge whether the difference between the adjusted predicted value and the real-time load is greater than the target threshold; If it is judged that the difference between the adjusted predicted value and the real-time load is greater than the target threshold, then update the original training data to update the prediction model; Generate a predicted value through the updated prediction model; Judge whether the difference between the predicted value generated by the updated prediction model and the actual load is greater than the target threshold; If it is judged that the difference between the predicted value generated by the updated prediction model and the actual load is greater than the target threshold, then continue to update the original training data to update the prediction model.

9. The method for scheduling computing power cloud resources based on adaptive learning according to claim 1, wherein The resources in the system are provided by several processors, and the computing power cloud resource scheduling method based on adaptive learning further includes: Retrieve the scheduling execution log of the resources from the system; Obtain the idle rate and overload rate of the resources through the scheduling execution log; Adjust the scheduling frequency of the resources according to the idle rate and the overload rate.

10. A computing power cloud resource scheduling system based on adaptive learning, characterized in that, It includes a data processing unit, a training unit, a model operation unit, and a resource allocation unit. The data processing unit is used to obtain the original training data, and then use the sliding window method to divide the original training data by time to obtain a training data set. The original training data includes the historical load task volume and the historical real-time task request volume. The training unit is used to train the long short-term memory network through the training data set to obtain a prediction model. The model operation unit is used to predict the predicted value of the short-term resource demand through the prediction model, and then combine the current resource usage status of the system to judge whether there is a trend of sudden increase or decrease in resource demand to obtain an adjusted predicted value, and judge whether it is necessary to increase the resources in the system according to the adjusted predicted value. The resource allocation unit is used to combine the constraint conditions for resource scheduling, optimize the preliminary resource allocation plan to obtain an optimized resource allocation plan, and schedule resources according to the optimized resource allocation plan.

Citation Information

Patent Citations

  • Service prediction based load balancing method

    CN102711177A

  • Method and system for improving computing power efficiency

    CN118550711A

  • Memory management method of virtual machine and computing device

    CN118689588A

  • Task allocation method and device, computer equipment, readable storage medium and program product

    CN119127419A

  • Computing resource scheduling method and device, storage medium and electronic equipment

    CN119473567A