A Resource Dynamic Management Method and Management System for Hybrid Loads
The dynamic resource management method addresses the challenge of mixed workloads by predicting resource demands and reservations for delay-sensitive tasks, improving resource utilization and reducing scheduling delays and waste.
Patent Information
- Application Number
- CN202211130391.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-16
AI Technical Summary
The prior art is difficult to accurately reserve resources in clusters deployed with hybrid loads, resulting in increased scheduling delays and waste of resources for latency-sensitive tasks, affecting application performance and resource utilization.
By predicting the resource requirements of delay-sensitive tasks and the time period of insufficient cluster idle resources, resources are accurately reserved and released after the demand is over, a time series prediction model such as the Prophet model predicts resource utilization and demand, and optimizes resource management.
It realizes the resource requirements of delay-sensitive tasks, reduces scheduling overhead, improves resource utilization, and avoids resource waste.
Smart Images

Figure CN115562853B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of cloud computing cluster resource management, and more specifically, relates to a resource dynamic management method and management system for hybrid loads. Background Art
[0002] In order to reduce operation and maintenance costs and make full use of cluster resources, Internet companies such as Google, Alibaba, and Microsoft Bing have begun to use hybrid loads to deploy data center clusters. Since batch processing jobs require a large amount of resources when running and are usually executed at night or when cluster resources are idle, they can make full use of cluster resources and help improve resource utilization. In contrast, latency-sensitive loads for users require fewer resources, and the load volume changes with user activities. Based on the above runtime characteristics, a hybrid load composed of latency-sensitive applications for users and batch processing jobs has become the load composition paradigm in hybrid deployments. However, running multiple loads in the same cluster will make the cluster runtime environment more complex, thus bringing performance impacts to tasks on different levels of the software stack and damaging application performance. The huge performance requirement differences between different types of loads make these performance impacts more obvious and more harmful to a certain type of load.
[0003] In a cluster with a hybrid deployment of latency-sensitive applications and batch processing jobs, latency-sensitive applications have performance characteristics such as high throughput and low latency. Therefore, it is necessary to strictly control the completion time of latency-sensitive tasks; while batch processing jobs do not have strict performance requirements, as long as the processing work can be completed as soon as possible. Therefore, in contrast, hybrid load deployment will cause greater performance losses to latency-sensitive applications. And tasks, as the "ultimate executors" of jobs, the performance loss of applications is essentially that the tasks of the corresponding jobs are interfered with in terms of performance. Existing research has found through analyzing cluster scheduling logs and load runtime information that task scheduling delays sometimes account for 60% of the task completion time. Therefore, ensuring task scheduling efficiency and avoiding unnecessary scheduling overheads are crucial for improving task performance and enhancing application service quality. In a hybrid load deployment cluster, since multiple loads share cluster resources, there will be fewer idle resources in the cluster. In this situation, if a large number of latency-sensitive loads are submitted to the cluster, these loads will enter the queuing state because the idle resources in the cluster cannot meet the resource requirements of the tasks, waiting for the running tasks in the cluster to release resources, resulting in a sharp increase in the scheduling delay of latency-sensitive tasks and seriously damaging the service quality of latency-sensitive applications.
[0004] Existing work mainly meets the resource requirements of the load by reserving sufficient cluster resources for latency-sensitive applications, thereby avoiding unnecessary scheduling overhead caused by insufficient idle resources. According to the way of implementing resource reservation, existing work can be divided into two categories, namely static resource reservation and dynamic resource reservation. Among them, static resource reservation is to specify a certain amount of resources in the cluster specifically for running specific loads. In a cluster with stable load, static resource reservation can well ensure sufficient available resources for specific loads. However, due to the rapid development of information technology and the diversity of user needs, the current data center load has high volatility and uncertainty, so it is difficult to determine the appropriate amount of reserved resources. In a cluster with a hybrid workload deployment, excessive reserved resources are contrary to the original intention of resource sharing and cause waste of cluster resources; while shortages of reserved resources will lead to some latency-sensitive tasks waiting for resources, increasing task scheduling latency. Dynamic resource reservation is to dynamically adjust the amount of reserved resources in the cluster according to the resource requests of users in the cluster or the resource requirements of waiting tasks. Although the dynamic resource reservation method can further ensure the accuracy of reserved resources by dynamically adjusting the reserved resources, it is usually difficult for ordinary users to accurately configure the resource requirements for the load, so the situation of inappropriate reserved resources will inevitably occur. In addition, existing work dynamically adjusts the reserved resources according to the resource requests of users or the resource requirements of waiting tasks, and the operation lag will also lead to an increase in task scheduling latency. Therefore, the dynamic resource reservation method still faces great challenges in determining the operation execution time. Summary of the Invention
[0005] In view of the defects and improvement requirements of the prior art, the present invention provides a method and a management system for dynamic resource management for hybrid loads, aiming to accurately predict the resource requirements of latency-sensitive loads and the time when there is a shortage of idle resources in the cluster, so as to accurately reserve resources for latency-sensitive loads on demand, reduce scheduling overhead, and improve resource utilization.
[0006] To achieve the above object, according to one aspect of the present invention, there is provided a method for dynamic resource management for hybrid loads, including:
[0007] Cluster status prediction and monitoring step: predicting the time period during which the resource requirements of latency-sensitive tasks cannot be met in each management cycle as the resource reservation time period in the corresponding management cycle, and predicting the resource requirements of latency-sensitive tasks during the resource reservation time period;
[0008] Resource reservation management step: reserving resources according to the prediction result of the resource requirements of latency-sensitive tasks during the resource reservation time period before the start of the resource reservation time period, and canceling the reserved resources after the end of the resource reservation time period.
[0009] Further, the cluster status prediction and monitoring steps include:
[0010] Taking the moments in the management cycle at intervals of Δt to form a set of prediction time points S = {t n |t n = t0 + n×Δt, t0 ≤ t n ≤ t0 + T, n = 0, 1, 2…}; t0 represents the start time of the management cycle, and T represents the length of the management cycle;
[0011] Identifying the moments when the resource requirements of delay-sensitive tasks in the set of prediction time points S are not met as resource reservation moments, predicting the total resource requirements of delay-sensitive tasks in the cluster at the resource reservation moments as the resource reservation amounts corresponding to the resource reservation moments, and storing the resource reservation moments and the corresponding resource reservation amounts into set P and set TR respectively;
[0012] If set P is not empty, determining the time period between the smallest resource reservation moment and the largest resource reservation moment in it as the resource reservation time period, and taking the largest resource reservation amount in set TR as the resource requirements of delay-sensitive tasks within the resource reservation time period.
[0013] Further, identifying the resource reservation moments in the set of prediction time points S and predicting the total resource requirements of delay-sensitive tasks in the cluster at the resource reservation moments includes:
[0014] Traversing the moments in the set of prediction time points S, and for the traversed moment t n , performing the following steps:
[0015] (S1) Predicting the resource utilization rate U of the cluster and the total resource requirements R n of newly arrived delay-sensitive tasks at time t k ;
[0016] (S2) Calculating the idle resource amount in the cluster at time t n as: R f = R c ×(1 - U);
[0017] (S3) If R k > R f , then identifying the moment t n as a resource reservation moment and transferring to step (S4); otherwise, transferring to step (S5);
[0018] (S4) Predicting the total resource requirements R n of delay-sensitive tasks in the running state in the cluster at time t x , and calculating t nThe total resource demand of latency-sensitive tasks in the cluster at time t is R t = R k + R x ;
[0019] (S5) End the operation at time t n of.
[0020] Further, at time t n , the prediction method of the resource utilization rate U of the cluster includes: using the historical data of the resource utilization rate at each moment within a specified time period before time t n to construct a first prediction sequence, and inputting it into the trained resource utilization rate prediction model to obtain the resource utilization rate U at time t n ;
[0021] At time t n , the prediction method of the total resource demand R k of the newly arrived latency-sensitive tasks in the cluster includes: using the historical data of the total resource demand of the newly arrived latency-sensitive tasks in the cluster at each moment within a specified time period before time t n to construct a second prediction sequence, and inputting it into the trained new resource demand prediction model to obtain the total resource demand R n of the newly arrived latency-sensitive tasks in the cluster at time t k ;
[0022] At time t n , the prediction method of the total resource demand R x of the latency-sensitive tasks in the running state in the cluster includes: using the historical data of the total resource demand of the latency-sensitive tasks in the running state in the cluster at each moment within a specified time period before time t n to construct a third prediction sequence, and inputting it into the trained resource demand prediction model to obtain the total resource demand R n of the latency-sensitive tasks in the running state in the cluster at time t x ;
[0023] Among them, the resource utilization rate prediction model, the new resource demand prediction model, and the resource demand prediction model are all time series prediction models.
[0024] Further, the resource utilization rate prediction model, the new resource demand prediction model, and the resource demand prediction model are all Prophet models.
[0025] Further, in the resource reservation management step, before the start of the resource reservation time period, reserve resources according to the prediction results of the resource demands of the latency-sensitive tasks within the resource reservation time period, including:
[0026] (T1) When the cluster runs to ts - At time ΔT1, determine whether the resources reserved in the previous management cycle have been released. If so, set the reserved resource amount to R v = 0; otherwise, set the reserved resource amount R v to the resource amount reserved in the previous management cycle;
[0027] (T2) If the resource demand R of delay-sensitive tasks during the resource reservation time period in the current management cycle tmax = R v , then go to step (T5); if R tmax < R v , then go to step (T3); if R tmax > R v , then go to step (T4);
[0028] (T3) Release the reserved resources of size R v - R tmaxt , and go to step (T5);
[0029] (T4) Select resources from the resources occupied by delay-sensitive tasks in the running state in the cluster, the idle resources in the cluster, and the resources released by batch processing tasks in the cluster in sequence to perform the reservation operation until the newly added reserved resource amount reaches R tmax - R v ;
[0030] (T5) The resource reservation operation ends;
[0031] where t s is the start time of the resource reservation time period in the current management cycle, and ΔT1 is a preset time interval.
[0032] Furthermore, step (T4) includes:
[0033] (T41) Calculate the remaining resource amount to be reserved as R w = R tmax - R v ;
[0034] (T42) Count the resource amount R a occupied by delay-sensitive tasks in the running state in the cluster. If R a ≥ R w , then select resources of size R w from the resources occupied by delay-sensitive tasks in the running state in the cluster for the reservation operation, and then go to step (T5); if R a < R w , then perform the reservation operation on all the resources occupied by delay-sensitive tasks in the running state in the cluster, and according to R w = Rw -R a The remaining resource amount R to be reserved w ;
[0035] (T43) Statistically calculate the resource amount R in the idle state in the cluster idle , if R idle ≥R w , then select resources with a size of R w from the resources in the idle state in the cluster to perform the reservation operation, and then transfer to step (T5); otherwise, perform the reservation operation on all the resources in the idle state in the cluster, and calculate according to R w =R w -R idle to update the remaining resource amount R to be reserved w ;
[0036] (T44) Obtain the resource amount R released by the batch processing tasks in the cluster b , if R b ≥R w , then select resources with a size of R b from the resources R w released by the batch processing tasks in the cluster to perform the reservation operation, and then transfer to step (T5); if R b <R w , then perform the reservation operation on all the resources released by the batch processing tasks in the cluster, and calculate according to R w =R w -R b to update the remaining resource amount R to be reserved w , and then transfer to step (T44).
[0037] Furthermore, in the resource reservation management step, after the resource reservation time period ends, cancel the reserved resources, including:
[0038] (W1) When the cluster runs to time t e +ΔT2, determine whether the current time has entered the next management cycle. If so, transfer to step (W2); otherwise, transfer to step (W3);
[0039] (W2) Determine whether the resource reservation operation has been performed in the next management cycle at the current time. If so, transfer to step (W3); otherwise, transfer to step (W3);
[0040] (W3) Release all the reserved resources in the cluster;
[0041] (W4) End the cancellation of the resource reservation operation;
[0042] where t eis the end time of the resource reservation time period within the current management cycle, and ΔT2 is a preset time interval.
[0043] According to another aspect of the present invention, there is provided a resource dynamic management system for hybrid loads, including:
[0044] A cluster status prediction and monitoring module, configured to predict the time period during which the resource requirements of delay-sensitive tasks cannot be met in each management cycle, as the resource reservation time period within the corresponding management cycle, and predict the resource requirements of delay-sensitive tasks during the resource reservation time period;
[0045] And a resource reservation management module, configured to reserve resources according to the prediction result of the resource requirements of delay-sensitive tasks during the resource reservation time period before the start of the resource reservation time period, and cancel the reserved resources after the end of the resource reservation time period.
[0046] According to yet another aspect of the present invention, there is provided a computer-readable storage medium, including: a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the resource dynamic management method for hybrid loads provided by the present invention.
[0047] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0048] (1) The resource dynamic management method and management system for hybrid loads provided by the present invention divide the cluster operation process into management cycles. By means of prediction, the time period during which the resource requirements of delay-sensitive tasks cannot be met in each management cycle, that is, the resource reservation time period, and the resource requirements of delay-sensitive tasks during this time period are determined in advance. Before the cluster runs to this time period, resources are reserved according to the determined resource requirements, and after the end of this time period, the reserved resources are released, realizing accurate and on-demand resource reservation for delay-sensitive loads, reducing scheduling overhead, and at the same time being able to avoid resource waste and improve the resource utilization rate of the cluster.
[0049] (2) In the preferred embodiment of the resource dynamic management method and management system for hybrid loads provided by the present invention, when determining the resource reservation time period in the management cycle and the resource requirements of delay-sensitive tasks during this time period, specifically, detection time points are selected at fixed time intervals from the entire management cycle, and the detection time points at which the resource requirements of delay-sensitive tasks cannot be met are identified as the resource reservation times when resource reservation needs to be performed. Then, the resource requirements of delay-sensitive tasks at the corresponding times are predicted as the resource reservation amounts. Finally, the time period between the maximum resource reservation time and the minimum resource reservation time is determined as the resource reservation time period, and the maximum resource reservation amount is determined as the amount of resources to be reserved during this time period. This can ensure that after the resource reservation operation is executed, the resource requirements of delay-sensitive tasks can be met at all times in the management cycle, and it also avoids large scheduling overheads caused by frequent execution of resource reservation and release operations.
[0050] (3) In the preferred embodiment of the resource dynamic management method and management system for hybrid loads provided by the present invention, specifically, by predicting the resource utilization rate of the cluster, the total resource requirements of newly arrived delay-sensitive tasks, and the total resource requirements of delay-sensitive tasks in the running state at each detection time point, it is determined whether the resource requirements of delay-sensitive tasks will not be met at the corresponding detection time point. When this situation occurs, the amount of resources to be reserved at the corresponding time is calculated based on the prediction results. This can accurately predict the task execution situation and resource requirements at each detection time point, providing a strong decision-making basis for determining the resource reservation time period and the resource reservation amount during this time period.
[0051] (4) In the preferred embodiment of the resource dynamic management method and management system for hybrid loads provided by the present invention, when predicting the resource utilization rate of the cluster, the total resource requirements of newly arrived delay-sensitive tasks, and the total resource requirements of delay-sensitive tasks in the running state at each detection time point, specifically, a time series prediction model is used to predict based on historical time series data. This prediction method is consistent with the operating characteristics of the cluster, ensuring the accuracy of the prediction results.
[0052] (5) In the preferred embodiment of the resource dynamic management method and management system for hybrid loads provided by the present invention, the time to execute the reservation operation is the time t s before the start time t s of the resource reservation time period minus ΔT1. This can effectively avoid an increase in task scheduling delay caused by the lag of the reservation operation. Moreover, when the reserved resource amount is greater than the actual resource requirements of delay-sensitive tasks, the over-reserved resources will be released, avoiding resource waste while ensuring that the resource requirements of delay-sensitive loads are met.
[0053] (6) In the preferred embodiment of the resource dynamic management method and management system for hybrid loads provided by the present invention, the time to cancel the resource reservation operation is the end time t of the resource reservation period e after t e +ΔT2, and the reserved resources are released only when the resource reservation operation is not performed in the next management cycle. While avoiding resource waste, it reduces unnecessary resource reservation operations and operations to cancel resource reservations, effectively improving the utilization rate of cluster resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flowchart of the resource dynamic management method for hybrid loads provided by an embodiment of the present invention;
[0055] Figure 2 is a schematic diagram of the resource dynamic management method for hybrid loads provided by an embodiment of the present invention;
[0056] Figure 3 is a flowchart of the prediction model construction provided by an embodiment of the present invention;
[0057] Figure 4 is a flowchart of the cluster status prediction and monitoring steps provided by an embodiment of the present invention;
[0058] Figure 5 is a flowchart of the resource reservation operation provided by an embodiment of the present invention;
[0059] Figure 6 is a flowchart of the operation to cancel the reservation provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0061] In the present invention, terms such as "first", "second", etc. (if any) in the present invention and the accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0062] Aiming at the problem that the existing work cannot accurately and effectively reserve appropriate resources for delay-sensitive applications, introducing unnecessary scheduling overhead and resource waste, the present invention provides a method and a management system for dynamic resource management for hybrid loads. The overall idea is as follows: accurately predict the time period when the available resources in the cluster are insufficient and the resource requirements of the load, and based on the prediction results, reserve the cluster resources as needed before the situation of insufficient available resources occurs, and cancel the resource reservation in a timely manner after the situation ends, so as to ensure the resource requirements of delay-sensitive loads, avoid introducing scheduling overhead to tasks, reduce unnecessary resource reservation, and improve the utilization rate of cluster resources.
[0063] Before explaining the technical solution of the present invention in detail, the following explanations are made for relevant technical terms:
[0064] Job: The logical instance of an application or load in the cluster;
[0065] Task: The smallest unit for job execution. A job usually consists of one or more tasks that cooperate to process.
[0066] Cluster Manager (ResourceManager, RM): The process program in the cluster that manages the entire cluster resources, schedules and manages tasks.
[0067] The following are embodiments.
[0068] Embodiment 1:
[0069] A method for dynamic resource management for hybrid loads, as Figure 1 and Figure 2 shown, includes:
[0070] Cluster status prediction and monitoring step: Predict the time period when the resource requirements of delay-sensitive tasks cannot be met in each management cycle as the resource reservation time period in the corresponding management cycle, and predict the resource requirements of delay-sensitive tasks during the resource reservation time period;
[0071] Resource reservation management step: Before the start of the resource reservation time period, reserve resources according to the prediction results of the resource requirements of delay-sensitive tasks during the resource reservation time period, and cancel the reserved resources after the end of the resource reservation time period.
[0072] In this embodiment, in order to accurately predict the time period during which the resource requirements of delay-sensitive tasks are not met in each management cycle, that is, the resource reservation time period, and the resource requirements of delay-sensitive tasks during the resource reservation time period, detection time points will be selected at fixed time intervals from the entire management cycle, and the detection time points at which the resource requirements of delay-sensitive tasks are not met will be identified as the resource reservation times that need to perform resource reservation, and the resource requirements of delay-sensitive tasks at the corresponding times will be predicted as the resource reservation amounts. Finally, the time period between the maximum resource reservation time and the minimum resource reservation time will be determined as the resource reservation time period, and the maximum resource reservation amount will be determined as the resource amount that needs to be reserved during this time period. For the judgment of each detection time point, the resource utilization rate of the cluster, the total resource requirements of newly arrived delay-sensitive tasks, and the total resource requirements of delay-sensitive tasks in the running state are predicted using a machine learning model, and the relevant calculations are completed. Considering that the running characteristics of the cluster make the resource utilization rate and the resource requirements of delay-sensitive tasks have a large correlation with historical data, in this embodiment, a time series prediction model is selected. Specifically, a Prophet model is selected according to the data characteristics and their changing trends.
[0073] Before the Prophet model is specifically used for prediction, corresponding training is required. In this embodiment, three models are trained, namely, a resource utilization rate prediction model, a newly added resource requirement prediction model, and a resource requirement prediction model, which are respectively used to predict the resource utilization rate of the cluster at a specific time, the total resource requirements of newly arrived delay-sensitive tasks in the cluster, and the total resource requirements of delay-sensitive tasks in the running state in the cluster. Taking the resource utilization rate prediction model as an example, refer to Figure 3 , and its training process includes:
[0074] Extract the historical data related to the resource utilization rate from the cluster running logs, historical workload information, resource utilization rate and other data in the cluster resource manager to construct a training sample set; each training sample consists of the resource utilization rate of the cluster at a certain time and the time series composed of the resource utilization rates of the cluster at each time within a specified time period before that time.
[0075] Select Prophet according to the data characteristics and their changing trends and construct a prediction model. Use the time series in the training sample as the model input and the resource utilization rate of the cluster at the selected time as the label information to train the established Prophet prediction model; after the training is completed, a resource utilization rate prediction model is obtained.
[0076] To ensure the prediction accuracy of the model, in this embodiment, after the model training is completed, new training samples will be periodically constructed using newly generated data to strengthen the training of the resource utilization rate prediction model.
[0077] The prediction model for the newly added resource demand and the training method of the resource demand prediction model are similar to the training method of the above-mentioned resource utilization prediction model. The difference is that for the prediction model of the newly added resource demand, in the constructed training sample set, each training sample is composed of the total resource demand of the newly arrived latency-sensitive tasks in the cluster at a certain moment, and the time series composed of the total resource demands of the newly arrived latency-sensitive tasks in the cluster at each moment within a specified time period before that moment; for the resource demand prediction model, in the constructed training sample set, each training sample is composed of the total resource demand of the latency-sensitive tasks in the running state in the cluster at a certain moment, and the total resource demands of the latency-sensitive tasks in the running state in the cluster at each moment within a specified time period before that moment. Similarly, in this embodiment, after the prediction models for the newly added resource demand and the resource demand are trained, the models will also be intensively trained.
[0078] It should be noted that the Prophet model is only an optional time series prediction model in this embodiment and should not be construed as a limitation to the present invention. Other models that can accurately complete relevant predictions can also be used in the present invention.
[0079] Refer to Figure 4 , based on the established prediction model, in this embodiment, the steps of predicting and monitoring the cluster state specifically include:
[0080] Taking the moments in the management period at intervals of Δt to form a set of prediction time points D = {t n |t n = t0 + n×Δt, t0 ≤ t n ≤ t0 + T, n = 0, 1, 2...}; t0 represents the start moment of the management period, and T represents the length of the management period;
[0081] Identifying the moments when the resource demands of the latency-sensitive tasks in the set of prediction time points S cannot be met as the resource reservation moments, and predicting the total resource demand of the latency-sensitive tasks in the cluster at the resource reservation moments as the resource reservation amounts corresponding to the resource reservation moments, and storing the resource reservation moments and the corresponding resource reservation amounts into the set P and the set TR respectively;
[0082] If the set P is not empty, then determining the time period between the smallest resource reservation moment and the largest resource reservation moment in it as the resource reservation time period, and taking the largest resource reservation amount in the set TR as the resource demand of the latency-sensitive tasks within the resource reservation time period;
[0083] Identifying the resource reservation moments in the set of prediction time points S and predicting the total resource demand of the latency-sensitive tasks in the cluster at the resource reservation moments specifically includes:
[0084] Traverse the moments in the set S of prediction time points. For the traversed moment t n , perform the following steps:
[0085] (S1) Predict the resource utilization rate U of the cluster and the total resource demand R of the newly arrived latency-sensitive tasks at time t n ; k ;
[0086] Specifically, at time t n , the prediction method of the resource utilization rate U of the cluster includes: constructing a first prediction sequence using the historical data of the resource utilization rate at each moment within a specified time period before time t n , and inputting it into the trained resource utilization rate prediction model to obtain the resource utilization rate U at time t n ;
[0087] At time t n , the prediction method of the total resource demand R of the newly arrived latency-sensitive tasks in the cluster includes: constructing a second prediction sequence using the historical data of the total resource demand of the newly arrived latency-sensitive tasks in the cluster at each moment within a specified time period before time t k , and inputting it into the trained new resource demand prediction model to obtain the total resource demand R of the newly arrived latency-sensitive tasks in the cluster at time t n ; n ; k ;
[0088] Among them, the historical data can be extracted from data such as cluster operation logs and historical workload information;
[0089] (S2) Calculate the idle resource amount in the cluster at time t n as: R f = R c × (1 - U);
[0090] (S3) If R k > R f , it means that the idle resources in the cluster at time t n are not enough to meet the resource requirements of the latency-sensitive tasks. Then, identify the moment t n as the resource reservation moment and transfer to step (S4); otherwise, transfer to step (S5);
[0091] (S4) Predict the total resource demand R n of the latency-sensitive tasks in the running state in the cluster at time t x , and calculate the total resource demand of the latency-sensitive tasks in the cluster at time t n as R t = R k+R x ;
[0092] Specifically, at time t n , the total resource demand R x of the latency-sensitive tasks in the running state in the cluster can be predicted in the following ways: constructing a third prediction sequence using the historical data of the total resource demand of the latency-sensitive tasks in the running state in the cluster at each moment within a specified time period before time t n , inputting it into the trained resource demand prediction model, and obtaining the total resource demand R n of the latency-sensitive tasks in the running state in the cluster at time t x ;
[0093] (S5) End the operation on time t n .
[0094] Refer to Figure 5 . In this embodiment, in the resource reservation management step, before the start of the resource reservation time period, resources are reserved according to the prediction result of the resource demand of the latency-sensitive tasks within the resource reservation time period, including:
[0095] (T1) When the cluster runs to time t s - ΔT1, determine whether the resources reserved in the previous management cycle have been released. If so, set the reserved resource amount to R v = 0; otherwise, set the reserved resource amount R v to the resource amount reserved in the previous management cycle;
[0096] t s is the start time of the resource reservation time period in the current management cycle, and ΔT1 is a preset time interval;
[0097] (T2) If the resource demand R tmax of the latency-sensitive tasks within the resource reservation time period in the current management cycle is equal to R v , it means that the resources already reserved in the system can exactly meet the resource requirements of the latency-sensitive tasks, and then go to step (T5); if R tmax < R v , it means that the resources already reserved in the system exceed the resources actually required by the latency-sensitive tasks, and then go to step (T3); if R tmax > R v , it means that the resources already reserved in the system are not enough to meet the resource requirements of the latency-sensitive tasks, and then go to step (T4);
[0098] (T3) Set the size to R v - R tmaxtRelease the reserved resources to avoid resource waste and improve resource utilization rate, and then proceed to step (T5);
[0099] (T4) Select resources from the resources occupied by latency-sensitive tasks in the running state in the cluster, the idle resources in the cluster, and the resources released by batch processing tasks in the cluster in sequence to perform the reservation operation until the newly added reserved resource volume reaches R tmax -R v ;
[0100] In step (T4) of this embodiment, when resources are insufficient, resources occupied by latency-sensitive tasks in the running state, idle resources, and resources released by batch processing tasks will be selected in sequence to perform resource reservation operations, and only when one type of resource cannot meet the resource requirements, resources will be selected from the next type of resource for reservation, ensuring that the resource requirements of latency-sensitive tasks are met while minimizing the impact on other tasks in the cluster; step (T4) specifically includes:
[0101] (T41) Calculate the remaining resource volume to be reserved as R w =R tmax -R v ;
[0102] (T42) Statistically calculate the resource volume R a occupied by latency-sensitive tasks in the running state in the cluster. If R a ≥R w , then select resources with a size of R w from the resources occupied by latency-sensitive tasks in the running state in the cluster for reservation operation, and then proceed to step (T5); if R a <R w , then perform reservation operations on all the resources occupied by latency-sensitive tasks in the running state in the cluster, and update the remaining resource volume to be reserved R w =R w -R a ; w ;
[0103] (T43) Statistically calculate the resource volume R idle of the idle resources in the cluster. If R idle ≥R w , then select resources with a size of R w from the idle resources in the cluster for reservation operation, and then proceed to step (T5); otherwise, perform reservation operations on all the idle resources in the cluster, and update the remaining resource volume to be reserved R w =R w -R idle ; w ;
[0104] (T44) Obtain the amount of resources R released by the batch tasks in the cluster b , if R b ≥R w , then select resources with a size of R from the resources R released by the batch tasks in the cluster b to perform the reservation operation, and then transfer to step (T5); if R w <R b , then perform the reservation operation on all the resources released by the batch tasks in the cluster, and update the amount of resources R that still need to be reserved according to R w =R w -R w , and then transfer to step (T44); b w
[0105] (T5) The resource reservation operation ends;
[0106] Figure 6 Refer to
[0107] , in this embodiment, in the resource reservation management step, after the end of the resource reservation time period, cancel the reserved resources, including:
[0108] (W1) When the cluster runs to the time t e +ΔT2, judge whether the current time has entered the next management cycle. If so, transfer to step (W2); otherwise, transfer to step (W3);
[0109] (W2) Judge whether the resource reservation operation has been performed in the next management cycle at the current time. If so, transfer to step (W3); otherwise, transfer to step (W3);
[0109] (W3) Release all the reserved resources in the cluster;
[0110] (W4) The resource reservation cancellation operation ends;
[0111] Among them, t e is the end time of the resource reservation time period in the current management cycle, and ΔT2 is a preset time interval;
[0112] It is easy to understand that, in this embodiment, for the judgment of whether the resource reservation operation needs to be performed at each moment within a single management cycle, it is carried out according to the time interval Δt. In order to accurately judge whether the resources reserved in the previous management cycle have been released before the resource reservation operation starts in the current management cycle, in this embodiment, specifically set Δt<ΔT1<T. Similarly, in order to accurately judge whether the resource reservation operation has been performed in the next management cycle before releasing the reserved resources in the current management cycle, specifically set Δt<ΔT2<T.
[0113] Embodiment 2:
[0114] A resource dynamic management system for hybrid loads, comprising:
[0115] A cluster status prediction and monitoring module, configured to execute the cluster status prediction and monitoring steps in the above-mentioned Embodiment 1, that is, to predict the time period during which the resource requirements of delay-sensitive tasks cannot be met in each management cycle as the resource reservation time period in the corresponding management cycle, and to predict the resource requirements of delay-sensitive tasks during the resource reservation time period;
[0116] And a resource reservation management module, configured to execute the resource reservation management steps in the above-mentioned Embodiment 1, that is, to reserve resources according to the prediction result of the resource requirements of delay-sensitive tasks during the resource reservation time period before the start of the resource reservation time period, and to cancel the reserved resources after the end of the resource reservation time period;
[0117] In this embodiment, the specific implementation manners of each module may refer to the description in the above-mentioned Embodiment 1 and will not be repeated here.
[0118] Embodiment 3:
[0119] A computer-readable storage medium, comprising: a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the resource dynamic management method for hybrid loads provided in the above-mentioned Embodiment 1.
[0120] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for dynamic resource management for hybrid loads, characterized in that, Including: Cluster status prediction and monitoring steps: At intervals, take the moments in the management cycle to form a set of prediction time points ; represents the start time of the management cycle, and T represents the length of the management cycle; Identify the moments when the resource requirements of latency-sensitive tasks in the set S of prediction time points are not met as resource reservation moments, and predict the total resource requirements of latency-sensitive tasks in the cluster at the resource reservation moments as the resource reservation amounts corresponding to the resource reservation moments, and store the resource reservation moments and the corresponding resource reservation amounts into the set P and the set TR respectively; If the set P is not empty, determine the time period between the smallest resource reservation moment and the largest resource reservation moment in it as the resource reservation time period, and use the largest resource reservation amount in the set TR as the resource requirements of latency-sensitive tasks within the resource reservation time period; Resource reservation management steps: Before the start of the resource reservation time period, reserve resources according to the prediction results of the resource requirements of latency-sensitive tasks within the resource reservation time period, and cancel the reserved resources after the end of the resource reservation time period; Among them, identifying the resource reservation moments in the set S of prediction time points and predicting the total resource requirements of latency-sensitive tasks in the cluster at the resource reservation moments includes: Traverse the moments in the set S of prediction time points, and for the traversed moment , perform the following steps: (S1) Predict at the moment, the resource utilization rate of the cluster U and the total resource demand of the newly arrived latency-sensitive tasks ; (S2) Calculate The amount of idle resources in the cluster at the moment is: ; If (S3) , then identify the moment as the resource reservation moment and proceed to step (S4); otherwise, proceed to step (S5). (S4) Prediction The total resource demand of latency-sensitive tasks in the cluster that are in the running state at time , and calculate The total resource demand of latency-sensitive tasks in the cluster at time is ; (S5) End the operation at time .
2. The resource dynamic management method for hybrid loads according to claim 1, wherein At time, the prediction method of the resource utilization rate of the cluster U includes: using the historical data of the resource utilization rate at each moment within a specified time period before time to construct a first prediction sequence, inputting it into the trained resource utilization rate prediction model, and obtaining the resource utilization rate at time U ; At the moment, the total resource demand of newly arrived latency-sensitive tasks in the cluster is predicted by: using the historical data of the total resource demand of newly arrived latency-sensitive tasks in the cluster at each moment within a specified time period before the moment to construct a second prediction sequence, and inputting it into the trained prediction model for the newly added resource demand to obtain the total resource demand of newly arrived latency-sensitive tasks in the cluster at the moment; At the moment, the total resource demand of latency-sensitive tasks in the running state in the cluster is predicted by: using the historical data of the total resource demand of latency-sensitive tasks in the running state in the cluster at each moment within a specified time period before the moment to construct a third prediction sequence, and inputting it into the trained resource demand prediction model to obtain the total resource demand of latency-sensitive tasks in the running state in the cluster at Among them, the resource utilization prediction model, the new resource requirement prediction model, and the resource requirement prediction model are all time series prediction models.
3. The resource dynamic management method for hybrid loads according to claim 2, wherein The resource utilization prediction model, the new resource requirement prediction model, and the resource requirement prediction model are all Prophet models.
4. The resource dynamic management method for hybrid loads according to any one of claims 1 to 3, characterized in that In the resource reservation management steps, before the start of the resource reservation time period, reserving resources according to the prediction results of the resource requirements of latency-sensitive tasks within the resource reservation time period includes: (T1) When the cluster runs to the moment, determine whether the resources reserved in the previous management cycle have been released. If so, set the reserved resource quantity to ; otherwise, set the reserved resource quantity to the resource quantity reserved in the previous management cycle; (T2) If the resource requirement of delay-sensitive tasks during the resource reservation time period within the current management cycle , then go to step (T5); if , then go to step (T3); if , then go to step (T4); (T3) Release the reserved resource of size and proceed to step (T5); (T4) sequentially selects resources from the resources occupied by latency-sensitive tasks in the running state in the cluster, the idle resources in the cluster, and the resources released by batch tasks in the cluster to perform the reservation operation until the newly added reserved resource amount reaches ; (T5) The resource reservation operation ends; Among them, is the start time of the resource reservation time period within the current management cycle, is a preset time interval.
5. The resource dynamic management method for hybrid loads according to claim 4, wherein The step (T4) includes: The amount of resources that still need to be reserved is calculated as ; (T42) Statistically calculate the amount of resources occupied by latency-sensitive tasks in the running state in the cluster If then select the resources with a size of among the resources occupied by latency-sensitive tasks in the running state in the cluster for reservation operations, and then proceed to step (T5); if then perform reservation operations on all the resources occupied by latency-sensitive tasks in the running state in the cluster, and update the amount of resources still to be reserved according to ; (T43) Statistically determine the amount of idle resources in the cluster , if , then select resources with a size of from the idle resources in the cluster to perform a reservation operation, and then proceed to step (T5); otherwise, perform a reservation operation on all the idle resources in the cluster and update the amount of resources still to be reserved according to ; ; (T44) Obtain the amount of resources released by batch tasks in the cluster ,like , then the resources released from the batch processing tasks in the cluster Select the size of The resource is reserved, and then the process goes to step (T5); if , all resources released by batch processing tasks in the cluster are reserved, and Update the amount of resources that need to be reserved , then proceed to step (T44).
6. The resource dynamic management method for hybrid loads as claimed in claim 4, wherein, In the resource reservation management steps, after the end of the resource reservation time period, canceling the reserved resources includes: When the cluster runs to At the moment, it is judged whether the current moment has entered the next management cycle. If so, go to step (W2); otherwise, go to step (W3); (W2) Judge whether the resource reservation operation has been executed in the next management cycle at the current moment. If so, go to step (W3); otherwise, go to step (W3); (W3) Release all the reserved resources in the cluster; (W4) The resource reservation cancellation operation ends; Wherein, is the end time of the resource reservation time period within the current management cycle, is a preset time interval.
7. A resource dynamic management system for hybrid loads, characterized in that, Including: A cluster status prediction and monitoring module for performing: At intervals, take the moments in the management cycle to form a set of prediction time points ; denotes the start time of the management cycle, and T denotes the length of the management cycle; Identify the moments when the resource requirements of latency-sensitive tasks in the set S of prediction time points are not met as resource reservation moments, and predict the total resource requirements of latency-sensitive tasks in the cluster at the resource reservation moments as the resource reservation amounts corresponding to the resource reservation moments, and store the resource reservation moments and the corresponding resource reservation amounts into the set P and the set TR respectively; If the set P is not empty, determine the time period between the smallest resource reservation moment and the largest resource reservation moment in it as the resource reservation time period, and use the largest resource reservation amount in the set TR as the resource requirements of latency-sensitive tasks within the resource reservation time period; And a resource reservation management module for reserving resources according to the prediction results of the resource requirements of latency-sensitive tasks within the resource reservation time period before the start of the resource reservation time period, and canceling the reserved resources after the end of the resource reservation time period; Among them, identifying the resource reservation moments in the set S of prediction time points and predicting the total resource requirements of latency-sensitive tasks in the cluster at the resource reservation moments, including: Traverse the moments in the set S of prediction time points, and for the traversed moment , perform the following steps: (S1) Predict at the moment, the resource utilization rate of the cluster U and the total resource requirement of newly arrived latency-sensitive tasks ; (S2) Calculate The amount of idle resources in the cluster at the moment is: ; If (S3) then identify the time as the resource reservation time and proceed to step (S4); otherwise, proceed to step (S5); (S4) Prediction The total resource demand of latency-sensitive tasks in the cluster at a certain moment , and calculate The total resource demand of latency-sensitive tasks in the cluster at a certain moment is ; (S5) End the operation at time .
8. A computer-readable storage medium, characterized in that, Including: A stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the resource dynamic management method for hybrid loads according to any one of claims 1 to 6.
Citation Information
Patent Citations
Resource Allocation Method And Resource Borrowing Method
US20220156115A1
Cluster node load state prediction-based job scheduling method
WO2020206705A1