A Dynamic Scheduling Method for GPU Resources in Cloud Environments
The dynamic GPU resource scheduling method optimizes cloud environment resource allocation by predicting demands and adjusting resource pools, reducing operational costs and improving task execution efficiency.
Patent Information
- Application Number
- CN202510594377.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-09
AI Technical Summary
In the prior art, GPU resource dynamic scheduling strategies fail to effectively solve the problems of low resource utilization and increased response delay caused by fluctuations in task demand, hardware heterogeneity, and multi-tenant competition. Especially in cloud computing, edge computing and high-performance computing scenarios, the resource pool adjustment strategy is based on high operating costs and task scheduling delays caused by a single indicator.
By monitoring the operating indicators of GPU instances in real time, generating memory status evaluation results, predicting the total resource demand, combining resource pool capacity configuration, optimizing resource pool adjustment decisions, based on the matching degree list between tasks and GPUs, dynamically selecting target instances and performing task scheduling, continuously tracking resource consumption, and forming a closed-loop scheduling optimization link.
It improves resource utilization and task matching accuracy, reduces task interruption, alleviates the problem of long-tail task hunger, and achieves a balance between resource utilization and task fairness.
Smart Images

Figure CN120104357B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of dynamic scheduling of GPU resources, and particularly to a method for dynamically scheduling GPU resources in a cloud environment. Background Art
[0002] The technical field of dynamic scheduling of GPU resources focuses on the real-time allocation, optimization, and rebalancing of graphics processing unit (GPU) resources in a distributed computing environment. Especially in cloud computing, edge computing, and high-performance computing scenarios, it solves problems such as low resource utilization and increased response latency caused by task demand fluctuations, hardware heterogeneity, and multi-tenant competition.
[0003] In the prior art, resource pool adjustment strategies mostly trigger scaling based on a single metric (such as the total video memory gap), ignoring the dynamic coupling relationship between the cold start time of instances and the release cost, which is prone to cause operation oscillations. For example, after quickly expanding to high-configured instances to cope with short-term peak loads, due to the high release cost, redundant resources are forced to be held for a long time, driving up the operating cost. Task scheduling usually allocates resources based on static priorities or polling mechanisms, without incorporating the instance state change time into the scheduling decision, resulting in high-priority tasks being forced to wait due to the startup delay of the target instance. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the drawbacks existing in the prior art, and to propose a method for dynamically scheduling GPU resources in a cloud environment.
[0005] To achieve the above purpose, the present invention adopts the following technical solution. A method for dynamically scheduling GPU resources in a cloud environment includes the following steps:
[0006] Real-time monitor the running metrics of each GPU instance in the cloud environment, collect GPU video memory allocation records, GPU video memory fragmentation degree values, and page migration activity counts, and generate a video memory state evaluation result by statistically predicting the total task resource requirements;
[0007] Based on the video memory state evaluation result, compare the predicted total resource requirements with the total capacity configuration of the current cloud environment GPU resource pool to obtain a resource capacity gap or redundancy. Based on the resource capacity gap or redundancy, query and evaluate the startup time parameters of newly added GPU instances and the associated cost parameters of releasing GPU instances, and establish a GPU resource pool adjustment decision;
[0008] Based on the video memory status evaluation result, access the task list to be scheduled, extract the GPU video memory requirements of each task, generate a task-GPU matching degree list, and based on the task-GPU matching degree list, combine the time information of the instance state change in the GPU resource pool adjustment decision to select a target GPU instance for each task in the task queue to be processed, and establish a task scheduling sequence to be executed;
[0009] Based on the task scheduling sequence to be executed, send the task allocation information and the corresponding GPU instance identifier to the GPU resource manager in the cloud environment, issue a GPU video memory pre-allocation or locking instruction according to the sequence content, obtain the task distribution execution result list, and based on the task distribution execution result list, trigger the start of the task process on the target GPU instance, continuously track the GPU resource consumption during the task operation, and obtain the GPU scheduling execution efficiency record.
[0010] Preferably, the steps for obtaining the video memory status evaluation result are as follows:
[0011] Call the video memory allocation monitoring interface of the GPU instance to obtain the video memory allocation record of each GPU instance within the time window, generate a video memory fragmentation degree value by calculating the standard deviation of the free area between consecutive video memory blocks, traverse the kernel log to count the number of page migrations per unit time to obtain the page migration activity count, and form a basic feature set;
[0012] Based on the basic feature set, construct an autoregressive prediction model to calculate the total task resource requirements in the future time window, and the expression is:
[0013] ;
[0014] Among them, is the total task resource requirement, is the video memory allocation record of the previous time window, is the change rate of the video memory fragmentation degree value over time, is the current page migration activity count, is the autoregressive coefficient fitted through historical data, is the white noise error term;
[0015] Compare the total task resource requirement with the confidence interval threshold If , where is the prediction benchmark value calculated by the moving average of historical data, then it is determined that the prediction is effective, and a video memory status evaluation result is generated.
[0016] Preferably, the steps for obtaining the resource capacity gap or redundancy are as follows:
[0017] Obtain the total amount of task resource requirements recorded in the video memory status evaluation result, and at the same time access the resource pool configuration database to extract the total capacity configuration parameters of the current GPU resource pool, and generate a resource requirement-capacity matching data set;
[0018] Based on the resource requirement-capacity matching data set, calculate the absolute difference between the total amount of task resource requirements and the total capacity configuration, and compare the difference result with the preset capacity fluctuation threshold. If the absolute value of the difference exceeds the capacity fluctuation threshold and the total demand is greater than the total capacity, it is determined as a resource capacity gap. If the absolute value of the difference exceeds the capacity fluctuation threshold and the total demand is less than the total capacity, it is determined as a resource capacity redundancy, and a resource capacity gap or redundancy is generated.
[0019] Preferably, the step of obtaining the GPU resource pool adjustment decision is as follows:
[0020] Access the instance specification meta-database provided by the cloud service provider, and according to the numerical range of the resource capacity gap or redundancy, extract the list of instance models that meet the current gap or redundancy range and the associated startup time parameters and unit time holding cost parameters, and generate an instance specification-cost data set;
[0021] Based on the instance specification-cost data set, sort the candidate instance models in ascending order of startup time parameters, and at the same time sort the instance models to be released in descending order of unit time holding cost parameters. Set the priority weight ratio of startup time to holding cost to 3:1, and screen out the top three expansion instance models with the shortest startup time and the top three contraction instance models with the highest holding cost to generate a priority ranking list;
[0022] According to the priority ranking list, calculate the number of expansion instances by rounding up the gap amount divided by the single instance capacity specification, and calculate the number of contraction instances by rounding down the redundancy amount divided by the minimum release unit of the contraction instance, and generate a GPU resource pool adjustment decision including instance models, quantities, and execution timings.
[0023] Preferably, the step of obtaining the task-GPU matching degree list is as follows:
[0024] Traverse the GPU instance video memory allocation records in the video memory status evaluation result, calculate the number of consecutive free video memory blocks of each GPU instance, and count the coefficient of variation of the free block size. Define the ratio of the coefficient of variation to the total video memory capacity of the instance as the fragmentation factor. At the same time, extract the average latency time of page migration operations from the system log to generate an instance fragmentation-migration feature set;
[0025] Based on the instance fragmentation-migration feature set, calculate the matching degree between the task and the GPU instance;
[0026] For each task, filter the GPU instances based on the matching degree, and sort them in descending order to generate a task and GPU matching degree list.
[0027] Preferably, the steps for obtaining the scheduling sequence of tasks to be executed are:
[0028] Based on the task and GPU matching list, the scheduling priority of the task on the instance is calculated;
[0029] Traverse all tasks, filter out instances with the highest scheduling priority and a state change time less than or equal to the maximum tolerable delay threshold of each task, and generate a scheduling sequence of tasks to be executed in descending order of scheduling priority.
[0030] Preferably, the steps for obtaining the task distribution execution result list are:
[0031] Parse each record in the to-be-executed task scheduling sequence, extract the task ID, GPU instance ID, pre-allocated video memory capacity and planned execution timestamp, verify the current video memory availability and online status of the GPU instance ID, and generate a task-instance binding relationship table with validity identification;
[0032] Based on the task-instance binding relationship table, a video memory operation instruction message is constructed, and instructions are sent in batches according to the planned execution timestamp sequence through the resource manager's REST API to generate an instruction delivery log with a time sequence mark;
[0033] Listen to the asynchronous response message of the resource manager, match the transaction ID in the instruction delivery log, parse the operation status code in the response message, retry the failed instruction up to 3 times according to the error type, and generate a task distribution execution result list.
[0034] Preferably, the steps of obtaining the GPU scheduling execution performance record are:
[0035] Parsing the successful records in the task distribution execution result list, extracting the task ID, GPU instance ID, and occupied video memory block address, sending a process start command to the target GPU instance, and generating a task process handle mapping table;
[0036] Based on the task process handle mapping table, the video memory usage, computing power utilization and process survival time of the task are collected at intervals of 500 milliseconds, and the three are aligned by timestamp and merged to generate a resource consumption trajectory data set with time series marks;
[0037] According to the resource consumption trajectory data set, the absolute deviation between the video memory peak and the pre-allocated capacity, and the relative deviation between the actual execution time and the planned time are calculated, and the deviation value is associated with the resource consumption trajectory and written into the database to generate a GPU scheduling execution performance record containing deviation indicators and original monitoring data.
[0038] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0039] By jointly optimizing the resource pool adjustment decision and the task matching degree list, the present invention screens the instance models and quantities under the dual constraints of the instance startup time parameter and the release cost parameter, reducing the operation delay and additional costs caused by blind scaling. Based on the strong binding mechanism between the video memory pre-allocation instruction and the startup of the task process, combined with the continuous tracking of the resource consumption trajectory, the peak video memory deviation and the execution time deviation are fed back in real time to form a closed-loop scheduling optimization link. This logic quantifies the video memory continuity through the fragmentation factor and evaluates the hardware stability through the page migration overhead, improving the matching accuracy between tasks and instances and reducing task interruptions caused by video memory fragmentation or migration delay. At the same time, a scheduling sequence is generated driven by a priority sorting list. On the premise of ensuring the priority execution of high-demand tasks, the scheduling strategy is dynamically corrected using the waiting time factor to alleviate the starvation problem of long-tail tasks and achieve a balance between resource utilization and task fairness. Description of the Drawings
[0040] Figure 1 It is a step schematic diagram of the present invention. Detailed Embodiment
[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0042] Please refer to Figure 1 , the present invention provides a technical solution, a method for dynamically scheduling GPU resources in a cloud environment, including the following steps:
[0043] Monitor the running metrics of each GPU instance in the cloud environment in real time, collect GPU video memory allocation records, GPU video memory fragmentation degree values, and page migration activity counts, and generate a video memory status evaluation result by statistically predicting the total task resource requirements;
[0044] Based on the video memory status evaluation result, compare the predicted total resource requirements with the total capacity configuration of the current cloud environment GPU resource pool to obtain the resource capacity gap or redundancy. Based on the resource capacity gap or redundancy, query and evaluate the startup time parameter of the newly added GPU instance and the associated cost parameter of releasing the GPU instance, and establish a GPU resource pool adjustment decision;
[0045] Based on the video memory status evaluation result, access the list of tasks to be scheduled, extract the GPU video memory requirements of each task, generate a task-GPU matching degree list, and based on the task-GPU matching degree list, combine the time information of the instance status change in the GPU resource pool adjustment decision to select a target GPU instance for each task in the task queue to be processed, and establish a scheduling sequence of tasks to be executed;
[0046] Based on the scheduling sequence of tasks to be executed, send the task allocation information and the corresponding GPU instance identifier to the GPU resource manager in the cloud environment, issue GPU video memory pre-allocation or locking instructions according to the sequence content, obtain the task distribution execution result list, and based on the task distribution execution result list, trigger the start of the task process on the target GPU instance, continuously track the GPU resource consumption during the task running, and obtain the GPU scheduling execution efficiency record.
[0047] The steps for obtaining the video memory status evaluation result are as follows:
[0048] Call the video memory allocation monitoring interface of the GPU instance to obtain the video memory allocation record of each GPU instance within the time window, generate the video memory fragmentation degree value by calculating the standard deviation of the free area between consecutive video memory blocks, traverse the kernel log to count the number of page migrations per unit time to obtain the page migration activity count, and form a basic feature set;
[0049] Based on the basic feature set, construct an autoregressive prediction model to calculate the total task resource requirements in the future time window. The expression is:
[0050] ;
[0051] where, is the total task resource requirements, is the video memory allocation record of the previous time window, is the change rate of the video memory fragmentation degree value over time, is the current page migration activity count, is the autoregressive coefficient fitted through historical data, is the white noise error term;
[0052] Compare the total task resource requirements with the confidence interval threshold If , where is the prediction benchmark value calculated by the moving average of historical data, then it is determined that the prediction is effective and the video memory status evaluation result is generated.
[0053] Specifically, by calling the monitoring API interfaces of each GPU instance provided by the cloud platform, such as using the function interfaces provided by NVIDIA's NVML library, setting the time window to 60 seconds, periodically obtaining the detailed video memory usage information of each active GPU instance, recording the start address, size, and allocation timestamp of the currently allocated video memory blocks, forming a video memory allocation record list for each instance. At the same time, the system traverses this list to calculate the sizes of all consecutive free video memory blocks, obtaining an array of free block sizes, and calculating the standard deviation of this array. This standard deviation is the measure of video memory fragmentation. For example, if an instance has free blocks [100MB, 50MB, 150MB], with an average of 100MB and a standard deviation of , then the fragmentation measure value is 40.82. Next, the system accesses the Linux kernel logging system (such as reading / var / log / kern.log or using the journalctl command), filters the page migration event log entries related to the target GPU device, and counts the total number of page migrations that occurred within the same 60-second time window, obtaining the page migration activity count. For example, by searching for specific kernel message patterns (such as log lines containing "pagemigration" and the GPU device ID), the count is obtained as 5 migrations. Finally, the video memory allocation records of each instance collected, the calculated video memory fragmentation measure value, and the counted page migration activity count are integrated and organized in the format of [instance ID, timestamp, video memory allocation record list, video memory fragmentation measure value, page migration activity count] to form a basic feature set for subsequent predictive analysis.
[0054] Formula: , The benefit of the formula is that by combining the historical video memory usage ( ), the dynamic changes in video memory fragmentation ( ), and the intensity of page migration activities ( ), an autoregressive model is constructed to predict the total GPU task resource requirements within a short future time window. Compared with relying only on historical allocation amounts, this model takes into account the health status and usage pressure of the video memory, improving the accuracy and foresight of the prediction, and helping to adjust the resource pool size more timely.
[0055] The steps to obtain the parameter are as follows. This parameter represents the total video memory allocation amount of the GPU instance within the previous time window. The obtaining method is to query the video memory allocation record list of the previous time window (for example, the 60-second window corresponding to the In a time window, the memory allocation record of a certain GPU instance shows that a total of 5 tasks are allocated, and the occupied memory is 2.5GB, 1.8GB, 3.2GB, 0.5GB, and 2.0GB respectively. Then GB.
[0056] The steps to obtain the parameter are as follows. This parameter represents the rate of change of the memory fragmentation degree value over time, reflecting whether the memory fragmentation is tending to be severe or alleviating. The acquisition method is to first obtain the memory fragmentation degree value in the current time window (for example, getting 55.3) and the memory fragmentation degree value in the previous time window (for example, 40.82), and then calculate the difference between the two and divide it by the length of the time window (for example, 60 seconds). The calculation formula is: . For example, MB / second. This positive value indicates that the fragmentation degree is increasing.
[0057] The steps to obtain the parameter are as follows. This parameter represents the page migration activity count observed within the current time window . The acquisition method is to query the page migration activity count value in the current time window (for example, the 60-second window corresponding to the moment) recorded in the formed basic feature set. For example, within the current time window , if the page migration count is 8 times statistically obtained from the kernel log, then .
[0058] The steps to obtain the parameter are as follows. These parameters are autoregressive coefficients obtained by fitting historical data, reflecting the influence weights of each feature on the total future resource demand. The acquisition method is to collect the basic feature set data of the past N time windows (including , , , ) and the actual total resource demand of each window, construct a linear regression model, and use the least squares method to solve this multiple linear regression equation to obtain the optimal coefficient estimation value. For example, collect the data of the past 100 time windows, conduct a regression analysis, and obtain the fitting coefficients: , , . The setting basis of these coefficients is the correlation strength between each factor and the future demand in historical data. The regression analysis will automatically find the optimal weights to minimize the prediction error; these coefficients will be periodically refitted and updated as new data is added to adapt to environmental changes.
[0059] The steps to obtain the parameter are as follows. This parameter represents the random fluctuations or white noise error terms that the model fails to explain, and its expected value is 0. In actual prediction, usually the expected value 0 is taken for point estimation, or the possible range is estimated according to the residual distribution obtained from regression analysis (for example, calculating the standard deviation of the residuals ). For example, if the standard deviation of the residuals obtained from regression analysis is 0.5 GB, then it can be considered to follow a normal distribution with a mean of 0 and a standard deviation of 0.5 . In this calculation, point estimation is used, and let .
[0060] Calculation process: Substitute specific values for calculation: Given GB, MB / second = GB / second, ;
[0061] Calculate ;
[0062] Calculate the total task resource requirements :
[0063] ;
[0064] ;
[0065] ;
[0066] ;
[0067] This result indicates that, based on the video memory allocation in the previous time window, the increasing trend of current video memory fragmentation, and the activity level of page migration activities, the model predicts that in the next time window, the total resource requirements of GPU tasks will be approximately 9.635 GB. This value will be used as the key basis for evaluating whether the current resource pool capacity is sufficient.
[0068] In the formula , the steps to obtain the parameter are as follows. This parameter is the predicted value of the total task resource requirements in the future time window calculated by the autoregressive prediction model. It can be calculated that GB.
[0069] The steps to obtain the parameter are as follows. This parameter is a predicted benchmark value calculated using the moving average method based on the historical total task resource demand data. It represents the average level of recent resource demand. The calculation method is to select the actual resource demand (or the verified predicted value) of the most recent M time windows , and calculate their average value. For example, when M = 10, that is, using the data of the past 10 time windows. For example, the actual demand (in GB) of these 10 windows are: [9.8, 9.5, 10.1, 9.9, 9.7, 9.6, 9.4, 10.0, 9.8, 9.7], then GB.
[0070] The steps to obtain the parameter are as follows. This parameter is the confidence interval threshold, which is used to define the acceptable deviation range between the predicted value and the benchmark value. Its setting is usually based on the statistical characteristics of historical prediction errors. For example, it can be set as k times the standard deviation of historical prediction errors. Collect the and at multiple time points in the historical data, calculate the absolute value sequence of the difference between them, and then calculate the standard deviation of this sequence . Set the confidence level. For example, the k value corresponding to the 95% confidence level in the normal distribution is approximately 1.96. For example, if the standard deviation GB is calculated based on historical error data, and k = 2 (approximately corresponding to 95% confidence level, taking an integer for easy calculation), then . Calculation process: GB. The setting of this threshold means that if the deviation between the predicted value of the model and the recent average demand level is within 0.5 GB, this prediction is considered reliable. This threshold can be adjusted according to the requirements for prediction stability. The higher the requirements, the smaller the value.
[0071] Make a comparison and judgment: Given GB GB GB;
[0072] Calculate the absolute value of the difference:
[0073] ;
[0074] Compare the calculated absolute value of the difference with the threshold :
[0075] ;
[0076] The judgment condition holds. This result indicates that the total task resource demand predicted by the model (9.635 GB) and the predicted benchmark value calculated based on the historical data moving average The absolute difference (0.115 GB) between (9.75 GB) is less than the preset confidence interval threshold (0.5 GB). Therefore, it is determined that the prediction result obtained by the autoregressive model this time is valid and reliable. Based on this valid prediction result, the system generates a video memory status evaluation result, and this evaluation result will include the total predicted value of the task resource requirements that have been confirmed to be valid GB, as well as other possible status information (such as the current fragmentation level, page migration activity, etc.), for subsequent resource capacity analysis and task scheduling decision-making.
[0077] The steps to obtain the resource capacity gap or redundancy are as follows:
[0078] Obtain the total task resource requirements recorded in the video memory status evaluation result, and at the same time access the resource pool configuration database, extract the total capacity configuration parameters of the current GPU resource pool, and generate a resource requirement-capacity matching data set;
[0079] Based on the resource requirement-capacity matching data set, calculate the absolute difference between the total task resource requirements and the total capacity configuration, compare the difference result with the preset capacity fluctuation threshold. If the absolute value of the difference exceeds the capacity fluctuation threshold and the total requirement is greater than the total capacity, it is determined as a resource capacity gap. If the absolute value of the difference exceeds the capacity fluctuation threshold and the total requirement is less than the total capacity, it is determined as a resource capacity redundancy, and generate a resource capacity gap or redundancy.
[0080] Specifically, the system first extracts the key data items from the video memory status evaluation result generated in the previous stage, that is, the total task resource requirements for the future time window that have been verified , for example, the obtained value is GB. At the same time, the system sends a query request to the configuration management database (CMDB) or the cloud platform resource management API, accesses the resource pool configuration database that records the detailed information of the current GPU resource pool. This database stores the specifications (such as model, video memory size) and status (such as running, stopped) of all GPU instances in the resource pool. The system traverses all GPU instances with the status of "running", accumulates their video memory capacity parameters, and obtains the total capacity configuration parameters of the current GPU resource pool , for example, if there are 2 running instances in the resource pool, and each instance is configured with 16 GB of video memory, then the total capacity configuration GB. Finally, the total task resource requirements obtained and the calculated total capacity configuration Combine to form a structured record, such as in the key-value pair format {'predicted_demand_Q': 9.635, 'total_capacity_C': 32.0}, and this record is the resource demand-capacity matching dataset.
[0081] Based on the resource demand-capacity matching dataset generated in the previous step, such as containing the record {'predicted_demand_Q': 9.635, 'total_capacity_C': 32.0}, the system first calculates the total task resource demand and the total capacity configuration The absolute difference between them is calculated by directly taking the absolute value of the difference between the two, that is GB. Then, this absolute difference result (22.365 GB) is compared with a preset capacity fluctuation threshold. The setting of this capacity fluctuation threshold is to avoid frequent adjustment of the resource pool due to minor and temporary supply-demand mismatches. Its value can be determined based on the statistical analysis of historical supply-demand differences. For example, collect the absolute differences per minute in the past month , calculate the standard deviation of these difference data. For example, the standard deviation is obtained as 1.5 GB. To allow a certain normal fluctuation range, the threshold is set to 1.5 times this standard deviation, that is, the capacity fluctuation threshold = GB. The system compares the calculated absolute difference of 22.365 GB with the threshold of 2.25 GB and finds that , indicating that the difference is significant and exceeds the normal fluctuation range. It is necessary to further determine whether it is a shortage or a redundancy. At this time, check the original difference GB. Since the difference is negative (i.e., the total demand is less than the total capacity), it is determined that there is a resource capacity redundancy at present, and its magnitude is 22.365 GB. If the original difference is positive and the absolute difference exceeds the threshold, it is determined as a resource capacity shortage. Finally, the system generates the resource capacity shortage or redundancy amount according to the determination result. In this case, the generated result is "Resource capacity redundancy amount: 22.365 GB".
[0082] The steps to obtain the GPU resource pool adjustment decision are as follows:
[0083] Access the instance specification metadata database provided by the cloud service provider. According to the numerical range of the resource capacity shortage or redundancy amount, extract the list of instance models that meet the current shortage or redundancy amount range and the associated startup time parameters and unit time holding cost parameters, and generate the instance specification-cost dataset;
[0084] Based on the instance specification-cost dataset, sort the candidate instance models in ascending order of the start time parameter, and at the same time sort the instance models to be released in descending order of the holding cost per unit time. Set the priority weight ratio of the start time to the holding cost to 3:1. Screen out the top three instance models for capacity expansion with the shortest start time and the top three instance models for capacity reduction with the highest holding cost, and generate a priority ranking list.
[0085] According to the priority ranking list, calculate the number of instances for capacity expansion by rounding up the gap volume divided by the single-instance capacity specification, and calculate the number of instances for capacity reduction by rounding down the redundancy volume divided by the minimum release unit of the instance for capacity reduction, and generate a GPU resource pool adjustment decision including the instance model, quantity, and execution timing.
[0086] Specifically, the system first calls the resource query interface provided by the cloud service provider to access its instance specification metadata database, which contains the detailed technical specifications and pricing information of various GPU instances. Subsequently, based on the specific value of the resource capacity gap or redundancy determined in the previous steps, for example, there is a resource capacity redundancy of 22.365 GB, the system screens out the instances that can be considered for release from the list of currently running GPU instances in the current resource pool (this list contains information such as instance ID, instance type, start time, and cost per unit time, usually obtained from the internal resource management system or cloud platform API). The screening condition is the instance type and its associated holding cost parameter per unit time. For example, all running instances and their types and costs are screened out, such as [{'instance_id': 'i-1', 'type': 'TypeQ', 'capacity_gb': 16, 'cost_per_hour': 1.0}, {'instance_id': 'i-2', 'type': 'TypeQ', 'capacity_gb': 16, 'cost_per_hour': 1.0}], and the specification information (type, capacity, start time, holding cost) of these instances is organized and aggregated into a structured list, which is the instance specification-cost dataset.
[0087] Based on the instance specification-cost dataset generated in the previous step, for example, containing running instance information in a redundant scenario [{'instance_id': 'i-1', 'type': 'TypeQ', 'capacity_gb': 16,'startup_time_min': 8, 'cost_per_hour': 1.0}, {'instance_id': 'i-2', 'type': 'TypeQ', 'capacity_gb': 16,'startup_time_min': 8, 'cost_per_hour': 1.0}], the system executes different sorting and filtering logics according to the currently determined resource capacity redundancy (need to downsize) or resource capacity gap (need to upsize). In the current redundant scenario (downsizing), the system will sort the involved instances (or instance types, if summarizing costs by type) in descending order according to the per-unit-time holding cost parameter in the instance specification-cost dataset, and preferentially select the instance with the highest cost for release. If there are multiple instances with the same cost, a secondary sorting can be performed according to the longest running time or other strategies (such as the latest startup time). In this example, the costs of the two instances are the same, and sorting by instance ID gives ['i-2', 'i-1'] or ['i-1', 'i-2']. Then, the top three instance models (if multiple models are involved) or instance IDs with the highest holding cost (in this example, all candidate instances) are selected as candidates for release. In this case, the candidate instances are i-1 and i-2. If it is an upsizing scenario, the instances are sorted in ascending order according to the startup time parameter in the instance specification, and the top three instance models with the shortest startup time are selected as candidates. The priority weight ratio of startup time to holding cost is set to 3:1. This weight setting is based on business considerations. For example, when upsizing, obtaining computing power quickly (startup time weight is 3) is more important than cost (weight is 1), and when downsizing, preferentially releasing instances with high costs (cost weight is 3) is more important than their shutdown speed (startup time, weight is 1). However, in the specific filtering of this step, according to the description, the top three are directly selected after sorting by a single indicator (ascending by time for upsizing and descending by cost for downsizing), and finally, a priority sorting list containing the preferred instances (models for upsizing and specific instance IDs for downsizing) and their key parameters (startup time or cost) is formed.
[0088] Based on the priority sorting list generated in the previous step and the determined resource capacity gap or redundancy, the system performs specific instance quantity calculation and decision-making. In the scenario of capacity reduction where the current resource capacity redundancy is 22.365 GB, the system selects the top-ranked instance from the priority sorting list (including candidate instances i-1 and i-2, both of TypeQ, with a cost of $1.0 / hr) for processing. Select instance i-1 (or i-2), determine its capacity specification as 16 GB. The minimum release unit for the capacity reduction instance is usually the capacity of a single instance, i.e., 16 GB. Then, calculate the number of instances to be reduced. The calculation method is to divide the redundancy (22.365 GB) by the minimum release unit of the capacity reduction instance (16 GB) and round down the result, i.e., floor(22.365 / 16)=floor(1.397)=1, which means 1 instance of TypeQ needs to be released. If the scenario is a resource capacity gap, for example, the gap amount is 20 GB, and the preferred instance model for capacity expansion determined by the priority sorting list is TypeQ (with a capacity of 16 GB), then the calculation method for the number of instances to be expanded is to divide the gap amount (20 GB) by the single instance capacity specification (16 GB) and round up the result, i.e., ceil(20 / 16)=ceil(1.25)=2, indicating that 2 instances of TypeQ need to be added. Finally, the system integrates the calculation results to generate a clear decision on GPU resource pool adjustment. This decision includes the adjustment action (add or remove), instance model (when expanding capacity) or instance ID (when reducing capacity), the calculated number of instances, and the planned execution time (e.g., execute immediately or execute within a specified time window). For the current capacity reduction case, the generated decision is: {Adjustment action: Remove, Instance ID: i-1, Instance type: TypeQ, Quantity: 1, Execution timing: Immediately}.
[0089] The steps to obtain the task-GPU matching degree list are as follows:
[0090] Traverse the GPU instance video memory allocation records in the video memory status evaluation results, calculate the number of consecutive free video memory blocks for each GPU instance, and statistically calculate the coefficient of variation of the free block size , where is the standard deviation, is the mean value. Define the ratio of the coefficient of variation to the total video memory capacity of the instance as the fragmentation factor , and at the same time extract the average latency time of page migration operations from the system log to generate an instance fragmentation-migration feature set;
[0091] Based on the instance fragmentation-migration feature set, calculate the matching degree between the task and the GPU instance. The expression is:
[0092] ;
[0093] Among them, is the matching degree, is the fragmentation factor of the GPU instance , is the available video memory capacity of the instance , is the task 's video memory requirement, is the attenuation coefficient for fitting the video memory allocation efficiency, is the instance 's average page migration latency time;
[0094] For each task, filter the GPU instances that meet and , where is the preset matching degree threshold, and generate a task-GPU matching degree list sorted in descending order of the matching degree.
[0095] Specifically, in the formula: Among them, the steps for obtaining the parameter are that this parameter represents the standard deviation of the sizes of all current consecutive free video memory blocks of the GPU instance . First, obtain the total free memory by calling the video memory monitoring interface, and obtain the information of each free block by analyzing the internal video memory management data structure or using a specific tool, and obtain the list of the sizes of all consecutive free video memory blocks on the instance , for example, obtain the list [1.5GB, 0.5GB, 2.0GB]. Then, calculate the standard deviation of this list. Calculation process: First calculate the mean GB, then calculate the variance , and finally take the square root to get the standard deviation GB.
[0096] The steps for obtaining the parameter are that this parameter represents the average value of the sizes of all current consecutive free video memory blocks of the GPU instance . The acquisition method is the same as 's first step. After obtaining the list of the sizes of consecutive free video memory blocks, directly calculate its average value. For example, for the list [1.5GB, 0.5GB, 2.0GB], GB.
[0097] The steps for obtaining the parameter are that this parameter represents the total video memory capacity of the GPU instance . This parameter is a fixed specification parameter of the instance and can be obtained by querying the instance specification metadata database provided by the cloud service provider or through the monitoring interface when the instance is started. For example, for an NVIDIA A100 40GB instance, its total video memory capacity GB.
[0098] Calculation process: Substitute specific values to calculate the fragmentation factor : Given GB GB GB;
[0099] Calculate the coefficient of variation:
[0100] ;
[0101] Calculate the fragmentation factor :
[0102] ;
[0103] This step also requires obtaining the average latency time of page migration operations . The acquisition method is as follows: Continuously monitor the system kernel logs, filter the page migration event log entries specific to the GPU device (which can be filtered according to the device PCI ID or driver log identifier), record the start and end timestamps of each migration operation, calculate the duration, and then calculate the average value of all recorded migration durations within a set time window (such as the past 5 minutes). For example, if 3 page migrations are observed in the past 5 minutes, with durations of 15ms, 25ms, and 20ms respectively, then the average latency time ms.
[0104] The finally generated instance has an instance fragment - migration feature set of {'instance_id': k, 'F_k': 0.0117, 'M_k': 20}.
[0105] This result indicates that the fragmentation factor of instance is 0.0117, and the average latency of page migration is 20ms. The smaller the value, the lower the degree of fragmentation (relative to its total capacity), and the smaller the value, the higher the page migration efficiency. These two metrics together constitute the basis for evaluating whether the current state of this GPU instance is suitable for hosting new tasks.
[0106] Formula: ;
[0107] (The integral calculation result is: );
[0108] Then ;
[0109] The benefit of the formula is that it comprehensively considers multiple key factors affecting the running efficiency of tasks on a specific GPU instance and designs a comprehensive matching degree Computing model. It is inversely proportional to the fragmentation factor and the page migration latency , meaning that the lower the fragmentation and the faster the migration, the higher the matching degree. At the same time, through an integral term that takes into account the available video memory , task requirements and the decay of allocation efficiency , the possibility and efficiency of successfully allocating the required video memory are quantified, and the situation where the available space is close to the task requirements is penalized, so that the matching degree can more accurately reflect the potential for the task to start and run smoothly on this instance.
[0110] The obtaining step of the parameter is that this parameter is the fragmentation factor of the GPU instance , and its value is obtained from the instance fragment - migration feature set generated by calculation. For example, obtain .
[0111] The obtaining step of the parameter is that this parameter represents the current real - time available video memory capacity of the GPU instance . This value is obtained by calling the real - time video memory monitoring interface. For example, the currently monitored available video memory of instance is GB.
[0112] The obtaining step of the parameter is that this parameter represents the declared GPU video memory requirement of the task to be scheduled . This value is specified by the task submitter when submitting the task, or estimated through static analysis of the task code and analysis of historical operation data. For example, when task is submitted, it is declared that it needs GB of video memory.
[0113] The obtaining step of the parameter is that this parameter is the decay coefficient for fitting the video memory allocation efficiency, which reflects the speed at which the allocation success rate or efficiency decreases when the requested video memory size is close to the available video memory. This coefficient needs to be fitted through experiments or historical data analysis according to the behavior of a specific GPU model and driver / memory manager. Collect historical data and record whether the allocation is successful or the allocation time for tasks requesting different sizes of video memory under different available video memories . Analyze the trend of the decrease in the allocation success rate or the increase in the allocation time when decreases, fit an exponential decay model, and solve to obtain . For example, through non - linear regression analysis of the historical allocation data of A100 GPUs under specific loads, fit to obtain (unit 1 / GB). A higher The value indicates that the allocation efficiency drops sharply when the video memory is nearly full.
[0114] The steps to obtain the parameter are as follows. This parameter represents the GPU instance 's average page migration latency time, and its value is obtained from the instance fragmentation - migration feature set generated by calculation. For example, we get ms.
[0115] Calculation process: Substitute specific values to calculate the matching degree : Given GB GB (1 / GB) ms;
[0116] Calculate the exponential term:
[0117] ;
[0118] ;
[0119] Calculate the value of the integral part:
[0120] ;
[0121] Calculate the logarithmic term:
[0122] ;
[0123] Calculate the matching degree :
[0124] ;
[0125] ;
[0126] ;
[0127] ;
[0128] The result shows that the task (requirement 8GB) and the GPU instance (available 15GB, fragmentation factor 0.0117, migration latency 20ms) have a matching degree score of 10.43. This value itself has no absolute unit, and its relative size is used to compare the adaptability of different instances to the same task, or the adaptability of the same instance to different tasks. The higher the score, the more suitable the current state of the instance (low fragmentation, sufficient possibility of continuous space, low migration overhead) is for running the task .
[0129] Formula: and ;
[0130] The steps to obtain the parameter are that it is the video memory requirement of the task For example GB.
[0131] The steps to obtain the parameter are that it is the current available video memory capacity of the instance For example GB.
[0132] The steps to obtain the parameter are that this parameter is the matching degree of the calculated task and the instance For example .
[0133] The steps to obtain the parameter are that this parameter is a preset matching degree threshold, used to filter out instances with too low matching degrees. This threshold should be set based on the analysis of the relationship between historical task running performance and matching degree scores Collect the running data (such as completion time, GPU utilization, etc.) of a large number of tasks on different instances (with different scores), and plot a scatter plot or curve of the performance metrics changing with . Observe whether there is a value. When the score is lower than this value, the task performance significantly deteriorates or the failure rate significantly increases. Select a value near this critical point as the threshold . For example, through the analysis of historical data, it is found that when is lower than 6.0, the task running time significantly increases, so is set.
[0134] Filter the task and the instance : Given that GB GB
[0135] Check the first condition:
[0136] ;
[0137] The condition holds.
[0138] Check the second condition:
[0139] ;
[0140] The condition holds.
[0141] Since both conditions are met, the instance is considered a qualified candidate instance for the task . The system needs to repeat this process for all available GPU instances to find all instances that meet the conditions for the task . For example, there is also an instance whose GB . Then (8 <= 10) holds (7.5 > 6.0) holds, so the instance is also qualified. For example, for the instance whose GB . Then (8 <= 20) holds, but (5.5 > 6.0) does not hold, so the instance is unqualified. After finding all qualified instances, sort them in descending order according to their matching degree scores.
[0142] For example, if the qualified instances are (score 10.43) and (score 7.5), then the sorted list is [(k, 10.43), (k', 7.5)]. This result indicates that for the task , the instances and are both candidates that meet the basic video memory requirements and have good status. Among them, the instance has a higher matching degree and is a better choice. The finally generated task-GPU matching degree list provides a list of qualified GPU instances sorted in descending order of matching degree for each task to be scheduled. For example, the list for the task is [(k, 10.43), (k', 7.5)], which is used for subsequent scheduling decisions.
[0143] The steps to obtain the execution task scheduling sequence are as follows:
[0144] Based on the task-GPU matching degree list, calculate the scheduling priority of the task on the instance . The calculation formula is:
[0145] ;
[0146] where is the scheduling priority, is the matching degree between the task and the instance , is the instance of the state change time, is the system base scheduling overhead constant (calibrated through stress testing), is the task 's cumulative waiting time in the current queue, is the total number of tasks to be processed in the queue, is a very small positive number to prevent zero waiting time overflow, is the task 's cumulative waiting time in the current queue;
[0147] Traverse all tasks, filter for each task the instance with the highest scheduling priority and the maximum tolerance delay threshold of the task, and generate a to-be-executed task scheduling sequence sorted in descending order of scheduling priority.
[0148] Specifically, the formula: , the benefit of the formula is that it provides a quantitative priority calculation method for task scheduling decisions. This method not only considers the task and the instance 's technical matching degree (reflecting operation efficiency and stability), but also combines the immediate availability of the instance (reflected by the state change time and the scheduling overhead ), and introduces the waiting time of the task as a factor to measure the urgency of the task. By combining the waiting time of the task with its relative waiting situation in the entire queue to be processed (normalized by the square root sum of the waiting times of other tasks), it realizes a comprehensive consideration of the fairness and urgency of the task, enabling tasks with high priorities (high matching degree, fast instance preparation, long waiting time) to be scheduled first.
[0149] The steps to obtain the parameter are as follows. This parameter is the matching degree score between the task and the GPU instance , and its value comes from the calculation result of the previous step "task-GPU matching degree list". For example, according to the previous calculation, the matching degree between task and the instance is .
[0150] The steps to obtain the parameter are as follows. This parameter represents the state change time of the GPU instance , that is, the time remaining until the instance can accept new tasks. If the instance is currently in a stable running state and idle, then . If the instance is in the process of starting according to the resource pool adjustment decision, is the estimated remaining startup time, which can be obtained from the cloud platform API or the resource manager. For example, if the instance is currently ready, then seconds.
[0151] The steps to obtain the parameter are as follows: This parameter is the system's basic scheduling overhead constant, which represents the minimum time overhead inherent in the scheduling system to process a task allocation request and is independent of specific tasks and instances. This constant is calibrated by stress testing the scheduling system: Under different loads, send a large number of simple task scheduling requests with negligible execution time to the system, measure the end-to-end time from the request sent to the confirmation of allocation completion (excluding the actual execution time of the task), and statistically calculate the average value or P99 value of these times. For example, through stress testing, it is determined that the average time taken for the system to complete a basic scheduling operation is seconds.
[0152] The steps to obtain the parameter are as follows: This parameter is the cumulative waiting time of the task in the pending queue from the time of submission to the current moment. The scheduler records the timestamp of the task when it enters the queue and subtracts this timestamp from the current time each time the priority is calculated. For example, if task has been waiting in the queue for 120 seconds, then seconds.
[0153] The steps to obtain the parameter are as follows: This parameter represents the total number of tasks in the current pending task queue. The scheduler maintains the pending task queue in real time, and this value can be obtained directly by counting the number of tasks in the queue. For example, if there are 5 tasks waiting for scheduling in the current queue, then .
[0154] The steps to obtain the parameter are as follows: This is a very small positive constant used to prevent the situation where the denominator is zero when all task waiting times are zero (for example, when the queue is just established) or during the calculation process. Its value should be much smaller than the typical waiting time or scheduling time unit. A fixed small value can be set. For example, set seconds.
[0155] The steps to obtain the parameter are as follows: This parameter represents the cumulative waiting time of the th task in the queue. The scheduler needs to traverse all tasks in the current pending queue to obtain the cumulative waiting time of each task. For example, the waiting times of 5 tasks in the queue are as follows: Task 1( ) seconds, Task 2 seconds, Task 3 seconds, Task 4 seconds, Task 5 seconds.
[0156] Calculation process: Substitute specific values to calculate the tasks In the instance Scheduling priority : Given that seconds, seconds, seconds, , the waiting times of all tasks in the pending queue are respectively [120, 60, 180, 30, 90] seconds, seconds.
[0157] Calculate the terms related to waiting time: Calculate the total waiting time:
[0158] ;
[0159] Calculate the square root sum of waiting times:
[0160] ;
[0161] Calculate the relative waiting time factor of Task :
[0162] ;
[0163] Calculate the terms related to matching degree and instance preparation time:
[0164] ;
[0165] Calculate the final scheduling priority :
[0166] ;
[0167] ;
[0168] This result indicates that Task has a scheduling priority score of 114.29 on the GPU instance . This score comprehensively reflects the matching degree between the task and the instance, the readiness of the instance, and the waiting urgency of the task itself. The scheduler will calculate the priority scores of Task on all eligible instances and combine the priority scores of other tasks to determine the final scheduling allocation. The task-instance pairing with a higher priority score is more likely to be executed first.
[0169] The scheduling priorities of each task calculated in the previous step on its candidate GPU instances , the system then traverses all tasks to be scheduled, and for each task , from its candidate instance list (i.e., instances that satisfy and ), a final screening is performed. The screening criteria include two aspects: one is that the status change time of the instance must be less than or equal to the maximum tolerance delay threshold allowed for this task, and the other is that this instance must have the highest scheduling priority among all candidate instances that satisfy the first time constraint condition. The maximum tolerance delay threshold of a task is specified when the task is submitted or preset according to the task type (for example, the threshold for interactive tasks is low, and the threshold for batch tasks is high). It defines the longest time that a task can wait for an instance to be ready starting from the current moment. For example, the maximum tolerance delay threshold of task j is 300 seconds, the seconds of instance k, the seconds of instance k', both satisfy , if the of instance k and the of instance k', then select instance k as the target execution instance for task j because it has the highest priority. After the system determines the unique best target instance for each task to be scheduled, these triples of (task ID, target instance ID, scheduling priority) are collected. Finally, they are sorted globally in descending order according to the value of the scheduling priority to form the final task scheduling sequence to be executed.
[0170] The steps to obtain the task distribution execution result list are as follows:
[0171] Parse each record in the task scheduling sequence to be executed, extract the task ID, GPU instance ID, pre-allocated video memory capacity, and planned execution timestamp, verify the current video memory availability and online status of the GPU instance ID, and generate a task-instance binding relationship table with a validity flag;
[0172] Based on the task-instance binding relationship table, construct video memory operation instruction messages, and send the instructions in batches in the order of the planned execution timestamp through the REST API of the resource manager to generate an instruction issuance log with a timing mark;
[0173] Listen for the asynchronous response messages of the resource manager, match the transaction IDs in the instruction issuance log, parse the operation status codes in the response messages, retry failed instructions up to 3 times according to the error type, and generate a task distribution execution result list.
[0174] Specifically, the system processes each scheduling record in the task scheduling sequence to be executed in sequence. For example, when processing the record (task_j, instance_k, P_jk), it first extracts the task identifier (task ID) as 'j', the target GPU instance identifier (GPU instance ID) as 'k', and the pre-allocated video memory capacity required by the task (i.e. the demand for the task). , for example, 8GB), and determine a planned execution timestamp based on the sequence order or priority, such as setting it to be executed within a few seconds after the current time. Then, before actually issuing the binding instruction, the system performs a final verification. The system initiates a real-time query to the resource manager or cloud platform monitoring interface to obtain the latest status of instance 'k', verify whether it is online (for example, the status is 'Running' or 'Active') and operational, and check its current real-time available video memory capacity Is it still greater than or equal to the pre-allocated video memory capacity (8GB) required by task 'j'? The binding relationship between the task and the instance is considered valid only when the instance is online and the video memory meets the requirements. The system records this information (task ID, instance ID, requested video memory, and planned timestamp) together with the validity verification result (Boolean value True or False) to form a task-instance binding relationship table with validity identification.
[0175] Based on the task-instance binding relationship table with validity identifiers generated in the previous step, the system filters out all the binding records marked as valid. For each valid record, such as task 'j' being bound to instance 'k' and requesting 8GB of video memory, the system constructs a specific video memory operation instruction message accordingly. This message usually adopts JSON or XML format and contains all the parameters required for the execution operation. For example, {"transaction_id": "unique-tx-id-123", "operation": "reserve_gpu_memory", "target_instance": "k", "requester_task": "j", "memory_size_gb": 8, "lock_mode": "exclusive"}, which includes a uniquely generated transaction ID for subsequent tracking. Subsequently, the system sorts these instruction messages according to the planned execution timestamps in the records and organizes the instructions with the same or similar timestamps into batches (if the resource manager API supports batch operations). By calling the RESTful API interface provided by the resource manager (for example, sending a request to the POST / v1 / gpu / instances / {instance_id} / memory / reservations endpoint), these constructed instruction messages are sent in chronological order. For each successfully sent instruction (or batch), the system records an entry in the instruction distribution log, including the exact timestamp at the time of sending, the used transaction ID, the target instance ID, the task ID, and the complete content of the instruction message, generating an instruction distribution log with chronological marks.
[0176] The system starts a background process or service that continuously listens for asynchronous response messages sent back by the resource manager through a specified mechanism (such as a message queue, callback URL, or polling status interface). Upon receiving each response message, the system first parses the message content, extracts the transaction ID therein, and uses this transaction ID to search for the matching original instruction record in the instruction issuance log to associate the response with its corresponding task and instance. Then, it parses the operation status code indicating the operation result in the response message (for example, HTTP status code 200 indicates success, and 4xx or 5xx indicates failure) and the error message text that may be included. If the status code indicates that the operation has failed, the system will determine whether the failure is retryable based on predefined error types. For example, errors indicating that the resource is temporarily unavailable (such as status code 503) or temporary network problems are set as retryable errors, while errors indicating invalid parameters (such as status code 400) or non-existent resources (such as status code 404) are set as non-retryable errors. For retryable failed instructions, as long as the cumulative retry count has not reached the set upper limit (for example, at most 3 times), the system will schedule to resend the instruction at a later time (either at a fixed interval or using an exponential backoff strategy) and update the retry count. If the instruction is finally successfully executed (either successfully on the first try or after retries), or fails after reaching the maximum retry count, or encounters a non-retryable error, the system records this final result (including the task ID, instance ID, success or failure status, the video memory block address allocated when successful, the error details when failed, and the retry count), summarizes the final distribution processing results of all tasks, and generates a task distribution execution result list.
[0177] The steps to obtain the GPU scheduling execution performance record are as follows:
[0178] Parse the successful records in the task distribution execution result list, extract the task ID, GPU instance ID, and the occupied video memory block address, send a process start command to the target GPU instance, and generate a task process handle mapping table;
[0179] Based on the task process handle mapping table, collect the video memory occupancy, computing power utilization, and process survival time during task operation at 500 - millisecond intervals, align the three by timestamp and merge them to generate a time - series - marked resource consumption trajectory data set;
[0180] According to the resource consumption trajectory data set, calculate the absolute deviation between the peak video memory and the pre - allocated capacity, and the relative deviation between the actual execution time and the planned time. After associating the deviation values with the resource consumption trajectory, write them into the database to generate a GPU scheduling execution performance record containing deviation metrics and the original monitoring data.
[0181] Specifically, the system first parses the task distribution execution result list, filters out the records with the status of "success", and for each successful record, extracts the corresponding task ID, target GPU instance ID, and the address of the successfully pre-allocated or locked video memory block (if any) returned by the resource manager. Subsequently, the system uses this information, combined with the execution logic of the task (such as the executable file path of the task, startup parameters, required environment variables, etc.), to construct a process startup command. This command is sent to the node corresponding to the specified GPU instance ID through an appropriate remote execution mechanism (such as executing a command line after an SSH connection, calling a container orchestration system API like KubernetesJobAPI, or through a proxy service deployed on the target instance). The command needs to include necessary environment variable settings, such as specifying the specific GPU device index allocated to the task through `CUDA_VISIBLE_DEVICES`, and passing the pre-allocated video memory block address (if directly used by the application) to the task process. After successfully starting the task process, the execution environment on the target instance will return a unique handle identifying the process (such as the process ID (PID) of the operating system, container ID, or job ID). The system collects these task IDs, corresponding process handles, and instance IDs, and generates and maintains a task process handle mapping table, such as `{'task_id_j': {'process_handle': 'pid_12345', 'instance_id': 'k'},...}`.
[0182] Based on the maintenance-based task process handle mapping table, the system starts a periodic monitoring process. For each active task process handle recorded in the table, at a fixed time interval of 500 milliseconds, it collects the real-time resource consumption data of the process on the specified GPU instance through remote query or by using the monitoring agent deployed on the instance. The specific collection content includes: querying the actual GPU memory occupancy (unit such as GB) associated with the process handle by calling the NVIDIA Management Library (NVML) interface or executing the `nvidia-smi` command, and the utilization rate of the computing units of the GPU core (percentage). At the same time, it confirms whether the process is still alive by querying the process management information of the operating system (such as checking the ` / proc / [PID]` directory or using the `ps` command) and calculates its cumulative running time since startup (process survival time). The system associates the memory occupancy, computing power utilization, process survival status (such as boolean value True / False), and cumulative running time collected each time with the exact timestamp when the collection occurs, and merges the data collected at the same timestamp into one record. These records are continuously accumulated to form a complete, time-sequentially marked resource consumption trajectory data set for each task, such as `[{'timestamp': 'ts1', 'task_id': 'j', 'instance_id': 'k','memory_gb': 7.5, 'utilization_pct': 85, 'alive': True, 'runtime_sec': 60.5}, {'timestamp': 'ts2',...}]`.
[0183] After the system finishes executing a task (i.e., monitors that the process is no longer alive) or reaches the preset maximum monitoring duration, it processes the resource consumption trajectory dataset generated by the task. First, it traverses the sequence of video memory occupancy in the dataset to find its maximum value, which is the video memory peak. Then, it compares this video memory peak with the pre-allocated capacity initially requested by the task (for example, 8GB requested by task j), and calculates the absolute deviation between the two. For example, `|peak 7.8GB - pre-allocated 8GB| = 0.2GB`. Next, it determines the actual execution time of the task from the dataset, that is, the time difference between the process start time and the timestamp of the last detected alive time. At the same time, it obtains the planned execution time of the task (this planned time usually refers to the estimated execution duration, which needs to be provided by the task submitter or estimated based on historical data, for example, planned execution for 1200 seconds), and calculates the relative deviation between the actual execution time (for example, 1150 seconds) and the planned execution time. For example, `(1150 - 1200) / 1200 = -0.0417` or -4.17%. If the planned execution time is not available, this deviation is not calculated. Finally, the system associates the two performance metrics, the absolute deviation between the calculated video memory peak and the pre-allocated capacity, and the relative deviation between the actual execution time and the planned time, with the basic information of the task (task ID, instance ID, start / end time, etc.) and the complete resource consumption trajectory dataset (or a reference to its storage location), and writes them as a whole record into the specified performance database or log storage, generating a GPU scheduling execution performance record containing deviation metrics and the original monitoring data.
[0184] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for dynamically scheduling GPU resources in a cloud environment, characterized in that, It includes the following steps: Real-time monitor the running metrics of each GPU instance in the cloud environment, collect GPU video memory allocation records, GPU video memory fragmentation degree values, and page migration activity counts, and generate video memory status evaluation results by statistically predicting the total task resource requirements; Based on the video memory status evaluation results, compare the predicted total resource requirements with the total capacity configuration of the current cloud environment GPU resource pool to obtain the resource capacity gap or redundancy. Based on the resource capacity gap or redundancy, query and evaluate the startup time parameters of newly added GPU instances and the associated cost parameters of released GPU instances, and establish a GPU resource pool adjustment decision; Based on the video memory status evaluation results, access the task list to be scheduled, extract the GPU video memory requirements of each task, generate a task-GPU matching degree list, and based on the task-GPU matching degree list, combine the time information of the instance status change in the GPU resource pool adjustment decision to select a target GPU instance for each task in the task queue to be processed, and establish an execution task scheduling sequence; Based on the execution task scheduling sequence, send the task allocation information and the corresponding GPU instance identifier to the GPU resource manager of the cloud environment, issue GPU video memory pre-allocation or locking instructions according to the sequence content, obtain the task distribution execution result list, and based on the task distribution execution result list, trigger the startup of the task process on the target GPU instance, continuously track the GPU resource consumption during the task operation, and obtain the GPU scheduling execution efficiency record; The steps for obtaining the video memory status evaluation results are as follows: Call the video memory allocation monitoring interface of the GPU instance to obtain the video memory allocation records of each GPU instance within the time window, generate the video memory fragmentation degree value by calculating the standard deviation of the free areas between consecutive video memory blocks, traverse the kernel log to count the number of page migrations per unit time to obtain the page migration activity count, and form a basic feature set; Based on the basic feature set, construct an autoregressive prediction model to calculate the total task resource requirements in the future time window. The expression is: ; Among them, is the total task resource requirement, is the video memory allocation record of the previous time window, is the change rate of the video memory fragmentation degree value over time, is the current page migration activity count, is the autoregressive coefficient fitted through historical data, is the white noise error term; Compare the total task resource requirement with the confidence interval threshold If , where is the predicted benchmark value calculated by moving average based on historical data, it is determined that the prediction is effective, and a video memory state evaluation result is generated.
2. The dynamic scheduling method for GPU resources in a cloud environment according to claim 1, wherein The steps for obtaining the resource capacity gap or redundancy are as follows: Obtain the total task resource requirements recorded in the video memory status evaluation results, and at the same time access the resource pool configuration database to extract the total capacity configuration parameters of the current GPU resource pool, and generate a resource requirement-capacity matching data set; Based on the resource requirement-capacity matching data set, calculate the absolute difference between the total task resource requirements and the total capacity configuration, compare the difference result with the preset capacity fluctuation threshold. If the absolute value of the difference exceeds the capacity fluctuation threshold and the total requirement is greater than the total capacity, it is determined as a resource capacity gap. If the absolute value of the difference exceeds the capacity fluctuation threshold and the total requirement is less than the total capacity, it is determined as a resource capacity redundancy, and a resource capacity gap or redundancy is generated; 3. The dynamic scheduling method for GPU resources in a cloud environment according to claim 1, wherein The steps for obtaining the GPU resource pool adjustment decision are as follows: Accessing the instance specification metadata database provided by the cloud service provider, extracting a list of instance models that meet the current gap or redundancy range and associated startup time parameters and unit time holding cost parameters according to the numerical range of the resource capacity gap or redundancy, and generating an instance specification-cost data set; Based on the instance specification-cost data set, the candidate instance models are arranged in ascending order according to the startup time parameter, and the instance models to be released are arranged in descending order according to the unit time holding cost parameter, and the priority weight ratio of the startup time to the holding cost is set to 3:1, and the top three expansion instance models with the shortest startup time and the top three reduction instance models with the highest holding cost are screened out to generate a priority sorting list; According to the priority sorting list, the number of instances for expansion is calculated by dividing the gap by the single instance capacity specification and rounding up, and the number of instances for reduction is calculated by dividing the redundancy by the minimum release unit of the reduction instance and rounding down, and a GPU resource pool adjustment decision including instance model, quantity and execution timing is generated.
4. The dynamic scheduling method of GPU resources in a cloud environment according to claim 1, wherein The steps for obtaining the task and GPU matching list are as follows: Traversing the GPU instance memory allocation records in the memory status evaluation result, calculating the number of continuous free memory blocks of each GPU instance, counting the coefficient of variation of the free block size, defining the ratio of the coefficient of variation to the total memory capacity of the instance as a fragmentation factor, and extracting the average delay time of page migration operations from the system log to generate an instance fragmentation-migration feature set; Based on the instance fragmentation-migration feature set, the matching degree between the computing task and the GPU instance; For each task, filter the GPU instances based on the matching degree, and sort them in descending order to generate a task and GPU matching degree list.
5. The dynamic scheduling method of GPU resources in a cloud environment according to claim 1, characterized in that The steps for obtaining the scheduling sequence of tasks to be executed are: Based on the task and GPU matching list, calculate the scheduling priority of the task on the instance; Traverse all tasks, filter out instances with the highest scheduling priority and a state change time less than or equal to the maximum tolerable delay threshold of each task, and generate a scheduling sequence of tasks to be executed in descending order of scheduling priority.
6. The dynamic scheduling method of GPU resources in a cloud environment according to claim 1, characterized in that, The steps for obtaining the task distribution execution result list are as follows: Parse each record in the to-be-executed task scheduling sequence, extract the task ID, GPU instance ID, pre-allocated video memory capacity and planned execution timestamp, verify the current video memory availability and online status of the GPU instance ID, and generate a task-instance binding relationship table with validity identification; Based on the task-instance binding relationship table, a video memory operation instruction message is constructed, and instructions are sent in batches according to the planned execution timestamp sequence through the resource manager's REST API to generate an instruction delivery log with a time sequence mark; Listen to the asynchronous response message of the resource manager, match the transaction ID in the instruction delivery log, parse the operation status code in the response message, retry the failed instruction up to 3 times according to the error type, and generate a task distribution execution result list.
7. The dynamic scheduling method of GPU resources in a cloud environment according to claim 1, characterized in that, The steps for obtaining the GPU scheduling execution performance record are as follows: Parse the successful records in the task distribution execution result list, extract the task ID, GPU instance ID, and occupied video memory block address, send a process startup command to the target GPU instance, and generate a task process handle mapping table; Based on the task process handle mapping table, collect the video memory occupancy, computing power utilization, and process survival time during task execution at intervals of 500 milliseconds, align the three by timestamp and merge them to generate a resource consumption trajectory data set with time sequence tags; According to the resource consumption trajectory data set, calculate the absolute deviation between the peak video memory and the pre-allocated capacity, and the relative deviation between the actual execution time and the planned time. After associating the deviation values with the resource consumption trajectory, write them into the database to generate a GPU scheduling execution efficiency record containing deviation indicators and original monitoring data.
Citation Information
Patent Citations
Dynamic resource scheduling method for GPU (Graphics Processing Unit) cluster
CN114647515A
Method and apparatus for dynamic resource allocation of processing units
US20120079498A1