Resource allocation and task scheduling method and system for heterogeneous GPU (Graphics Processing Unit) cluster
Through the resource allocation and task scheduling method for heterogeneous GPU clusters, the affinity perception algorithm and market-oriented resource scheduling algorithm are used to solve the problems of low resource allocation efficiency and high service latency in heterogeneous GPU environment, and efficient resource utilization and stable performance are achieved.
Patent Information
- Application Number
- CN202510148457.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
The existing inference resource scheduling methods fail to fully utilize the performance advantages of different models of GPUs in heterogeneous GPU environments, and it is difficult to cope with the instantaneous peak resource requirements in real-time burst edge inference scenarios, resulting in a decrease in resource allocation efficiency and an increase in service delay.
A resource allocation and task scheduling method for heterogeneous GPU clusters is proposed. By dynamically calculating the average response time, triggering a affinity perception algorithm based on resource utility, building a unit resource utility matrix, and based on market-oriented resource scheduling algorithm, a new resource allocation and task scheduling scheme is obtained to realize multi-task sharing GPU resources.
On the premise of ensuring user service quality, the overall resource utilization rate of heterogeneous GPU clusters is improved, the GPU utilization rate is improved, and the performance is relatively stable under high loads.
Smart Images

Figure CN120066784A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of inference performance optimization and heterogeneous resource scheduling, and particularly to a resource allocation and task scheduling method and system for heterogeneous GPU clusters. Background Art
[0002] With the continuous progress of communication technology and the innovation of artificial intelligence, real-time inference services close to the data source are shared among a wide user group. Users have put forward strict quality requirements for these latency-sensitive and computationally intensive services. However, many terminal devices often lack in computing and storage performance and are difficult to meet the needs of real-time inference. Therefore, migrating inference services to edge or cloud servers has become an important solution.
[0003] With the continuous upgrade of hardware devices, the computing resources in the network infrastructure show ubiquity and heterogeneity. In inference services, GPUs, as the main computing power, are distributed across multiple nodes. These GPU models differ in the number of computing cores, memory capacity, and power consumption. This enables inference services to flexibly schedule computing resources, intelligently select suitable inference models and the deployment quantity of optimized models to cope with the real-time changing traffic demands. Therefore, reasonably utilizing these heterogeneous GPU resources and improving their utilization rate have become the key to efficient inference services.
[0004] However, most of the current inference resource scheduling methods have two obvious deficiencies. First, in terms of heterogeneous GPU expansion, the applicability of these methods is still limited, and the performance advantages of different GPU models have not been fully exploited. Especially when there are significant performance preference differences between inference models and different GPU models, this limitation may lead to a decline in resource allocation efficiency, thereby affecting the overall performance of inference services. For example, some models may perform well on specific GPUs but poorly on other models, resulting in unnecessary resource waste. Second, the existing scheduling methods perform poorly in handling real-time bursty edge inference scenarios, are difficult to cope with the instantaneous peak resource demands, often cause an increase in service latency, and ultimately affect the user experience. Therefore, it is particularly important to design a new scheduling method. It not only needs to ensure the quality of user services and ensure that the response time meets user expectations, but also should have a high degree of flexibility and dynamics to achieve more efficient resource scheduling. Summary of the Invention
[0005] Aiming at the limitations of the prior art, we propose a resource allocation and task scheduling method and system for heterogeneous GPU clusters. This method makes full use of the performance advantages of heterogeneous GPUs, reasonably schedules inference models and resource allocation, and realizes the efficient utilization of computing resources on the premise of ensuring the quality of user services.
[0006] This application provides a resource allocation and task scheduling method for heterogeneous GPU clusters, including the following steps:
[0007] S1. Dynamically calculate the average response time of the system based on the current number of queued tasks, the number of running tasks, and the task arrival rate in the system; determine whether the average response time exceeds the latest service delay; if so, execute step S2;
[0008] S2. Trigger an affinity-aware algorithm based on resource utility, and construct a unit resource utility matrix of each inference model on heterogeneous GPUs;
[0009] S3. Based on the unit resource utility matrix, obtain a new resource allocation and task scheduling scheme based on a market-oriented resource scheduling algorithm;
[0010] S4. According to the new resource allocation and task scheduling scheme, each inference module implements configuration switching to achieve multi-task sharing of GPU resources, and return to step S1;
[0011] Among them, the latest service delay is the acceptable latest service delay set by the user.
[0012] Preferably, calculating the average response time of the system based on Little's theorem and determining whether it exceeds the set latest service delay includes the following steps:
[0013] (1) Estimate the average response time of the system. In the time interval [t 0 , t 1 , integrate the arrival rate λ t to calculate the average arrival rate Further, according to Little's theorem, calculate the average response time of the current system Its value is the product of the number of tasks not yet processed in the system and the average arrival rate , that is
[0014] Among them, in real-time inference services, the data source side uses filtering technology to preprocess the input data stream, extract key data units (such as key video frames, important data packets, or other key information), so as to reduce the inference load and remove redundant information. At the same time, due to the random distribution of key content in the data stream, the occurrence frequency of data units containing objects of interest (such as specific image frames, sensor events, or data packets) usually conforms to the Poisson distribution, satisfying the basic assumptions of Little's theorem in queuing theory.
[0015] (2) Compare the average response time with the latest service delay. If the response time If the latest service delay LC is exceeded, it indicates that the current resource configuration cannot meet the quality of service requirements of the task, and the affinity-aware algorithm based on resource utility is triggered.
[0016] Preferably, when the calculated average response time exceeds the set latest service delay, the affinity-aware algorithm based on resource utility is triggered. First, it is necessary to calculate the affinity of each inference model for different types of GPUs. For this purpose, an affinity-aware algorithm based on resource utility is proposed to quantify the performance preference differences of inference models for different GPUs. For each model, a unit resource utility matrix A for different types of GPUs is constructed. Each element in the matrix quantifies the performance improvement that can be brought when a specific model consumes unit resources on a certain GPU model, providing a clear quantitative basis for subsequent resource allocation optimization. The steps of triggering the affinity-aware algorithm based on resource utility and constructing the unit resource utility matrix of each inference model on heterogeneous GPUs are as follows:
[0017] S21. Collect the processing delay T and unit time throughput ω of the inference model i on the GPU j under the current resource allocation ratio x ij to obtain the resource utility ij under the current resource allocation ratio x
[0018] S22. Calculate the difference between the resource utility μ ij and the latest service delay LC and arrival rate λ, and dynamically adjust the resource allocation ratio x according to the difference ij ;
[0019] S23. Determine whether to perform resource adjustment max_iters times. If so, execute step S24; if not, return to step S21;
[0020] S24. Calculate the change rate of the resource utility μ ij in the last two times to obtain the unit resource utility a ij ;
[0021] S25. Determine the unit resource utility matrix A according to the unit resource utility a ij ;
[0022] where max_iters is the number of resource adjustment times for the inference model to execute cyclically.
[0023] Preferably, calculate the difference between the resource utility μ ij and the latest service delay LC and arrival rate λ, and dynamically adjust the resource allocation ratio proportionally
[0024] Preferably, calculate the resource utility μ of the last two times ij change rate to obtain the unit resource utility where x ij [max_iters], μ ij [max_iters] refers to the resource allocation ratio of the penultimate time and the resource utility obtained under this configuration; x ij [max_iters - 1], μ ij [max_iters - 1] refers to the resource allocation ratio of the second-to-last time and the resource utility obtained under this configuration.
[0025] Preferably, according to the unit resource utility matrix A, based on the market-oriented resource scheduling algorithm, obtain a new resource allocation and task scheduling scheme, including the following steps:
[0026] S31. Initialize the budget B held by the inference model i i , the initial price p j of GPU j ;
[0027] S32. In the current t-th iteration, according to the modeling result of the Fisher Market mechanism, determine the resource allocation ratio x j (t) of each model at the current price p ij (t);
[0028] S33. Calculate the GPU price of the next iteration where σ is the minimum adjustment unit of the price;
[0029] S34. Compare the price difference between the current t-th iteration and the (t + 1)-th iteration. If the price difference is greater than the set threshold, return to S32; otherwise, output the final resource allocation ratio x ij .
[0030] Among them, Fisher Market is one of the most basic models for resource allocation problems in economic theory. The Fisher Market mechanism is used to model the heterogeneous resource allocation scenario, where each buyer (i.e., inference task) has a preference relationship with a set of bundled goods (i.e., heterogeneous GPUs), and this relationship can be represented by a utility function. Specifically, the buyer always chooses the GPU that can maximize its utility.
[0031] In the process of resource allocation, the market model needs to determine the price p j of each GPU and the resource allocation ratio x ij allocated to each buyer. And these factors are all affected by the consumption budget B of each buyeri and their utility a for different GPUs ij ∈A. Therefore, the core objective is to maximize the obtained resource utility u under the budget B i constraint, based on the unit utility a of each purchaser for heterogeneous GPUs ij , that is: ij :
[0032] Maximize u ij = ∑ j∈[m] (a ij x ij ) ρ
[0033]
[0034] Lagrange + KKT conditions:
[0035]
[0036] It can be known from the results that when the resource allocation ratio is x ij is , the resource utility obtained by the model reaches the maximum.
[0037] Preferably, the new resource allocation and task scheduling scheme includes two steps: reloading the model and adjusting the resource allocation ratio; among them, the process of reloading the model is divided into three stages: importing the model, building an inference engine, and creating a context; among them, the CPU asynchronously builds the inference engine, and at the same time continues to run the original inference task on the GPU to ensure the efficient use of resources; when the construction of the inference engine is completed, the system loads the new model onto the GPU, and at the same time adjusts the resource allocation ratio, and completes the optimization of resource allocation and task scheduling according to the new resource allocation ratio. Through the overlapping parallel method, this method realizes the effective reuse of inference and loading time, completes the reallocation of resources, and reduces the resource waste caused by reloading.
[0038] This application also provides an effective resource allocation and task scheduling system for a heterogeneous GPU cluster, including the following modules:
[0039] Latency overrun trigger module: used to calculate the current average response time of the system, and determine whether the average response time exceeds the latest service latency; when the latest service latency is exceeded, trigger the inference performance monitoring module;
[0040] Preferably, the latency overrun trigger module calculates the average response time of the system. In the time interval [t 0 , t 1 , integrate the arrival rate λ t to calculate the average arrival rate Furthermore, according to Little's Law, calculate the average response time of the current system Its value is the number of tasks that have not been processed in the system Multiplied by the average arrival rate That is
[0041] Preferably, the latency overrun trigger module compares the average response time with the latest service latency. If the response time Exceeds the latest latency LC, it indicates that the current resource configuration cannot meet the quality of service requirements of the tasks, and triggers the affinity-aware algorithm based on resource utility
[0042] Inference performance monitoring module: Trigger the affinity-aware algorithm based on resource utility, construct the unit resource utility matrix of each inference model on heterogeneous GPUs; and feedback the unit resource utility matrix to the heterogeneous resource configuration analysis module
[0043] Preferably, the inference performance monitoring module performs the following steps: S21. Collect the processing latency T and unit time throughput ω of model i and GPU j Under the current resource allocation ratio x ij To obtain the resource utility under the current resource allocation ratio x ij ;
[0044] S22. Calculate the difference between the resource utility μ ij And the latest service latency LC and arrival rate λ, and dynamically adjust the resource allocation ratio x according to the difference ij ;
[0045] S23. Determine whether to perform resource adjustment max_iters times. If so, execute step S24; if not, return to step S21
[0046] S24. Calculate the change rate of the resource utility μ in the last two times ij To obtain the unit resource utility a ij ;
[0047] S25. According to the unit resource utility a ij Determine the unit resource utility matrix A, and feedback the unit resource utility matrix A to the heterogeneous resource configuration analysis module
[0048] Heterogeneous resource configuration analysis module: Based on the unit resource utility matrix, obtain a new resource allocation and task scheduling scheme based on the market-oriented resource scheduling algorithm, and issue the new resource allocation and task scheduling scheme to each working node
[0049] Preferably, the heterogeneous resource configuration analysis module performs the following steps: S31. Initialize the budget B held by each inference model i , and the initial price p of each type of GPU j ;
[0050] S32. In the current t-th iteration, according to the modeling result of the Fisher Market mechanism, respectively determine the resource allocation ratio x j (t) of each model at the current price p ij (t);
[0051] S33. Calculate the GPU price for the next iteration where σ is the minimum adjustment unit of the price;
[0052] S34. Compare the price difference between the current t-th iteration and the (t + 1)-th iteration. If the price difference is greater than the set threshold, return to S32; otherwise, output the final resource configuration x ij .
[0053] Dynamic batch inference module: used to implement configuration switching for each inference module according to the new resource allocation and task scheduling scheme, so as to realize the sharing of GPU resources by multiple tasks.
[0054] Preferably, the dynamic batch inference module reloads the model and adjusts the resource allocation ratio; among them, the process of reloading the model is divided into three stages: importing the model, building the inference engine, and creating the context; among them, the CPU asynchronously builds the inference engine, and at the same time continues to keep the original inference task running on the GPU to ensure the efficient use of resources; when the construction of the inference engine is completed, the system loads the new model onto the GPU, adjusts the resource allocation ratio at the same time, and completes the optimization of resource allocation and task scheduling according to the new resource allocation ratio.
[0055] This application deeply perceives the affinity between the model and the hardware, and fully explores the performance differences of different task models on heterogeneous GPUs. At the same time, by introducing the market configuration mechanism, a resource allocation framework based on the dynamic balance of supply and demand is established, realizing the optimization of resource configuration efficiency. On the premise of ensuring the task service quality, this method effectively improves the utilization rate of the overall resources of the GPU cluster.
[0056] Experiments prove that after adopting this method, the efficient resource allocation and task scheduling optimization method and system for model affinity perception for heterogeneous GPU clusters can improve the GPU utilization rate by 1.3 - 2.2 times compared with the original system while ensuring the single-task service quality. In addition, the system performs relatively stably under high load. These experimental results show that this method effectively utilizes the affinity difference of the inference model for heterogeneous GPUs, realizes a higher cost-effective resource configuration, and further improves the utilization rate of GPU resources.
[0057] This application has the characteristics of high resource utilization rate and strong adaptability, and can be widely applied to various heterogeneous resource allocation scenarios.
[0058] Through the overlapping parallel method, this application realizes the effective reuse of inference and loading time, completes resource reallocation while reducing resource waste caused by reloading. Brief Description of the Drawings
[0059] Figure 1 It is a flowchart of a resource allocation and task scheduling method for a heterogeneous GPU cluster according to this application;
[0060] Figure 2 It is a schematic diagram of the affinity-aware algorithm process based on resource utility provided by this application;
[0061] Figure 3 It is a schematic diagram of the process of obtaining a new resource allocation and task scheduling scheme based on the market-oriented resource scheduling algorithm provided by this application;
[0062] Figure 4 It is a schematic diagram of the structure of a resource allocation and task scheduling system for a heterogeneous GPU cluster provided by this application. Detailed Embodiments
[0063] This application discloses an efficient resource allocation and task scheduling optimization method and system for model affinity perception in a heterogeneous GPU cluster. This application has the characteristics of high resource utilization rate and strong adaptability, and can be widely applied to various heterogeneous resource allocation scenarios.
[0064] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the specific embodiments will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.
[0065] In one embodiment, as Figure 1 It is a flowchart of an efficient resource allocation and task scheduling method for model affinity perception in a heterogeneous GPU cluster according to this application, as shown in the figure:
[0066] Step 100: Dynamically calculate the average response time of the system using Little's theorem based on the current number of queued tasks, the number of running tasks, and the task arrival rate in the system; determine whether the average response time exceeds the latest service delay; if so, execute Step 101;
[0067] Step 101: Trigger the affinity awareness algorithm based on resource utility, and construct the unit resource utility matrix of each inference model on heterogeneous GPUs;
[0068] Step 102: Based on the unit resource utility matrix and the market-oriented resource scheduling algorithm, obtain a new resource allocation and task scheduling scheme;
[0069] Step 103: According to the new resource allocation and task scheduling scheme, each inference module performs configuration switching to realize multi-task sharing of GPU resources, and return to Step 100.
[0070] In one embodiment, as Figure 2 shown, the affinity awareness algorithm based on resource utility includes the following steps:
[0071] Step 200: Collect the performance under the current resource configuration, and calculate the resource utility of the model;
[0072] Step 201: Calculate the difference between the performance and the user requirement standard, and dynamically adjust the resource allocation proportionally;
[0073] Step 202: Determine whether all rounds of resource adjustment are completed. If so, execute Step 203; if not, return to Step 200;
[0074] Step 203: Calculate the resource utility change rate under the last two similar resource adjustments, and quantitatively evaluate the unit resource utility obtained by the model;
[0075] Step 204: Determine the final unit resource utility matrix according to the unit resource utility.
[0076] Among them, all inference models need to perform resource adjustment for a fixed number of times in a loop.
[0077] In one embodiment, the affinity awareness algorithm based on resource utility includes the following steps:
[0078] Step 2000: Collect the processing delay T and the unit time throughput ω of model i on GPU j under the current resource configuration x ij to obtain the resource utility under the current resource allocation proportion x ij ;
[0079] Step 2001: Calculate the difference between the resource utility μ ij and the latest service delay LC and arrival rate θ, and dynamically adjust the resource allocation proportion x according to the difference ij ;
[0080] Step 2002: Determine whether to perform resource adjustment for max_iters times. If so, execute Step 2003; if not, return to Step 2000.
[0081] Step 2003: Calculate the resource utility change rate μ for the last two times and obtain the unit resource utility a. ij ij ;
[0082] Step 2004: Determine the unit resource utility matrix A according to the unit resource utility a. ij
[0083] Among them, all inference models need to perform resource adjustment for a fixed number of times in a loop.
[0084] In one embodiment, as Figure 3 shown, according to the unit resource utility matrix, based on the market-oriented resource scheduling algorithm, obtain a new resource allocation and task scheduling scheme, including the following steps:
[0085] Step 300: Obtain the constructed unit resource utility matrix, uniformly initialize the budgets held by each inference model, and the initial prices of each type of GPU.
[0086] Step 301: In the current iteration, according to the modeling results of the Fisher Market mechanism, respectively determine the resource allocation ratios of each model at the current price.
[0087] Step 302: Calculate the GPU price for the next iteration according to the resource allocation situation.
[0088] Step 303: Compare the price difference between the next iteration and the current iteration, and determine whether the price difference is greater than the set threshold. If so, return to Step 301; otherwise, enter Step 304.
[0089] Step 304: Output the final resource configuration.
[0090] In one embodiment, according to the unit resource utility matrix A, based on the market-oriented resource scheduling algorithm, obtain a new resource allocation and task scheduling scheme, including the following steps:
[0091] Step 3000: Initialize the budget B held by each inference model i , and the initial price p j of each type of GPU;
[0092] Step 3001: In the current t-th iteration, according to the modeling results of the Fisher Market mechanism, respectively determine the resource allocation ratio x j (t) of each model at the current price p ij (t).
[0093] Step 3002, calculate the GPU price for the next iteration where σ is the minimum adjustment unit of the price;
[0094] Step 3003, compare the price difference between the current t-th iteration and the (t + 1)-th iteration. If the price difference is greater than the set threshold, return to Step 3001; otherwise, go to Step 3004;
[0095] Step 3004, output the final resource allocation ratio x ij 。
[0096] In one embodiment, as Figure 4 shown, an efficient resource allocation and task scheduling optimization system for model affinity awareness for heterogeneous GPU clusters is provided:
[0097] Latency overrun trigger module 41: used to calculate the current average response time of the system, and determine whether the average response time exceeds the latest service latency; when the latest service latency is exceeded, trigger the inference performance monitoring module;
[0098] Inference performance monitoring module 42: trigger the affinity awareness algorithm based on resource utility, construct the unit resource utility matrix of each inference model on heterogeneous GPUs; and feedback the unit resource utility matrix to the heterogeneous resource configuration analysis module;
[0099] Heterogeneous resource configuration analysis module 43: based on the unit resource utility matrix, obtain a new resource allocation and task scheduling scheme based on the market-oriented resource scheduling algorithm, and send the new resource allocation and task scheduling scheme to each working node;
[0100] Dynamic batch inference module 44: used to implement configuration switching for each inference module according to the new resource allocation and task scheduling scheme, and achieve multi-task sharing of GPU resources.
[0101] The embodiments of the present application disclose a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above method embodiments, including: calculating the average response time based on Little's Law according to the current load status of the system and the data stream arrival rate, and determining whether it exceeds the set latest service delay. If it exceeds, trigger resource reallocation to ensure the quality of service. The resource reallocation process includes the following steps: constructing the unit resource utility matrix of the inference model through an affinity-aware algorithm based on resource utility, measuring the performance preference of the model for heterogeneous GPUs, and calculating a new resource allocation and task scheduling method based on the Fisher Market mechanism to adjust the model deployment and resource allocation ratio.
[0102] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0103] Finally, it should be noted that although the embodiments of the present application mainly show specific cases in the display and design parts and the software involved is in the form of a system. However, each component involved in this system is still protected by patents, including but not limited to modules such as resource monitors, task schedulers, and their codes. Therefore, any modifications, improvements, and uses of the implementation of the present invention should be covered by the content of the claims and the specification of the present invention.
Claims
1. A resource allocation and task scheduling method for a heterogeneous GPU cluster, characterized in that: The method comprises the following steps: S1. Based on the number of currently queued tasks, the number of running tasks and the task arrival rate of the system, dynamically calculate the average response time of the system; determine whether the average response time exceeds the latest service delay; if so, execute step S2; S2, trigger the resource utility-based affinity perception algorithm to build a unit resource utility matrix for each inference model on the heterogeneous GPU; S3. According to the unit resource utility matrix and based on a market-based resource scheduling algorithm, a new resource allocation and task scheduling scheme is obtained; S4, according to the new resource allocation and task scheduling scheme, each inference module implements configuration switching to achieve multi-task sharing of GPU resources, and returns to step S1; The latest service delay is an acceptable latest service delay set by a user.
2. The resource allocation and task scheduling method according to claim 1, characterized in that: The average response time of the system described in step S1 is dynamically calculated using Little's theorem, including the following steps: obtaining the average number of tasks in the system within the time interval t0 to t1 Including running tasks and queued tasks; and the average arrival rate of tasks Get the average response time 3. The resource allocation and task scheduling method according to claim 2, characterized in that: The resource utility-based affinity perception algorithm described in step S2 constructs a unit resource utility matrix for each inference model on a heterogeneous GPU, including the following steps: S21, acquisition inference model i, GPU j In the current resource allocation ratio x ij The processing delay T and unit time throughput ω under the current resource configuration x are obtained. ij Resource utility under S22, computing resource utility μ ij The difference between the latest service delay LC and the arrival rate λ, and the resource allocation ratio x is dynamically adjusted according to the difference ij ; S23, determine whether to perform max_iters resource adjustments, if yes, execute step S24; if no, return to step S21; S24, calculate the last two resource utilities μ ij Change rate, obtain unit resource utility a ij ; S25, according to the unit resource utility a ij , determine the unit resource utility matrix A; Among them, max_iters is the number of resource adjustments executed in the inference model loop.
4. The resource allocation and task scheduling method according to claim 3, characterized in that: Computational resource utility μ ij The difference between the latest service delay LC and the arrival rate λ, in proportion Dynamically adjust resource allocation ratio 5. The resource allocation and task scheduling method according to claim 4, characterized in that: Calculate the last two resource utilities μ ij Change rate, obtain unit resource utility Among them, x ij [max_iters], μ ij [max_iters] refers to the penultimate resource allocation ratio and the resource utility obtained under this configuration; x ij [max_iters-1], μ ij [max_iters-1] refers to the penultimate resource allocation ratio and the resource utility achieved under this configuration.
6. The resource allocation and task scheduling method according to claim 5, characterized in that: According to the unit resource utility matrix A, based on the market-based resource scheduling algorithm, a new resource allocation and task scheduling scheme is obtained, including the following steps: S31. Initialize the budget B held by inference model i i , GPU j The initial price p j ; S32. In the current iteration t, based on the modeling results of the FisherMarket mechanism, determine the inference model i at the current price p j (t) resource allocation ratio x ij (t); S33, GPU for calculating the next iteration j price Where σ is the minimum adjustment unit of price; S34. Compare the price difference between the current iteration t and the iteration t+1. If the price difference is greater than the set threshold, return to S32; otherwise, output the final resource configuration x ij ; Among them, the Fisher Market mechanism takes into account the price p of the j-th model GPU j And GPU resource allocation ratio x ij Depends on consumption budget B i and the unit resource utility a ij ∈A; the core goal is to i Under the constraints, based on the unit resource utility a of each buyer for heterogeneous GPUs ij , maximize the resource utility u obtained ij .
7. The resource allocation and task scheduling method according to claim 6, characterized in that: The resource utility u obtained by maximizing ij : Maximize u ij =∑ j∈[m] (a ij x ij ) ρ Lagrangian + KKT conditions: When the resource allocation ratio x ij for When the resource utility u obtained by the model ij Reach the maximum.
8. The resource allocation and task scheduling method according to claim 7, characterized in that: The new resource allocation and task scheduling scheme includes two steps: reloading the model and adjusting the resource allocation ratio. The model reloading process is divided into three stages: importing the model, building the inference engine, and creating the context. The CPU asynchronously builds the inference engine while continuing to run the original inference tasks on the GPU to ensure efficient use of resources. After the inference engine is built, the system loads the new model onto the GPU and completes the resource allocation and task scheduling optimization according to the new resource allocation ratio.
9. An efficient resource allocation and task scheduling system for heterogeneous GPU clusters according to any one of claims 1 to 8, characterized in that: include: Delay overlimit trigger module: used to calculate the current average response time of the system and determine whether the average response time exceeds the latest service delay; When the latest service delay is exceeded, the inference performance monitoring module is triggered; Inference performance monitoring module: triggers the resource utility-based affinity perception algorithm to construct a unit resource utility matrix for each inference model on a heterogeneous GPU; and feeds the unit resource utility matrix back to the heterogeneous resource configuration analysis module; Heterogeneous resource configuration analysis module: According to the unit resource utility matrix and based on the market-oriented resource scheduling algorithm, a new resource allocation and task scheduling scheme is obtained, and the new resource allocation and task scheduling scheme is sent to each working node; Dynamic batch inference module: used to implement configuration switching of each inference module according to the new resource allocation and task scheduling scheme, so as to realize multi-task sharing of GPU resources.
Citation Information
Cited By
Computing task allocation method based on end-cloud fusion, computer equipment and medium
CN120723488A
Software performance load test method and system for large-scale video conference
CN120872493A
Task scheduling method and device and electronic equipment
CN122220117A