LLM model resource allocation method and device under GPU cluster based on mixed integer programming
By building a resource allocation model based on hybrid integer programming, the problem of low efficiency of GPU resource allocation in the existing technology is solved, and more efficient resource utilization and cost-effectiveness are achieved.
Patent Information
- Application Number
- CN202510474851.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the configuration and scheduling of GPU resources based on manual experience leads to a low resource utilization rate of the LLM model.
The resource allocation model is constructed based on mixed integer programming, and the solution variables are obtained through the optimization algorithm framework to obtain the global optimal target allocation plan, thereby improving resource utilization.
The utilization rate of LLM model resources under the GPU cluster is improved, costs are reduced, automated resource allocation is realized, and efficiency is improved.
Smart Images

Figure CN119988043A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method and device for allocating LLM model resources in a GPU cluster based on mixed integer programming. Background Art
[0002] With the rapid development of artificial intelligence technology, large language models (LLM models) such as Generative Pre-Trained Transformer 4 (GPT-4), Gemini, Large Language Model Meta AI 3 (Llama3), Claude, Mixtral, and DeepSeek-V3 have demonstrated their excellent performance in a wide range of practical applications. These applications cover multiple fields such as chatbots, education, and medical care, and have a profound impact on people's daily lives and work styles. However, as the capabilities of LLM models continue to improve, their service requirements have become increasingly diverse and complex. This diversity is not only reflected in the different types of requests, but also in the different requirements for computing resources.
[0003] In related technologies, users can rent cloud platform resources and deploy LLM services on the cloud platform. Specifically, the Graphic Processing Unit (GPU) resources provided by the cloud platform are used to process the service requirements of different LLM models. GPU resources are usually configured and scheduled based on manual experience to match the computing resource requirements of different LLM models. However, this manual method has low resource utilization. Summary of the invention
[0004] In view of the above problems, the present application aims to provide a method and device for allocating resources of an LLM model in a GPU cluster based on mixed integer programming, so as to improve resource utilization.
[0005] In a first aspect, the present application provides a method for allocating LLM model resources in a GPU cluster based on mixed integer programming, the method comprising:
[0006] Receive a computing resource allocation request from a user, and determine resource requirement information and a budget upper limit of at least one LLM model based on the computing resource allocation request; wherein the resource requirement information includes at least one of the following: computing requirements, memory requirements, and response time upper limit of each LLM model;
[0007] Obtain available resource information of a GPU cluster; wherein the GPU cluster includes multiple heterogeneous GPUs;
[0008] The available resource information, the resource demand information and the budget upper limit are used as inputs of a pre-constructed resource allocation model, and the decision variables of the resource allocation model are solved by a pre-prepared optimization method to obtain a target allocation solution; wherein the resource allocation model is constructed based on a mixed integer programming algorithm framework;
[0009] Computational resources are allocated to the at least one LLM model based on the target allocation scheme.
[0010] In one possible implementation, the constraints of the resource allocation model include at least one of the following: model allocation constraint, GPU utilization constraint, load balancing constraint, computing resource constraint, and budget constraint; the objective function of the resource allocation model includes at least one of the following: cost penalty item, system throughput item, and response delay penalty item.
[0011] In a possible implementation manner, the resource requirement information further includes: resource allocation relationship parameters when two different LLM models are deployed in the same GPU;
[0012] The constraint conditions of the resource allocation model also include: resource allocation constraints when two different LLM models are deployed in the same GPU.
[0013] In a possible implementation, the resource allocation constraint is determined based on the computing resources respectively allocated by the GPU to the two different LLM models and resource allocation relationship parameters.
[0014] In a possible implementation, the computing resource constraint includes at least one of the following: a computing capability constraint, a memory requirement constraint, and a response time constraint.
[0015] In a possible implementation, the available resource information includes at least one of the following: computing power, memory capacity, unit time operation cost, utilization upper limit, utilization lower limit of each GPU in the GPU cluster, and load imbalance upper limit of the GPU cluster.
[0016] In a second aspect, the present application provides a LLM model resource allocation device under a GPU cluster based on mixed integer programming, the device comprising:
[0017] A receiving unit, configured to receive a computing resource allocation request from a user, and determine resource requirement information and a budget upper limit of at least one LLM model based on the computing resource allocation request; wherein the resource requirement information includes at least one of the following: computing requirement, memory requirement, and response time upper limit of each LLM model;
[0018] An acquisition unit, used to acquire available resource information of a GPU cluster; wherein the GPU cluster includes a plurality of heterogeneous GPUs;
[0019] A decision unit, configured to use the available resource information, the resource demand information and the budget upper limit as inputs of a pre-constructed resource allocation model, solve the decision variables of the resource allocation model by a pre-constructed optimization method, and obtain a target allocation solution; wherein the resource allocation model is constructed based on a mixed integer programming algorithm framework;
[0020] An allocation unit is used to allocate computing resources to the at least one LLM model based on the target allocation scheme.
[0021] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;
[0022] The memory stores computer-executable instructions;
[0023] The processor executes the computer-executable instructions stored in the memory to implement the method in any possible implementation of the first aspect above.
[0024] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer execution instructions, and when the computer execution instructions are executed by a processor, they are used to implement the method in any possible implementation of the above-mentioned first aspect.
[0025] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the method in any possible implementation manner of the first aspect.
[0026] The present application provides a method and device for allocating resources of an LLM model under a GPU cluster based on mixed integer programming, the method comprising: receiving a computing resource allocation request from a user, and determining resource requirement information and a budget upper limit of at least one LLM model based on the computing resource allocation request; wherein the resource requirement information includes at least one of the following: computing requirements, memory requirements and response time upper limit of each LLM model; obtaining available resource information of a GPU cluster; wherein the GPU cluster includes multiple heterogeneous GPUs; using available resource information, resource requirement information and budget upper limit as inputs of a pre-constructed resource allocation model, solving the decision variables of the resource allocation model through a prepared optimization method, and obtaining a target allocation scheme; wherein the resource allocation model is constructed based on a mixed integer programming algorithm framework; and computing resources are allocated to at least one LLM model based on the target allocation scheme. The present scheme pre-constructs a resource allocation model based on a mixed integer programming algorithm framework, and then the resource allocation model can be used to obtain a globally optimal target allocation scheme, and computing resource allocation using the target allocation scheme improves resource utilization. And the present scheme also takes the budget upper limit into account, which improves cost efficiency. And the present scheme can automatically complete the allocation of computing resources, with high efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0028] Figure 1 A schematic diagram of a process flow of a LLM model resource allocation method in a GPU cluster based on mixed integer programming provided in Example 1 of the present application;
[0029] Figure 2 A flowchart of another LLM model resource allocation method in a GPU cluster based on mixed integer programming provided in Example 2 of the present application;
[0030] Figure 3 A schematic diagram of the structure of an LLM model resource allocation device in a GPU cluster based on mixed integer programming provided in Example 3 of the present application;
[0031] Figure 4 A hardware structure diagram of an electronic device provided in Example 4 of the present application.
[0032] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments.
[0033] Description of reference numerals:
[0034] 300-LLM model resource allocation device under GPU cluster based on mixed integer programming; 301-receiving unit; 302-acquisition unit; 303-decision making unit; 304-allocation unit;
[0035] 401 - processor; 402 - memory; 403 - communication interface; 404 - communication bus. DETAILED DESCRIPTION
[0036] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0037] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way. In the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more.
[0038] Users need certain computing resources during the training, deployment, and operation of LLM models, especially the training phase of LLM models requires more computing resources. Usually, users can rent GPU resources from the cloud platform to complete LLM training. Different types of LLM models have different requirements for the computing power and memory of GPU resources. Therefore, a single type of GPU resource often cannot achieve optimal resource utilization in all scenarios. For example, a compute-intensive LLM model may fully utilize the computing power of high-performance GPUs, but a memory-intensive LLM model may not be able to fully utilize the computing power of these GPUs, resulting in resource waste.
[0039] Heterogeneous GPU resources refer to the use of GPUs of different models and specifications in the same cluster, each with different computing power, memory capacity, and memory capacity. By properly configuring and scheduling these heterogeneous GPUs, the resource requirements of different LLM models can be better matched, thereby improving resource utilization and cost efficiency.
[0040] In the related art, GPU resources are usually configured and scheduled based on manual experience to match the computing resource requirements of different LLMs. However, this manual method has low resource utilization.
[0041] In order to solve the above technical problems, the embodiment of the present application provides a method and device for resource allocation of LLM model under GPU cluster based on mixed integer programming. The resource allocation model is pre-built based on the mixed integer programming algorithm framework, and then the resource allocation model can be used to obtain the global optimal target allocation scheme. The target allocation scheme is used to allocate computing resources to improve resource utilization. And this scheme also takes the budget upper limit into account, which improves cost efficiency. And this scheme can automatically complete the allocation of computing resources with high efficiency.
[0042] Figure 1 A flowchart of a method for allocating resources of an LLM model in a GPU cluster based on mixed integer programming is provided in the first embodiment of the present application. This embodiment can be applicable to the allocation scenario of computing resources. The method can be executed by an LLM model resource allocation device in a GPU cluster based on mixed integer programming, which can be implemented by software and / or hardware and specifically configured in an electronic device. Figure 1 As shown, the method includes:
[0043] Step 101, receiving a computing resource allocation request from a user, and determining resource requirement information and a budget upper limit of at least one LLM model based on the computing resource allocation request; wherein the resource requirement information includes at least one of the following: computing requirement, memory requirement and response time upper limit of each LLM model.
[0044] Specifically, a computing resource allocation request from a user may be received, wherein the computing resource allocation request from the user includes resource demand information of at least one LLM model and a budget cap corresponding to the at least one LLM model. The budget cap refers to a budget cap pre-planned by the user for leasing computing resources for the service demand of the at least one LLM model.
[0045] In implementation, this solution can simultaneously allocate computing resources to LLM models of various types and sizes.
[0046] Step 102, obtaining available resource information of a GPU cluster; wherein the GPU cluster includes a plurality of heterogeneous GPUs.
[0047] The GPU cluster is a computer cluster, each node in which is equipped with a GPU. The GPU cluster can provide computing resources for the LLM model.
[0048] Specifically, the GPU cluster includes multiple heterogeneous GPUs, which means that the GPU cluster includes multiple GPUs of different models and specifications, and each GPU has different computing power, memory capacity, and other characteristics.
[0049] In implementation, this solution can design a GPU combination solution to provide computing resources for LLM models based on resource characteristics such as computing power, memory capacity, and content bandwidth of different types of GPUs. For example, computationally intensive LLM models can be preferentially assigned to high-performance GPUs, memory-intensive LLM models can be assigned to GPUs with larger memory capacity, and cost-sensitive LLM models can be assigned to GPUs with high cost performance. The respective advantages of heterogeneous GPUs can be fully utilized to improve overall resource utilization and system performance.
[0050] The available resource information refers to information related to the computing resources that the current GPU cluster can provide. For example, the available resource information of the GPU cluster may include the computing power, memory capacity, memory bandwidth, etc. that each GPU in the GPU cluster can provide. In actual implementation, the available resource information usually changes in real time.
[0051] Step 103, using available resource information, resource demand information and budget cap as inputs of a pre-built resource allocation model, solving the decision variables of the resource allocation model through a prepared optimization method, and obtaining a target allocation plan; wherein the resource allocation model is constructed based on a mixed integer programming algorithm framework.
[0052] Specifically, based on the mixed integer programming algorithm framework, the constraints and objective functions of the resource allocation model can be pre-built. The available resource information, resource demand information and budget upper limit are input into the constraints and objective functions, and the minimum value of the objective function is searched under the premise of satisfying the constraints, so as to determine the corresponding target allocation scheme when the objective function is the minimum value.
[0053] The prepared optimization method refers to the optimization method prepared in advance. This solution does not restrict the optimization method. For example, the optimization method can be a solver or a heuristic method. The solver can use a mature commercial solver, such as the Cardinal Optimizer (COPT).
[0054] Step 104: Allocate computing resources to at least one LLM based on the target allocation scheme.
[0055] The target allocation scheme at least includes: the corresponding relationship between the LLM model and the GPU, and the computing resources allocated to the LLM model by the GPU having the corresponding relationship with the LLM model.
[0056] Among them, computing resources include computing power, memory capacity, memory bandwidth, etc.
[0057] The LLM model resource allocation method based on mixed integer programming under the GPU cluster provided in the above embodiment pre-builds a resource allocation model based on the mixed integer programming algorithm framework, and then the resource allocation model can be used to obtain the global optimal target allocation plan, and the target allocation plan is used to allocate computing resources to improve resource utilization. This solution also takes the budget upper limit into account, which improves cost efficiency. This solution can automatically complete the allocation of computing resources with high efficiency. Furthermore, the LLM model uses the appropriate computing resources provided by this solution, and its training speed is faster at the same time or machine cost.
[0058] Figure 2 A flowchart of another LLM model resource allocation method for a GPU cluster based on mixed integer programming is provided in Example 2 of this application. Figure 1 On the basis of the illustrated embodiment, an improvement is made to the LLM model resource allocation method in a GPU cluster based on mixed integer programming.
[0059] like Figure 2 As shown, a LLM model resource allocation method in a GPU cluster based on mixed integer programming may include the following steps:
[0060] Step 201, receiving a computing resource allocation request from a user, and determining resource requirement information and a budget upper limit of at least one LLM model based on the computing resource allocation request; wherein the resource requirement information includes at least one of the following: computing requirement, memory requirement and response time upper limit of each LLM model.
[0061] In an implementable manner, the resource requirement information further includes: resource allocation relationship parameters when two different LLM models are deployed in the same GPU.
[0062] Among them, the resource allocation relationship parameter can characterize the resource competition relationship when two different LLM models are deployed in the same GPU.
[0063] Step 202, obtaining available resource information of a GPU cluster; wherein the GPU cluster includes a plurality of heterogeneous GPUs.
[0064] In an implementable manner, the available resource information includes at least one of the following: computing power, memory capacity, unit time running cost, utilization upper limit, utilization lower limit of each GPU in the GPU cluster, and load imbalance upper limit of the GPU cluster.
[0065] The unit time operation cost refers to the unit time cost of renting a GPU.
[0066] Optionally, the available resource information may further include the memory bandwidth of each GPU in the GPU cluster.
[0067] During implementation, a variety of available resource information is used to make decisions and determine the target allocation plan, ensuring a high resource utilization rate.
[0068] Step 203, using available resource information, resource demand information and budget upper limit as inputs of a pre-built resource allocation model, solving the decision variables of the resource allocation model through a prepared optimization method, and obtaining a target allocation plan; wherein the resource allocation model is constructed based on a mixed integer programming algorithm framework; the constraints of the resource allocation model include at least one of the following: model allocation constraint, GPU utilization constraint, load balancing constraint, computing resource constraint, budget constraint; the objective function of the resource allocation model includes at least one of the following: cost penalty item, system performance item.
[0069] Specifically, the model assignment constraint is used to ensure that each LLM model must be assigned to at least one GPU.
[0070] The model allocation constraint can be expressed as follows:
[0071]
[0072] Among them, w represents the LLM model, W represents the set of LLM models; represents the GPU in the GPU cluster, G represents the GPU cluster; if w is assigned to , then let is 1, otherwise, let is 0.
[0073] In implementation, model allocation constraints can ensure that resources are allocated to each LLM model to prevent omissions.
[0074] Specifically, the GPU utilization constraint is used to ensure that the utilization of each GPU is within a certain range, thereby avoiding resource overload or resource idleness.
[0075] The GPU utilization constraint can be expressed as follows:
[0076]
[0077] Among them, w represents the LLM model, W represents the set of LLM models; Represents the GPU in the GPU cluster, and G represents the GPU cluster; express The lower limit of utilization rate; express The lower limit of utilization rate; if is used, then is 1, otherwise, let is 0; express The computational resources allocated to w.
[0078] Specifically, the load balancing constraint is used to ensure that the loads of the GPUs in the GPU cluster are as balanced as possible, to avoid overloading some GPUs while idling other GPUs.
[0079] The load balancing constraint can be expressed using the following formula:
[0080]
[0081] Among them, w represents the LLM model, W represents the set of LLM models; and Both represent GPUs in a GPU cluster, and G represents a GPU cluster; express The computing resources allocated to w; express The computing resources allocated to w; Indicates the upper limit of the load imbalance of the GPU cluster.
[0082] Specifically, the computing resource constraint is used to ensure that the resource requirement corresponding to the LLM model allocated to each GPU does not exceed the computing resources that the GPU can provide.
[0083] In one achievable manner, the computing resource constraint includes at least one of the following: a computing capability constraint, a memory requirement constraint, and a response time constraint.
[0084] The computing power constraint is used to ensure that the computing power required by the LLM model allocated to each GPU does not exceed the computing power that the GPU can provide.
[0085] The computing capacity constraint can be expressed using the following formula:
[0086]
[0087] Among them, w represents the LLM model, W represents the set of LLM models; Represents the GPU in the GPU cluster, and G represents the GPU cluster; represents the computational requirements of the LLM model w, The unit can be expressed as the total number of floating-point operations (FLO); if w is assigned to , then let is 1, otherwise, let is 0; express The computing power of The unit can be expressed as floating-point operations per second (FLOPS); if is used, then is 1, otherwise, let is 0.
[0088] The memory requirement constraint is used to ensure that the memory requirement of the LLM model assigned to each GPU does not exceed the memory capacity that the GPU can provide.
[0089] The memory requirement constraint can be expressed using the following formula:
[0090]
[0091] Among them, w represents the LLM model, W represents the set of LLM models; represents the GPU in the GPU cluster, G represents the GPU cluster; if w is assigned to , then let is 1, otherwise, let is 0; if is used, then is 1, otherwise, let is 0; represents the memory requirement of w; express The memory capacity of the content. The unit of memory requirement and content capacity can be GB.
[0092] The response time constraint is used to ensure that the response time of the LLM model meets the preset response time requirement. Assuming that the response time is inversely proportional to the computing resources, the response time constraint can be expressed as follows:
[0093]
[0094] Among them, w represents the LLM model, W represents the set of LLM models; represents the GPU in the GPU cluster, G represents the GPU cluster; if w is assigned to , then let is 1, otherwise, let is 0; express The computing resources allocated to w; represents the computational requirements of the LLM model w; express computing power; Indicates the upper limit of the response time of the LLM model w, and its unit can be seconds.
[0095] Specifically, the budget constraint is used to ensure that the total operating cost, i.e., the resource rental cost, does not exceed the budget.
[0096] The budget constraint can be expressed as follows:
[0097]
[0098] in, represents the GPU in the GPU cluster, G represents the GPU cluster; if is used, then is 1, otherwise, let is 0; express The unit time operation cost; Indicates the budget cap.
[0099] In an implementable manner, the constraint conditions of the resource allocation model also include: resource allocation constraints when two different LLMs are deployed in the same GPU.
[0100] During implementation, in multi-model scenarios, considering the resource competition and collaborative optimization issues among different LLM models, this solution can ensure the coordination and efficiency of resource allocation when multiple LLM models are deployed on the same GPU through resource allocation constraints. It can efficiently support the coexistence of multiple models in heterogeneous GPU resource scenarios to avoid resource competition and improve the overall performance, stability and efficiency of the system.
[0101] Optionally, the resource allocation constraint is determined based on computing resources respectively allocated by the GPU to the two different LLMs and resource allocation relationship parameters.
[0102] Specifically, the resource allocation constraint can be expressed by the following formula:
[0103]
[0104] Among them, w and All represent LLM models, and W represents the set of LLM models; Represents the GPU in the GPU cluster, and G represents the GPU cluster; express The computing resources allocated to w; express Assigned to computing resources; represents w and Deployed at the same time The resource allocation relationship parameters in the middle, It can be expressed in matrix form.
[0105] Specifically, the resource allocation constraints can be determined conveniently through the above method.
[0106] Specifically, the objective function of the resource allocation model can be expressed using the following formula:
[0107]
[0108] Among them, min means solving the minimum value; represents the cost penalty term; Indicates system performance items; is the weight coefficient of the cost penalty term, is the weight coefficient of the system performance item, and the user can set the weight coefficient as needed; w represents the LLM model, and W represents the set of LLM models; Represents the GPU in the GPU cluster, and G represents the GPU cluster; express The unit time operating cost; if is used, then is 1, otherwise, let is 0; represents the computational requirements of the LLM model w; Represents the upper limit of the response time of LLM model w.
[0109] Specifically, the objective function provided by this solution can be used for multi-objective optimization. By setting weight coefficients, the two objectives of cost minimization and performance maximization can be integrated into a weighted objective function. In this way, while meeting budget constraints, the service response time and system throughput can be improved as much as possible, achieving the best balance between cost and performance.
[0110] Specifically, the resource allocation model can be constructed conveniently through the above method.
[0111] Step 204: Allocate computing resources to at least one LLM based on the target allocation scheme.
[0112] In practice, the principle and implementation of step 204 are similar to those of step 104 and will not be described in detail.
[0113] Figure 3 This is a schematic diagram of the structure of a LLM model resource allocation device for a GPU cluster based on mixed integer programming provided in the third embodiment of the present application. The device can be in the form of software and / or hardware. Figure 3 As shown, a LLM model resource allocation device 300 in a GPU cluster based on mixed integer programming includes: a receiving unit 301, an acquiring unit 302, a decision unit 303 and an allocation unit 304.
[0114] The receiving unit 301 is used to receive a computing resource allocation request from a user, and determine resource requirement information and a budget upper limit of at least one LLM model based on the computing resource allocation request; wherein the resource requirement information includes at least one of the following: computing requirement, memory requirement, and response time upper limit of each LLM model;
[0115] The acquisition unit 302 is used to acquire available resource information of a GPU cluster; wherein the GPU cluster includes a plurality of heterogeneous GPUs;
[0116] The decision unit 303 is used to use the available resource information, resource demand information and budget upper limit as inputs of a pre-built resource allocation model, solve the decision variables of the resource allocation model through a prepared optimization method, and obtain a target allocation plan; wherein the resource allocation model is constructed based on a mixed integer programming algorithm framework;
[0117] The allocation unit 304 is configured to allocate computing resources to at least one LLM model based on the target allocation scheme.
[0118] In one implementable manner, the constraints of the resource allocation model include at least one of the following: model allocation constraint, GPU utilization constraint, load balancing constraint, computing resource constraint, and budget constraint; the objective function of the resource allocation model includes at least one of the following: cost penalty item and system performance item.
[0119] In an implementable manner, the resource requirement information further includes: resource allocation relationship parameters when two different LLM models are deployed in the same GPU;
[0120] The resource allocation model constraints also include: resource allocation constraints when two different LLM models are deployed in the same GPU.
[0121] In an implementable manner, the resource allocation constraint is determined based on the computing resources respectively allocated by the GPU to the two different LLM models and resource allocation relationship parameters.
[0122] In one achievable manner, the computing resource constraint includes at least one of the following: a computing capability constraint, a memory requirement constraint, and a response time constraint.
[0123] In an implementable manner, the available resource information includes at least one of the following: computing power, memory capacity, unit time running cost, utilization upper limit, utilization lower limit of each GPU in the GPU cluster, and load imbalance upper limit of the GPU cluster.
[0124] The LLM model resource allocation device for a GPU cluster based on mixed integer programming provided in the embodiment of the present application has the same implementation principle and technical effects as the aforementioned embodiment of the LLM model resource allocation method for a GPU cluster based on mixed integer programming. For the sake of brief description, for matters not mentioned in the device embodiment, reference may be made to the corresponding contents in the aforementioned embodiment of the LLM model resource allocation method for a GPU cluster based on mixed integer programming.
[0125] Figure 4 A hardware structure diagram of an electronic device provided in Embodiment 4 of the present application. This embodiment provides an electronic device, comprising: at least one processor 401, and a memory 402 in communication with the at least one processor 401; the memory 402 stores computer-executable instructions; the processor 401 executes the computer-executable instructions stored in the memory 402, for implementing the LLM model resource allocation method under a GPU cluster based on mixed integer programming described in any of the above embodiments.
[0126] Figure 4 The electronic device shown also includes a communication interface 403 and a communication bus 404, wherein the processor 401, the memory 402 and the communication interface 403 are connected to each other via the communication bus 404. The communication bus 404 can be divided into an address bus, a data bus, a control bus, etc. Figure 4 In the figure, only one thick line is used to represent the communication bus 404, but it does not mean that there is only one communication bus 404 or one type of communication bus 404. The processor 401 may also be called a controller, and there is no limitation on the name.
[0127] In the embodiment of the present application, the memory 402 stores instructions that can be executed by at least one processor 401. The at least one processor 401 can execute the LLM model resource allocation method for GPU cluster based on mixed integer programming discussed above by executing the instructions stored in the memory 402. The processor 401 can implement Figure 4 The functions of each module in the device shown.
[0128] Among them, the processor 401 is the control center of the device, and can use various interfaces and lines to connect the various parts of the entire control device. By running or executing instructions stored in the memory 402 and calling the data stored in the memory 402, the various functions of the device and process data, the device can be monitored as a whole.
[0129] In one possible design, the processor 401 may include one or more processing units, and the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly processes an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the modem processor may not be integrated into the processor 401. In some embodiments, the processor 401 and the memory 402 may be implemented on the same chip or on separate chips.
[0130] Processor 401 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the LLM model resource allocation method under the GPU cluster based on mixed integer programming disclosed in the embodiments of the present application, it can be directly embodied as a hardware processor execution, or it can be executed by a combination of hardware and software modules in the processor.
[0131] The memory 402 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 402 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 402 is any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 402 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.
[0132] By designing and programming the processor 401, the code corresponding to the LLM model resource allocation method under the GPU cluster based on mixed integer programming described in the above embodiment can be fixed into the chip, so that the chip can execute it when running. Figure 1 or Figure 2 The steps of the LLM model resource allocation method for a GPU cluster based on mixed integer programming in the illustrated embodiment. How to design and program the processor 401 is a technique known to those skilled in the art and will not be described in detail here.
[0133] The embodiment of the present application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by the processor, they are used to implement the LLM model resource allocation method under the GPU cluster based on mixed integer programming described in any of the previous embodiments, so they will not be described in detail here. In addition, the description of the beneficial effects of using the same method will not be repeated. For technical details not disclosed in the computer storage medium embodiment involved in the present invention, please refer to the description of the method embodiment of the present invention.
[0134] In some possible implementations, various aspects of the LLM model resource allocation method for a GPU cluster based on mixed integer programming provided in the present application can also be implemented in the form of a program product, which includes a program code. When the program product is run on an apparatus, the program code is used to enable the control device to execute the steps of the LLM model resource allocation method for a GPU cluster based on mixed integer programming according to various exemplary implementations of the present application described above in this specification.
[0135] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0136] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0137] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0139] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A LLM model resource allocation method for a GPU cluster based on mixed integer programming, characterized in that: The method comprises: Receive a computing resource allocation request from a user, and determine resource requirement information and a budget upper limit of at least one LLM model based on the computing resource allocation request; wherein the resource requirement information includes at least one of the following: computing requirements, memory requirements, and response time upper limit of each LLM model; Obtain available resource information of a GPU cluster; wherein the GPU cluster includes multiple heterogeneous GPUs; The available resource information, the resource demand information and the budget upper limit are used as inputs of a pre-constructed resource allocation model, and the decision variables of the resource allocation model are solved by a pre-prepared optimization method to obtain a target allocation solution; wherein the resource allocation model is constructed based on a mixed integer programming algorithm framework; Computational resources are allocated to the at least one LLM model based on the target allocation scheme.
2. The method according to claim 1, characterized in that: The constraint conditions of the resource allocation model include at least one of the following: model allocation constraint, GPU utilization constraint, load balancing constraint, computing resource constraint, and budget constraint; the objective function of the resource allocation model includes at least one of the following: cost penalty item and system performance item.
3. The method according to claim 2, characterized in that The resource requirement information also includes: resource allocation relationship parameters when two different LLM models are deployed in the same GPU; The constraint conditions of the resource allocation model also include: resource allocation constraints when two different LLM models are deployed in the same GPU.
4. The method according to claim 3, characterized in that The resource allocation constraint is determined based on the computing resources respectively allocated by the GPU to the two different LLM models and resource allocation relationship parameters.
5. The method according to claim 2, characterized in that: The computing resource constraint includes at least one of the following: computing capacity constraint, memory requirement constraint and response time constraint.
6. The method according to claim 1, characterized in that The available resource information includes at least one of the following: computing power, memory capacity, unit time operation cost, utilization upper limit, utilization lower limit of each GPU in the GPU cluster, and load imbalance upper limit of the GPU cluster.
7. A LLM model resource allocation device for a GPU cluster based on mixed integer programming, characterized in that: The device comprises: A receiving unit, configured to receive a computing resource allocation request from a user, and determine resource requirement information and a budget upper limit of at least one LLM model based on the computing resource allocation request; wherein the resource requirement information includes at least one of the following: computing requirement, memory requirement, and response time upper limit of each LLM model; An acquisition unit, used to acquire available resource information of a GPU cluster; wherein the GPU cluster includes a plurality of heterogeneous GPUs; A decision unit, configured to use the available resource information, the resource demand information and the budget upper limit as inputs of a pre-constructed resource allocation model, solve the decision variables of the resource allocation model by a pre-constructed optimization method, and obtain a target allocation solution; wherein the resource allocation model is constructed based on a mixed integer programming algorithm framework; An allocation unit is used to allocate computing resources to the at least one LLM model based on the target allocation scheme.
8. An electronic device, characterized in that: comprising a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the LLM model resource allocation method in a GPU cluster based on mixed integer programming as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the method for allocating LLM model resources in a GPU cluster based on mixed integer programming is implemented as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the LLM model resource allocation method in a GPU cluster based on mixed integer programming as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Method and device for optimizing reasoning resources and electronic equipment
CN118796471A
Resource scheduling method and device, processing equipment and computer readable storage medium
CN118819754A
Resource scheduling method and device for large model reasoning request
CN119311423A
Distributed training method based on end-to-end adaption, and device
US20230169351A1
Cited By
Computing power resource allocation method and device, storage medium and program product
CN121008933A
Computing resource allocation method and device, storage medium and program product
CN121008933B