Computing power resource allocation method and device, storage medium and program product

By constructing a global resource view and optimization model, the optimal resource allocation scheme is automatically generated, which solves the problem of low efficiency in computing resource management in existing technologies, and realizes efficient and accurate resource allocation and task execution, adapting to diverse task requirements.

CN121008933AActive Publication Date: 2025-11-25BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511535918.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2025-11-25
Estimated Expiration
2045-10-27

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively manage and allocate computing resources in enterprise-level inference services that involve multiple models, multiple tenants, and long-term online operations, especially on orchestration platforms such as Kubernetes. This results in inefficient resource management and task scheduling, making it difficult to handle complex scenarios involving large models.

Method used

By constructing a global resource view, multiple resource allocation candidate schemes are generated, and the optimal resource allocation scheme is automatically generated using an optimization model and preset orchestration templates. The scheme is then solved using a constraint programming-Boolean satisfiability solver, achieving efficient and rational utilization of computing resources.

Benefits of technology

It improves the accuracy and efficiency of resource allocation, reduces manual configuration errors, can quickly adapt to the memory requirements and parallel strategies of different models, ensures the smooth execution of inference tasks, and reduces operation and maintenance costs and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121008933A_ABST
    Figure CN121008933A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a computing power resource allocation method and device, a storage medium and a program product. The method comprises the following steps: collecting current running state information of computing equipment in each node in a computing power cluster, and forming a global resource view according to the current running state information; generating a plurality of feasible resource allocation candidate schemes based on the global resource view, the memory demand of the target model and a parallel strategy, and establishing an optimization model based on the feasible resource allocation candidate schemes; solving the optimization model to obtain an optimal resource allocation candidate scheme; arranging the optimal resource allocation candidate schemes into a deployment configuration list based on a preset arrangement template; and based on the deployment configuration list, executing an inference task of the target model. According to the method, the global resource view is constructed, the optimal resource allocation scheme is automatically generated, and the reasoning task is executed in combination with the preset arrangement template, so that reasonable utilization of computing resources is realized, manual configuration errors and time are reduced, the execution efficiency is improved, and diversified task requirements can be flexibly met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, and particularly relates to a computing resource allocation method and device, a storage medium and a program product. BACKGROUND

[0002] In today's technological development, cluster resource allocation and task deployment management are crucial, and existing technical solutions mostly use manual configuration or simple script automation for management. With the rapid popularization of large models (LLM) in search, dialogue, code generation and other scenarios, enterprise-level inference services present the characteristics of multi-model, multi-tenant and long-period online. The current mainstream is to deploy services on Kubernetes and other orchestration platforms, and use Device Plugin to expose heterogeneous resources such as GPUs and ASICs to workloads.

[0003] However, manual configuration or simple script automation is difficult to cope with the resource dynamic allocation and task efficient deployment requirements in complex scenarios such as multi-model, multi-tenant and long-period online brought by large models. Although deploying services on Kubernetes and other platforms can achieve certain resource management and task scheduling, as large models become more popular, the shortcomings in resource awareness, high-throughput scheduling and optimization support for AI native workloads gradually emerge. SUMMARY

[0004] Therefore, the embodiments of the present disclosure provide a computing resource allocation method and device, a storage medium and a program product, which can build a global resource view, automatically generate an optimal resource allocation scheme and execute an inference task combined with a preset orchestration template, realize the rational use of computing resources, reduce manual configuration errors and time, improve execution efficiency, and flexibly cope with diversified task requirements.

[0005] In a first aspect, the embodiments of the present disclosure provide a computing resource allocation method, which adopts the following technical solution: Collect current running state information of computing devices in each node in a computing cluster, and form a global resource view based on the current running state information; Generate a plurality of feasible resource allocation candidate schemes based on the global resource view, memory requirements of a target model and a parallel strategy; Establish an optimization model based on all resource allocation candidate schemes; Solve the optimization model to obtain an optimal resource allocation candidate scheme; Based on a preset orchestration template, arrange the optimal resource allocation candidate scheme into a deployment configuration list; Based on the deployment configuration list, execute an inference task of the target model.

[0006] Optionally, the running state information of each node and the computing device in the node in the collection computing cluster is collected, a global resource view is constructed based on the running state information, and the global resource view includes: Real-time monitoring of the running state information of the computing device in each node in the computing cluster; When receiving an inference task of a target model, a control container group in the computing cluster uniformly pulls a lightweight snapshot within a preset collection window to obtain the current running state information of the computing device in each node; The current running state information is arranged into a global resource view in the form of a table.

[0007] Optionally, based on the global resource view, the memory requirement of the target model, and a parallel strategy, a plurality of feasible resource allocation candidate schemes are generated, and the method includes: Based on the global resource view, the memory requirement of the target model, and a parallel strategy, a plurality of feasible resource allocation candidate schemes are enumerated, and variable parameters of the resource allocation candidate schemes are obtained; The global variable parameters are configured, and the variable parameters of each resource allocation candidate scheme are bound to the global variable parameters.

[0008] Optionally, the optimization model includes a target function and a constraint condition. The constraint condition includes a global consistency constraint condition and a linear inequality constraint condition.

[0009] Optionally, the expression of the target function is: In the formula, represents a weight coefficient; represents a selection variable of the i-th node, when x i = 0, represents that the i-th node is not selected, and when x i = 1, represents that the i-th node is selected; represents available video memory of the j-th computing device; represents a selection variable of the j-th computing device, when x j = 0, represents that the j-th computing device is not selected, and when x j = 1, represents that the j-th computing device is selected; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; represents video memory that needs to be allocated to the j-th computing device according to the i-th resource allocation candidate scheme; a total number of actual required computing devices of the resource allocation candidate scheme; denotes the selection variable of the th resource allocation candidate scheme, denotes that the th resource allocation candidate scheme is not selected, and denotes that the th resource allocation candidate scheme is selected.

[0010] Optionally, an expression of the global consistency constraint condition is as follows: In the formula, denotes a total number of computing devices, ; denotes a total number of required computing devices globally; An expression of the linear inequality constraint condition is as follows: In the formula, denotes an actual allocated GPU memory amount of the computing device, denotes a constant.

[0011] Optionally, the solving of the optimization model to obtain the optimal resource allocation candidate scheme comprises: calling a constraint programming-Boolean satisfiability solver to solve the optimization model to obtain the optimal resource allocation candidate scheme.

[0012] In a second aspect, the embodiments of the present disclosure further provide a computing power resource allocation system, which adopts the following technical scheme: A collection module is configured to collect current running state information of computing devices in each node in a computing power cluster, and construct a global resource view based on the current running state information; A generation module is configured to generate a plurality of feasible resource allocation candidate schemes based on the global resource view, memory requirements of a target model, and a parallel strategy; An establishment module is configured to establish an optimization model based on all the resource allocation candidate schemes; An acquisition module is configured to solve the optimization model to obtain an optimal resource allocation candidate scheme; A collation module is configured to collate the optimal resource allocation candidate scheme into a deployment configuration list based on a preset orchestration template; An execution module is configured to execute an inference task of the target model based on the deployment configuration list.

[0013] In a third aspect, the embodiments of the present disclosure further provide a computer device, which adopts the following technical scheme: The computer device comprises:​ at least one processor; and a memory connected with the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the computing resource allocation method described in any one of the above.

[0014] In a fourth aspect, the embodiments of the present disclosure further provide a computer readable storage medium storing computer instructions for causing a computer to perform the computing resource allocation method described in any one of the above.

[0015] In a fifth aspect, the embodiments of the present disclosure further provide a computer program product comprising computer programs / instructions, which, when executed by a processor, implement the steps of the method described in any one of the above.

[0016] The computing resource allocation method provided by the embodiments of the present disclosure provides comprehensive and accurate basic information for subsequent resource allocation by collecting the current running state information of the computing devices in each node in the computing cluster and constructing a global resource view based on the information. By constructing the global resource view, the system can clearly understand the overall condition of the computing resources in the cluster, and in combination with the memory requirement and parallel strategy of the target model, the computing resources in the cluster can be more reasonably utilized, and resource waste can be avoided. Because only in the case of clear understanding of the resources, scientific allocation can be made according to the actual demand to prevent the situation of idle resources or over-allocation.

[0017] Based on the global resource view, the memory requirement and the parallel strategy of the target model, a plurality of feasible resource allocation candidate schemes are generated, an optimization model is established, and then the optimal resource allocation candidate scheme is obtained by solving. This series of steps realizes the process of automatically generating the optimal resource allocation scheme. In the traditional resource allocation, manual configuration not only consumes time, but also is prone to errors, while the present method reduces the time and errors of manual configuration through the automatic way, and improves the execution efficiency of the reasoning task. The system can quickly filter out the optimal scheme from a plurality of candidate schemes according to the preset rules and algorithms, so that the resource allocation is more accurate and efficient.

[0018] Based on the preset orchestration template, the optimal resource allocation candidate scheme is sorted into a deployment configuration list, and the inference task of the target model is executed based on this, which embodies the method's ability to quickly adapt to different model memory requirements and parallel strategies, and flexibly adjust resource allocation schemes. Different target models may have different memory requirements and parallel strategies. This method can dynamically generate appropriate deployment configuration lists according to these changes, thereby better responding to diverse task requirements. Whether facing simple models or complex models, this method can quickly respond, reasonably allocate resources, and ensure the smooth execution of inference tasks.

[0019] The above description is only a summary of the technical solutions of the present disclosure. In order to more clearly understand the technical means of the present disclosure, the contents of the specification can be implemented, and in order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the following preferred embodiments are described in detail below, and the accompanying drawings are described as follows. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0021] Figure 1 The flowchart of the computing power resource allocation method provided by the embodiments of the present disclosure is shown in the figure. Figure 2 The flowchart of the global resource view construction method provided by the embodiments of the present disclosure is shown in the figure. Figure 3 The flowchart of the resource allocation candidate scheme acquisition method provided by the embodiments of the present disclosure is shown in the figure. Figure 4 The principle block diagram of the computing power resource allocation system provided by the embodiments of the present disclosure is shown in the figure. Figure 5 The structural diagram of a computer device provided by the embodiments of the present disclosure is shown in the figure. DETAILED DESCRIPTION

[0022] The embodiments of the present disclosure will be described in detail below with reference to the drawings.

[0023] It should be apparent that the following description illustrates by way of example only a number of possible embodiments of the present disclosure. Those skilled in the art will readily understand other advantages and benefits of the present disclosure from the description that follows, without departing from the scope of the present disclosure. Obviously, the described embodiments are only a part of embodiments of the present disclosure, and are not all-inclusive of all embodiments. The present disclosure can also be implemented or applied by other different specific embodiments, and the details in the specification can be modified or changed based on different views and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present disclosure.

[0024] It should be noted that the various aspects of the embodiments described below are within the scope of the appended claims. It should be apparent that the aspects described herein can be embodied in a wide variety of forms and that any specific structure and / or function described herein is merely illustrative. Based on the teachings provided herein one skilled in the art will appreciate that one aspect described herein can be implemented independently of any other aspects and that the various aspects described herein can be combined in various ways. For example, an apparatus can be implemented or a method can be practiced using any number of the aspects set forth herein. In addition, such an apparatus can be implemented or such a method can be practiced using other structure and / or functionality in addition to or other than one or more of the aspects set forth herein.

[0025] It should also be noted that the drawings included in the following description are only schematic and that actual implementations can differ from those of the illustrations. It should also be noted that the drawings are provided only for purpose of illustration and description and that they are not to be used to exclude any implementation deemed to fall within the scope of the present disclosure.

[0026] In addition, in the following description, specific details are provided to thoroughly understand examples. However, one of ordinary skill in the art will understand that the described aspects can be practiced without these specific details.

[0027] Referring to Figure 1 The present disclosure provides a computing resource allocation method, comprising the following steps: S1: Collecting current running state information of computing devices in each node in the computing cluster, and constructing a global resource view based on the current running state information; S2: Generating a plurality of feasible resource allocation candidate schemes based on the global resource view, memory requirements of the target model, and parallel strategies; S3: Establishing an optimization model based on all resource allocation candidate schemes; S4: solving the optimization model to obtain an optimal resource allocation candidate scheme; S5: based on a preset orchestration template, arranging the optimal resource allocation candidate scheme into a deployment configuration list; S6: based on the deployment configuration list, performing an inference task of the target model.

[0028] The computing power resource allocation method provided by the present disclosure provides comprehensive and accurate basic information for subsequent resource allocation by collecting the current running state information of the computing devices in each node in the computing power cluster and forming a global resource view based on the same. By constructing the global resource view, the system can clearly understand the overall condition of the computing resources in the cluster, and in combination with the memory requirement and parallel strategy of the target model, the computing resources in the cluster can be more reasonably utilized, and resource waste can be avoided. Because only in the case of clear understanding of the resources, scientific allocation can be made according to the actual demand to prevent the situation of idle or excessive allocation of resources.

[0029] Based on the global resource view, the memory requirement and the parallel strategy of the target model, a plurality of feasible resource allocation candidate schemes are generated, an optimization model is established, and then the optimal resource allocation candidate scheme is obtained by solving. This series of steps realizes the process of automatically generating the optimal resource allocation scheme. In the traditional resource allocation, manual configuration is not only time-consuming, but also prone to errors. However, the present method reduces the time and errors of manual configuration and improves the execution efficiency of the inference task through automation. The system can quickly filter out the optimal scheme from a plurality of candidate schemes according to the preset rules and algorithms, so that the resource allocation is more accurate and efficient.

[0030] Based on the preset orchestration template, the optimal resource allocation candidate scheme is arranged into a deployment configuration list, and based on this, the inference task of the target model is performed, which shows that the method can quickly adapt to different model memory requirements and parallel strategies and flexibly adjust the resource allocation scheme. Different target models may have different memory requirements and parallel strategies, and the present method can dynamically generate appropriate deployment configuration lists according to these changes, so as to better cope with diversified task requirements. Whether it is a simple model or a complex model, the method can quickly respond and reasonably allocate resources to ensure the smooth execution of the inference task.

[0031] In S1, refer to Figure 2 The flowchart of the global resource view construction method shown in the figure comprises the following steps: S11: real-time monitoring of the running state information of the computing devices in each node in the computing power cluster; S12: When receiving the inference task of the target model, the control container group in the computing power cluster pulls a lightweight snapshot within a preset collection window to obtain the current running state information of the computing devices in each node; S13: The current running state information is arranged into a global resource view in table form.

[0032] In the above steps, the computing devices include GPUs, ASICs, TPUs, etc., and the present scheme is mainly aimed at resource allocation of GPUs, but can also be applied to other computing devices such as ASICs, TPUs, etc. On each node (that is, a computing device server, usually a GPU server), NVDIA DCGM (or dcgm-exporter issued by the computing device Operator) is deployed and constantly runs. NVDIA DCGM stands for NVDIA Data Center Computing Device Manager, which is a tool for managing and monitoring computing devices in a data center environment. The dcgm-exporter is a component related to NVDIA DCGM and can be issued by the computing device Operator. The main function of the dcgm-exporter is to expose the computing device running state indicators collected by DCGM in Prometheus format. The main function of these tools is to continuously maintain the computing device running state indicators, like a "little guard", constantly paying attention to the running of the computing devices in the node, and preparing for subsequent accurate resource information acquisition.

[0033] When receiving an inference task of a target model (such as a neural network, a Bayesian network or the like), that is, before resource allocation, the control container group in the computing cluster uniformly pulls a lightweight snapshot in a preset collection window, wherein the control container group refers to a control Pod, and specifically, an Inventory-Snapshotter control Pod can be selected, which is a basic deployable and manageable computing unit in Kubernetes (kube container orchestration engine, K8s for short). The control container group specifically accesses the HTTP index endpoint of the dcgm-exporter to obtain the UUID, total memory and available memory of each computing device, and associates the node name of the computing device. This step is based on continuous monitoring on the host side. Since the tools have been continuously maintaining the computing device state indicators, the control container group can directly obtain the required information from the HTTP index endpoint provided by these tools, and through one-time snapshot collection, the key resource information of all computing devices at a certain moment can be obtained, providing data support for subsequent generation of the resource table. The collected data is summarized to generate a temporary resource table with computing devices as the granularity. The core fields of the table only include the computing device gpu_id, node, mem_total and mem_free, which represent the unique identifier of the computing device, the node name, the total memory and the available memory, respectively. This step is based on the snapshot data collected in the foregoing, and the collected data is scattered. Through summarization and arrangement, it is converted into a structured temporary resource table for subsequent use. The temporary resource table is the global resource view. The method adopts a minimal implementation scheme of "host-side resident monitoring + centralized one-time snapshot", generates a temporary resource table with computing devices as the granularity, and in this process, there is no need for topology inference or time sequence consistency control. Moreover, with the global resource view, the information can be directly used for subsequent resource allocation candidate scheme generation and linear constraint modeling, avoiding complex topology inference and time sequence consistency control, and simplifying the resource allocation process.

[0034] In S2, referring to Figure 3 The flowchart of the resource allocation candidate scheme acquisition method is shown. Based on the global resource view, the memory requirement of the target model and the parallel strategy, multiple feasible resource allocation candidate schemes are generated. S21: Based on the global resource view, the memory requirement of the target model and the parallel strategy, multiple feasible resource allocation candidate schemes are enumerated, and the variable parameters of the resource allocation candidate schemes are obtained. S22: Configure global variable parameters, and bind the variable parameters of each resource allocation candidate scheme with the global variable parameters.

[0035] In S21, if the user interface receives the memory requirement of the target model, it is directly adopted; if it does not receive the memory requirement, it analyzes the target model in detail to determine the total amount of memory required at runtime, which can be estimated by the architecture of the model, the number of parameters, and the data size during training or inference. For deep learning models, the memory occupied by intermediate calculation results and gradient information also needs to be considered, and the tools provided by the model framework or manual calculation can be used to obtain the accurate memory requirement value. Similarly, if the user interface receives the parallel strategy of the target model, it is directly adopted; if it does not receive the parallel strategy, according to the characteristics of the target model and the system resource situation, a suitable parallel strategy is selected, and the parallel strategy includes tensor parallel and pipeline parallel. These two parallel strategies belong to the category of model parallel, and they are both ways of splitting the model to achieve parallel computing. Tensor parallel is to split the tensors in the model, while pipeline parallel is to split the layers of the model. According to the size, computational complexity and data distribution of the model, etc. to determine which parallel strategy to use. Based on the global resource view obtained before, the memory requirement of the target model and the parallel strategy, all possible resource allocation combinations, i.e. resource allocation candidate schemes, are generated. For example, if the tensor parallel strategy is used, the resource allocation candidate scheme can be to split the large tensors in the model by rows or columns in different proportions and allocate them to different numbers and combinations of computing devices; if the pipeline parallel strategy is used, the resource allocation candidate scheme can be to divide the different layers of the model in different ways to computing devices; tensor parallel and pipeline parallel can be used together to form a hybrid parallel strategy. In the enumeration process, it is necessary to ensure that each resource allocation candidate scheme meets the memory requirement of the target model and the resource constraints of the system.

[0036] Based on a pre-defined symbol and variable system, variable parameters for resource allocation candidate schemes are set. These variable parameters include model deployment method, tensor parallelism, pipeline parallelism, and the total number of computing devices actually required. The model deployment method includes single-node single-computing-device, single-node multi-computing-device, and multi-node multi-computing-device. Single-node single-computing-device, also known as single-machine single-card, refers to using only one computing device on a single node for computational tasks. Single-node multi-computing-device, also known as single-machine multi-card, refers to using multiple computing devices simultaneously on a single node to collaboratively complete computational tasks. Multi-node multi-computing-device, also known as multi-machine multi-card, refers to using computing devices on multiple nodes simultaneously to collaboratively complete computational tasks. Tensor parallelism refers to the number of computing devices or partitions used when large tensors in a model are split into rows, columns, or other components and distributed to different computing devices for parallel computation in a tensor parallelism strategy. It reflects the granularity and scale of tensor parallel computation. Pipeline parallelism refers to the number of pipeline stages used when different layers of a model are divided into different computing devices to form a pipeline for parallel computation in a pipeline parallelism strategy. It reflects the degree of stage division and parallelization level of pipeline parallel computation. The variable parameters of each resource allocation candidate scheme can be represented by a quadruple, which is... ,in, Indicates the first The model deployment method for each resource allocation candidate scheme; Indicates the first Tensor parallelism of resource allocation candidate schemes; Indicates the first Pipeline parallelism of each resource allocation candidate scheme; Indicates the first The actual total number of computing devices required for each resource allocation candidate scheme; , This represents the set of candidate resource allocation schemes, and can also be regarded as the total number of candidate resource allocation schemes.

[0037] The variable parameters for each resource allocation candidate scheme must satisfy the following rules: If the model deployment method is single-node single computing device, then tensor parallelism, pipeline parallelism, and the actual total number of computing devices required are all set to 1; if the model deployment method is single-node multi-computing device, then tensor parallelism is set to 1, and the actual total number of computing devices required is equal to the value of pipeline parallelism; if the model deployment method is multi-node multi-computing device, then the actual total number of computing devices required is equal to the product of the value of tensor parallelism and the value of pipeline parallelism. This rule can be expressed by the following formula: In the formula, Indicates the first The deployment method for each resource allocation candidate scheme is a single node and a single computing device; Indicates the first The deployment method for each resource allocation candidate scheme is a single node with multiple computing devices; Indicates the first The deployment mode of the resource allocation candidate scheme is multi-node and multi-computing device.

[0038] In S22, global variable parameters include the total number of computing devices required globally, the actual allocated GPU memory per computing device, and deployment type switches. The deployment type switches include a first type switch, a second type switch, and a third type switch, corresponding to single-node single-computing-device, single-node multiple-computing-device, and multi-node multiple-computing-device scenarios, respectively. To facilitate binding the variable parameters of each resource allocation candidate scheme with the global variable parameters, it is also necessary to obtain the GPU memory allocation per computing device when deploying the model according to the resource allocation candidate scheme. The calculation formula is as follows: In the formula, Indicates according to the first When deploying a model with multiple resource allocation candidate schemes, the amount of video memory usage needs to be evenly distributed across each computing device. This represents the floor function; This represents the estimated total video memory required for the target model.

[0039] The formula for binding the variable parameters of resource allocation candidate schemes to global variable parameters is as follows: In the formula, This indicates the total number of devices that need to be calculated globally; Indicates the first The selection variables for each resource allocation candidate scheme ,when When, it indicates the first When no resource allocation candidate was selected, When, it indicates the first One resource allocation candidate was selected; This indicates the actual allocated video memory of the computing device; This indicates the first type of switch, which corresponds to a single node and a single computing device. This indicates the second type of switch, which corresponds to a single-node multi-computing device. This indicates a third type of switch, which corresponds to multiple nodes and multiple computing devices. , The values of the three type switches are bound by the resource allocation candidate scheme, because there is only one model deployment mode of single-node single-computing device, single-node multi-computing device or multi-node multi-computing device for one resource allocation candidate scheme, and therefore, for the deployment type switch, only one of which is equal to 1 and the rest are 0; denotes the summation of the conditions that satisfy ; denotes the summation of the conditions that satisfy ; denotes the summation of the conditions that satisfy ; denotes the summation of the conditions that satisfy ; denotes the summation of the conditions that satisfy

[0040] The method is adapted to the requirements of the solver for linear constraints, and through the "resource allocation candidate scheme discretization" strategy, the nonlinearity of the ratio of the total estimated video memory required by the target model to the actual total number of computing devices is converted into a constant parameter input, so as to eliminate the nonlinearity brought by the "total video memory / card number" in the original problem. First, a plurality of candidate configurations including deployment type, TP (tensor parallelism), PP (pipeline parallelism), total card number and the like are enumerated in the discrete parallelism and deployment paradigm space, and the actual video memory allocated to each computing device is taken as a constant bound to the candidate to construct the input of the optimization model, which realizes the pure linear expression of the problem feasibility and the objective function, and lays a foundation for subsequent screening and fast solving. The actual total number of computing devices required and the video memory occupancy on each computing device allocated according to the resource allocation candidate scheme are pre-calculated for each resource allocation candidate scheme; then, in the preset solver, the linear binding of the global variable is completed with the unique activation of the resource allocation candidate scheme selection variable, so as to depict the total card number and the single-card video memory demand in a pure linear / Boolean form, avoid the subsequent solving difficulty caused by nonlinearity, and ensure seamless connection with the preset variable system.

[0041] In S3, the optimization model includes an objective function and constraint conditions, and the constraint conditions include global consistency constraint conditions and linear inequality constraint conditions. The expression of the objective function is: In the formula, w represents a weight coefficient; represents a selection variable of the i-th node, , when 0 < i < n, represents that the i-th node is not selected, and when 1 < i < n, represents that the i-th node is selected; , when 0 < i < n, represents that the i-th node is not selected, and when 1 < i < n, represents that the i-th node is selected; , when 0 < i < n, represents that the i-th node is not selected, and when 1 < i < n, represents that the i-th node is selected; , when 0 < i < n, represents that the i-th node is not selected, and when 1 < i < n, represents that the i-th node is selected; , when 0 < i < n, represents that the i-th node is not selected, and when 1 < i < n, represents that the i-th node is selected; , when 0 < i < n, represents that the i-th node is not selected, and when 1 < i < n, represents that the i-th node is selected; , when 0 < i < n, represents that the i-th node is not selected, and when 1 < i < n, represents that the i-th node is selected; , when 0 < i < n, represents that the i-th node is not selected, and when 1 < i < n, represents that the i-th node is selected; available memory of the computing device; represents the i-th computing device is selected, represents the i-th computing device is not selected, represents the i-th computing device is selected. The expression of the global consistency constraint is:

[0042] wherein, represents the total number of computing devices, In terms of global consistency, it is required that the number of selected computing devices is strictly equal to the total number of cards required by the candidate. This constraint ensures that the solver will not produce an inexecutable solution of "over / under selection", which is the basis for all subsequent feasibility and target computation.

[0043] The expression of the linear inequality constraint is: represents a constant, which is a large enough constant obtained by the Big-M () method, used to "relax" the constraint (make it automatically true) when , and to force the remaining memory on the computing device to be no less than the required memory per card when . With the constant precomputed in the candidate discretization stage, the entire constraint remains linear, so it can be efficiently processed by solvers such as CP-SAT.

[0044] ​​​​​The method adopts a constraint programming (CP) method to establish an optimization model containing decision variables and linear constraints. Constraint programming is a mathematical modeling and solving technique based on constraint satisfaction, which is particularly suitable for handling complex combinatorial optimization problems. The method models the problem in the form of "variables-constraints-objectives" and automatically searches for the optimal solution that meets the conditions. The optimization target adopts a lexicographic two-level strategy: the first-level target is to minimize the node occupancy, which is the primary target, and it prioritizes controlling cross-node communication and scheduling overhead; on the basis of meeting the first-level optimization, the second-level target is to minimize the waste of video memory (i.e., maximize the utilization of video memory), and linear indicators such as cross-node communication overhead can be extended and introduced, which is specifically defined as "the difference between the total available video memory of the selected GPU and the actual video memory required by the model". This strategy uniformly supports three deployment paradigms: single-machine single-card, single-machine multi-card, and multi-machine multi-card, and through the explicitly modeled node selection variable, it ensures that the "each node card consistency" constraint is met. As can be seen, the objective function is composed of the function of the first-level target and the function of the second-level target. The expression of the first-level target function is: The expression of the second-level target function is: For ease of engineering implementation, the lexicographic optimization can be approximated as a single-solution weighted model: by assigning a large enough weight coefficient to the first-level target, its numerical range is significantly larger than the maximum value of the second-level target, which can simulate the "first node number, then video memory waste" lexicographic effect in one optimization.

[0045] The consistency of the number of nodes and computing devices is modeled through the "node triggering mechanism + fixed number of cards per node" to uniformly support single-machine and multi-machine deployment paradigms. If the deployment type is single-machine paradigm (i.e. or ), the constraints of the optimization model ensure that only one node is enabled, and the number of computing devices allocated on this node is equal to the total number of cards in the single-machine mode in the candidate configuration (when , represent 1 or , respectively), and the number of cards on the remaining nodes is forced to be 0. The specific form is as follows: In the formula, represents the node number of the th computing device.

[0046] If the deployment type is multi-machine paradigm ( ), the number of enabled nodes is required to be exactly equal to the pipeline parallelism , and the number of allocated computing devices on each enabled node is required to be exactly equal to the tensor parallelism . The specific form is as follows: The method can combine multiple rounds of optimization into a single round of solving without changing the constraint conditions, which not only maintains the optimization effect but also improves the solving efficiency and is easier to deploy in actual systems.

[0047] Based on the above, after obtaining the computing power resource snapshot and the resource allocation candidate scheme, the two are integrated and converted into a linear model (i.e., an optimization model) that can be directly processed by a solver. The linear model follows the decision variables and constants defined in the foregoing, and explicitly binds each attribute of the candidate combination to a global variable, ensuring that the "total number of computing devices actually needed", "the actual allocated memory amount of the computing device", and "the deployment type" have consistent and unique input sources in the optimization model. In terms of constraint system construction, the CP (Constraint Programming) model mainly follows and implements the following three principles: global consistency, memory capacity feasibility, and consistency of node and computing device usage. Through the method, an optimization model containing decision variables and linear constraints can be constructed.

[0048] In S4, a constraint programming-Boolean satisfiability solver (CP-SAT solver for short) is called to solve the optimization model and obtain an optimal resource allocation candidate scheme. CP stands for Constraint Programming, SAT stands for Boolean Satisfiability Problem, and the CP-SAT solver combines the techniques of constraint programming and Boolean satisfiability problem solving to solve complex constraint satisfaction and optimization problems. The CP-SAT solver used in this method can efficiently handle linear constraints between Boolean variables and integer variables, and with the help of a conflict-driven search strategy and a cutting plane technique, it still maintains good solving performance in large-scale problems.

[0049] By calling the CP-SAT solver to search for the optimal device subset that meets the global target, a deployment list and parameter configuration (such as TP, PP, and device mapping relationship) suitable for mainstream inference frameworks can be automatically generated. Finally, the allocation result is sent to the orchestration system of the computing power cluster and the resource state is updated, realizing automatic and interpretable resource allocation. The results returned by the CP-SAT solver and related runtime parameters specifically include: the selected optimal resource allocation candidate scheme (including ), a set of assigned computing devices , a set of enabled nodes , and derived values such as mirror address, model path, service port, environment variable, instance replica number, and release policy; meanwhile, deployment parameters such as mirror address, model path, service port, environment variable, instance replica number, and release policy.

[0050] In constraint programming modeling, the symbol and variable system includes the following three categories: indices and sets, which are used to define the search space of the solution (computing device, node, and resource allocation candidate scheme, etc.), the variable parameters of these indices and sets include , (node set, ), , (tensor parallelism candidate set) and (pipe parallelism candidate set); input parameters are known quantities determined by inventory snapshots and parallel strategies, which are used to limit the feasible region, these input parameters include , , , , and ={0: single node single computing device, 1: single node multiple computing devices, 2: multiple node multiple computing devices} (deployment paradigm enumeration), wherein and can be tailored according to cluster capabilities, for example , ; decision variables depict the choices and switches (card selection, node usage, deployment paradigm, and candidate activation) that the solver needs to decide, the values of which will be determined under subsequent linear constraints and objectives, these decision variables include , , , , , , and , wherein means a positive integer, and for decision variables are all unique activation modes.

[0051] In S5 and S6, the final output of the CP-SAT solver is a standard deployment description file that can be directly issued to the orchestration system. The standard deployment description file is a deployment configuration list that is sorted based on the preset orchestration template and the optimal resource allocation candidate scheme, which includes two types of configuration information: one is the reasoning framework dedicated parameter, such as the required --tensor-parallel-size, --pipeline-parallel-size and device mapping relationship of vLLM or SGLang; the other is the cluster orchestration resource list, such as the Kubernetes Pod description or Helm Values configuration. The orchestration system automatically deploys the large model reasoning service through the deployment configuration list, thereby guiding the system to perform the reasoning task of the target model.

[0052] With the rapid popularization of large models in search, dialogue, code generation and other scenarios, enterprise-level reasoning services present the characteristics of multi-model, multi-tenant and long-period online. The current mainstream approach is to deploy services on orchestration platforms such as Kubernetes, expose heterogeneous resources through Device Plugin, use "whole card direct access" or virtualization isolation methods such as vGPU / MIG, combine reasoning framework configuration interfaces to complete deployment, and support monitoring and alarm systems on the platform side. After going online, the service stability is maintained by increasing or decreasing instances or stopping service release. To reduce idling, some platforms introduce GPU sharing / isolation and scheduling strategies at the container layer, and develop a general technology stack of "resource virtualization + scheduling algorithm". However, the current resource virtualization and containerized operation and maintenance scheme has obvious deficiencies in automatic algorithm arrangement for large model distributed reasoning. First, there is a lack of global optimization scheme that considers parallel strategy, cluster topology and resource availability uniformly, and device selection depends on experience or heuristic scripts, which can easily lead to increased memory fragmentation and increased cross-node bandwidth pressure. Second, the deployment scheme feasibility checking method is simple and lacks a unified mechanism, requiring multiple trial and error. Third, multi-objective collaborative optimization relies on fixed weights or heuristic strategies, making it difficult to optimize multiple objectives simultaneously. Fourth, the "resource allocation once, release only when the service terminates" operation and maintenance mode can exacerbate global resource fragmentation and waste. Fifth, the current process lacks an end-to-end automatic closed loop, resulting in low deployment delivery efficiency and high operation and maintenance cost.

[0053] To solve the above problems, the disclosure provides an automatic computing resource allocation system based on constraint programming for large model distributed inference. The system first constructs a global resource view with computing devices as the granularity through a one-time resource snapshot, shielding redundant information and real-time fluctuations. Then, the resource allocation candidate scheme is discretized and the required video memory of each card is pre-calculated, so that the constraint programming model can be solved in a linear range. Finally, based on the unified CP-SAT constraint model and the lexicographic optimization strategy, the number of occupied nodes is minimized first, and then the video memory waste is reduced. The deployment parameters and orchestration templates are automatically generated to realize the end-to-end automatic closed loop of "data collection-candidate generation-constraint modeling-solution-execution backfill". The system can realize global automatic optimization of distributed inference resources without changing the existing orchestration and inference framework, significantly reduce the cross-node communication overhead, improve the running stability, suppress the video memory fragmentation, improve the overall resource utilization, reduce the manual intervention and trial and error cost, shorten the service online time, and have good universality and scalability. For example, in the actual application of a certain telecommunications operating company, the system makes the service throughput more stable, the cross-node communication less, the deployment time significantly shortened, and the overall GPU utilization and resource reuse degree significantly improved in the multi-model and multi-tenant concurrent scene.

[0054] Reference Figure 4 The disclosure provides a computing resource allocation system, comprising: The acquisition module 101 is configured to acquire the current running state information of the computing devices in each node in the computing cluster, and construct a global resource view based on the current running state information. The generation module 102 is configured to generate a plurality of feasible resource allocation candidate schemes based on the global resource view, the memory requirement of the target model and the parallel strategy. The establishment module 103 is configured to establish an optimization model based on all the resource allocation candidate schemes. The acquisition module 104 is configured to solve the optimization model and obtain the optimal resource allocation candidate scheme. The arrangement module 105 is configured to arrange the optimal resource allocation candidate scheme into a deployment configuration list based on a preset orchestration template. The execution module 106 is configured to execute the inference task of the target model based on the deployment configuration list.

[0055] The various variations and specific examples of the computing resource allocation method provided above are also applicable to the computing resource allocation system provided by the disclosure. Through the foregoing detailed description of the computing resource allocation method, those skilled in the art can clearly understand the implementation method of the computing resource allocation system. For the sake of brevity of the specification, it will not be described here in detail.

[0056] A computer device according to an embodiment of the present disclosure includes a memory and a processor. The memory is configured to store non-transitory computer readable instructions. Specifically, the memory can include one or more computer program products, which can include various forms of computer readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory, etc. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.

[0057] The processor can be a central processing unit (CPU) or other form of processing unit having data processing and / or instruction execution capabilities, and can control other components in the computer device to perform desired functions. In one embodiment of the present disclosure, the processor is configured to execute the computer readable instructions stored in the memory, so that the computer device performs all or part of the steps of the computing resource allocation method according to the embodiments of the present disclosure.

[0058] Those skilled in the art will understand that, in order to solve the technical problem of how to obtain a good user experience effect, the present embodiment can also include well-known structures such as a communication bus, an interface, etc., which should also be included in the protection scope of the present disclosure.

[0059] As Figure 5 A structural schematic diagram of a computer device according to an embodiment of the present disclosure is shown. It shows a structural schematic diagram of a computer device suitable for use to implement the computer device in the embodiments of the present disclosure. Figure 5 The computer device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.

[0060] As Figure 5 As shown, the computer device can include a processor (such as a central processor, a graphics processor, etc.) that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) or loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the computer device are also stored. The processor, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0061] Generally, the following devices can be connected to the I / O interface: input devices including, for example, sensors or visual information acquisition devices, etc.; output devices including, for example, display screens, etc.; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. The communication devices can allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or by wire to exchange data. Although Figure 5Computer devices with various devices are shown, but it is understood that not all of the shown devices are required to be implemented or present. More or fewer devices can alternatively be implemented or present.

[0062] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the computing resource allocation method of the embodiments of the present disclosure are performed.

[0063] Detailed descriptions of the present embodiments can refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0064] The computer-readable storage medium according to the embodiments of the present disclosure has non-transitory computer-readable instructions stored thereon. When the non-transitory computer-readable instructions are run by a processor, all or part of the steps of the computing resource allocation method of the embodiments of the present disclosure described above are performed.

[0065] The computer-readable storage medium described above includes, but is not limited to, optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or mobile hard disk), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).

[0066] Detailed descriptions of the present embodiments can refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0067] The basic principles of the present disclosure are described above in combination with specific embodiments, but it should be noted that the advantages, advantages, effects, etc. mentioned in the present disclosure are only examples and are not limiting, and these advantages, advantages, effects, etc. cannot be considered as the must-have of each embodiment of the present disclosure. In addition, the above specific details of the disclosure are only for the purpose of example and for the purpose of understanding, and are not limiting, and the above details do not limit the present disclosure to the above specific details.

[0068] In this disclosure, relational terms such as first and second and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. The block diagram of the devices, apparatus, equipment, systems referred to in this disclosure is merely illustrative and not intended to imply the necessity or arrangement of the connections, arrangement, configuration as shown in the block diagram. As will be appreciated by those skilled in the art, the devices, apparatus, equipment, systems can be connected, arranged, configured in any manner. The words comprising, including, having and the like are to be open ended. As used in this document, the conjunction "or" is to be interpreted in the inclusive sense, i.e. as meaning one or the other, or both. As used in this document, the words "and" and "or" are to be interpreted as having the meaning indicated in the phrase "and / or". As used in this document, the word "such as" is to be interpreted as meaning "such as, but not limited to". As used in this document, the word "for example" is to be interpreted as meaning "by way of example, not by way of limitation".

[0069] Also, as used in this document, the word "or" in the cases used to introduce list items is to be interpreted in the exclusive sense, i.e. as meaning one or the other, but not both. In addition, the phrase "example of" does not mean an example of the preferred or only example, and the phrases "for example" and "such as" do not mean that a list of following items is an exhaustive list.

[0070] It is also important to note that the systems and methods of the present disclosure can be embodied in a variety of forms including, but not limited to, a data processor, a computer program product, a computer, one or more components of a computer, software, and combinations of the same. As used in this document, the term "data processor" encompasses one or more processors capable of manipulating data according to instructions, such as a microprocessor, microcontroller, central processing unit (CPU), digital signal processor (DSP), application specific integrated circuit (ASIC), field programmable gate array (FPGA), or any other circuit or combination of circuits capable of processing data.

[0071] Various changes, modifications and improvements in the herein described technologies can be made within the teachings of the technology, particularly in view of the foregoing descriptions, which are to be considered as examples only and not as limiting the scope of the technology. Accordingly, the disclosures of the present technology are intended to be illustrative, but not limiting, of the scope of the technology, which is set forth with particularity in the following claims. What is claimed is:

[0072] The previous description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0073] The foregoing description has been presented for the purposes of illustration and description. Furthermore, the description is not intended to limit the embodiments of the disclosure to the forms disclosed herein. Although the various example aspects and embodiments have been described herein with regard to particular aspects and embodiments, those skilled in the art will recognize that certain modifications, changes, substitutions, additions and sub-combinations can be made without departing from the spirit of the disclosure.

Claims

1. A method for allocating computing resources, characterized in that, include: Collect the current operating status information of computing devices in each node of the computing power cluster, and construct a global resource view based on the current operating status information; Based on the global resource view, the memory requirements of the target model, and the parallel strategy, several feasible resource allocation candidate schemes are generated. An optimization model is established based on all candidate resource allocation schemes. Solve the optimization model to obtain the optimal resource allocation candidate scheme; Based on the preset orchestration template, the optimal resource allocation candidate schemes are organized into a deployment configuration list; Based on the deployment configuration list, perform inference tasks for the target model.

2. The computing power resource allocation method according to claim 1, characterized in that, The system collects operational status information of each node and computing device within the computing power cluster, and constructs a global resource view based on this operational status information, including: Real-time monitoring of the operating status information of computing devices in each node of the computing power cluster; When the inference task of the target model is received, a lightweight snapshot is uniformly pulled within the preset acquisition window through the control container group in the computing power cluster to obtain the current running status information of the computing devices in each node; The current running status information is organized into a global resource view in tabular form.

3. The computing resource allocation method according to claim 1, characterized in that, Based on the global resource view, the memory requirements of the target model, and the parallel strategy, several feasible resource allocation candidate schemes are generated, including: Based on the global resource view, the memory requirements of the target model, and the parallel strategy, multiple feasible resource allocation candidate schemes are enumerated, and the variable parameters of the resource allocation candidate schemes are obtained. Configure global variable parameters to bind the variable parameters of each resource allocation candidate scheme to the global variable parameters.

4. The computing resource allocation method according to claim 1, characterized in that, The optimization model includes an objective function and constraints; The constraints include global consistency constraints and linear inequality constraints.

5. The computing power resource allocation method according to claim 4, characterized in that, The expression for the objective function is: In the formula, Indicates the weighting coefficient; Indicates the first The selection variables for each node ,when When, it indicates the first When no node is selected, When, it indicates the first One node was selected; Indicates the first Available video memory for each computing device; Indicates the first The selection variables for each computing device ,when When, it indicates the first A computing device was not selected when When, it indicates the first One computing device was selected; Indicates according to the first When deploying a model with multiple resource allocation candidate schemes, the amount of video memory usage needs to be evenly distributed across each computing device. Indicates the first The actual total number of computing devices required for each resource allocation candidate scheme; Indicates the first The selection variables for each resource allocation candidate scheme ,when When, it indicates the first When no resource allocation candidate was selected, When, it indicates the first One resource allocation candidate was selected.

6. The computing resource allocation method according to claim 5, characterized in that, The expression for the global consistency constraint is: In the formula, Indicates the total number of computing devices. ; This indicates the total number of devices that need to be calculated globally; The expression for the linear inequality constraint is: In the formula, This indicates the actual allocated video memory of the computing device; Represents a constant.

7. The computing resource allocation method according to claim 1, characterized in that, Solving the optimization model to obtain the optimal resource allocation candidate scheme includes: The constraint planning-Boolean satisfiability solver is invoked to solve the optimization model and obtain the optimal resource allocation candidate scheme.

8. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the computing resource allocation method according to any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the computing resource allocation method according to any one of claims 1-7.

10. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the computing resource allocation method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Resource allocation method, device and equipment for large model cluster, storage medium and program product

    CN119597368A

  • Multi-center heterogeneous computing power intelligent scheduling method, platform and equipment and storage medium

    CN119739502A

  • LLM model resource allocation method and device under GPU cluster based on mixed integer programming

    CN119988043A

  • Inference service resource configuration method, electronic equipment and readable storage medium

    CN120743558A

  • Resource allocation method and system

    US20250251976A1

Cited By

  • Resource planning method and device, electronic equipment, storage medium and program product

    CN122044892A