Resource allocation method and device, equipment and storage medium
By determining the hardware resource configuration in the bare metal node pool based on model information and performance goals, and finding and allocating bare metal nodes that meet the configuration, the problem of resource waste or insufficiency in the cloud platform is solved, and the matching allocation of resources and performance is realized, thereby improving resource utilization efficiency and model performance.
Patent Information
- Application Number
- CN202511536076.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2025-11-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing cloud platforms suffer from resource waste or insufficient allocation when distributing resources for deploying artificial intelligence models. This is especially true in virtualization technology, where physical resources cannot be fully used for deploying AI models, leading to resource waste. Alternatively, excessive resources may be allocated to improve performance, resulting in some resources being idle, while insufficient resources may be allocated to reduce resource consumption, causing performance to fall short of expectations.
By determining the required hardware resource configuration in the bare metal node pool based on the model information and performance goals of the artificial intelligence model, and then searching for nodes in the bare metal node pool that meet the configuration for allocation, including searching for nodes or multiple nodes with the least idle hardware resources in the bare metal node pool, in order to match the performance goals and avoid resource waste or insufficiency.
It achieves a reasonable allocation of hardware resources for artificial intelligence models, ensuring that resources match performance goals, avoiding resource waste and shortages, and improving resource allocation efficiency and model performance.
Smart Images

Figure CN121008934A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a resource allocation method and device, an electronic device, and a storage medium. BACKGROUND
[0002] Some cloud platforms can be used to uniformly manage and schedule resources of multiple servers. The resources of the servers can include, but are not limited to, memory, a central processing unit (CPU), a hard disk, a graphics processing unit (GPU), and the like in the servers. When a user needs to use the resources of the servers, the user can send a resource allocation request to the cloud platform through a client, and the cloud platform allocates the resources.
[0003] At present, with the development of artificial intelligence technology, many users are deploying artificial intelligence models required by the users. Before deploying the artificial intelligence models, the users can send a resource allocation request to the cloud platform, and the cloud platform allocates resources for deploying the artificial intelligence models. However, the resources allocated by the cloud platform for the artificial intelligence models can be unreasonable, and there can be problems of resource waste or resource deficiency. SUMMARY
[0004] The present application provides a resource allocation method, a resource allocation device, an electronic device, and a computer readable storage medium to at least solve the problems of hardware resource waste or hardware resource deficiency in the related art.
[0005] The present application provides a resource allocation method, comprising: receiving a resource allocation request; if the resource allocation request represents allocation of hardware resources to a to-be-deployed artificial intelligence model, obtaining model information and a performance target of the artificial intelligence model; determining a hardware resource configuration required by the artificial intelligence model based on the model information and the performance target; in a bare metal node pool, searching for a bare metal node satisfying the hardware resource configuration according to idle hardware resources of each bare metal node in the bare metal node pool, and allocating at least one bare metal node found to the artificial intelligence model, wherein the bare metal node pool comprises a plurality of bare metal nodes, and each bare metal node has a respective corresponding hardware resource.
[0006] The present application also provides a resource allocation device, comprising: a request receiving module configured to receive a resource allocation request; an information obtaining module configured to, if the resource allocation request represents allocation of hardware resources to a to-be-deployed artificial intelligence model, obtain model information and a performance target of the artificial intelligence model; a configuration determining module configured to determine a hardware resource configuration required by the artificial intelligence model based on the model information and the performance target; a node allocating module configured to search for a bare metal node satisfying the hardware resource configuration according to idle hardware resources of each bare metal node in a bare metal node pool, and allocate at least one bare metal node found to the artificial intelligence model, wherein the bare metal node pool comprises a plurality of bare metal nodes, and each bare metal node has a respective corresponding hardware resource.
[0007] The application further provides an electronic device, comprising a memory configured to store a computer program, and a processor configured to execute the computer program to implement the steps of any of the resource allocation methods.
[0008] The application further provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of any of the resource allocation methods.
[0009] In the technical solutions of some embodiments of the application, the hardware resource configuration required by the artificial intelligence model is determined according to the model information and the performance target of the artificial intelligence model, and the bare metal node satisfying the hardware resource configuration is searched for according to the idle hardware resources of each bare metal node in a bare metal node pool, and at least one bare metal node found is allocated to the artificial intelligence model. In this way, on the one hand, the hardware resources allocated to the artificial intelligence model can match the performance target and the actual parameter quantity of the artificial intelligence model, and thus the problem of insufficient hardware resources can be avoided. On the other hand, under the premise that the operation of the artificial intelligence model can reach the performance target, the problem of allocating too many hardware resources to the artificial intelligence model can be avoided, and thus the problem of resource waste can be prevented. In summary, when the hardware resources are allocated to the artificial intelligence model based on the resource allocation method of the application, the problems of hardware resource waste or insufficient hardware resources in some technologies can be solved. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the application, the drawings required to be used in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort.
[0011] Figure 1 An architecture schematic diagram of a cloud platform provided by some embodiments of the application is shown; Figure 2 A flowchart of a resource allocation method provided by some embodiments of the application is shown; Figure 3A module schematic diagram of a resource allocation apparatus provided for some embodiments of the present application is shown in the figure; Figure 4 A module schematic diagram of an electronic device provided for some embodiments of the present application is shown in the figure. DETAILED DESCRIPTION
[0012] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, any other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.
[0013] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0014] In order to enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0015] At present, some cloud platforms have problems of resource waste or resource deficiency when allocating resources for deploying artificial intelligence models, which are embodied in the following three aspects: 1) These cloud platforms construct multiple virtual machines based on virtualization technology and physical resources of servers, and the multiple virtual machines share server hardware resources. When receiving a request for allocating resources for an artificial intelligence model, the cloud platform allocates one or more virtual machines to the artificial intelligence model. In this virtualization technology, part of the physical resources of the server needs to be used for virtualization layer overhead, i.e. the physical resources of the server cannot be fully used for deployment of the artificial intelligence model, and therefore there is a problem of resource waste.
[0016] 2) In order to improve the performance of the artificial intelligence model, too many resources are allocated to the artificial intelligence model, resulting in that part of the resources are in an idle state, and therefore there is a problem of resource waste.
[0017] 3) In order to reduce resource consumption, the artificial intelligence model is allocated insufficient resources, resulting in that the performance of the artificial intelligence model cannot reach the expected target.
[0018] To solve the above problems, the present application first provides an OpenStack-based cloud platform. In combination with reference toFigure 1 An architecture diagram of a cloud platform provided for some embodiments of the present application. Figure 1 In some embodiments of the present application, the cloud platform includes a bare metal service, a bare metal node pool, a database table, and an image tool and a file library required for an artificial intelligence model.
[0019] In some embodiments of the present application, the bare metal node pool includes a plurality of bare metal nodes registered to the cloud platform. The bare metal node refers to a hardware resource provided by the cloud platform directly to a user in the form of a physical server. Each bare metal node represents the hardware resource of a server. For example, the bare metal node R1 represents the hardware resource of the server S1, the bare metal node R2 represents the hardware resource of the server S2, and so on. At least part of the bare metal nodes can be used to deploy an artificial intelligence model. Specifically, in the bare metal nodes used to deploy the artificial intelligence model, each bare metal node can be used to deploy one or more artificial intelligence models. For example, assuming that the bare metal node R1 includes 24 CPU cores, 32G memory, 100G hard disk, and 3 GPUs, 16 GPU cores, 16G memory, 50G hard disk, and 1 GPU of the bare metal node R1 can be allocated to the artificial intelligence model M1, and 8 GPU cores, 16G memory, 50G hard disk, and 2 GPUs of the bare metal node R1 can be allocated to the artificial intelligence model M2.
[0020] The tool image is used to provide a running environment for the artificial intelligence model in the bare metal node. Specifically, the running environment includes an inference framework, a driver, and a dependent library of the artificial intelligence model, etc. The tool image includes but is not limited to ollam, vlim, OpenAI, and GGUF (GPT-Generated Unified Format). The file library refers to various files required when training, deploying, and running the artificial intelligence model, such as a model weight file, a model architecture file, and a model configuration file, etc.
[0021] The database table includes a bare metal resource information table, a bare metal information table, and a bare metal resource usage table. In combination with Table 1, a data structure of the bare metal resource information table provided for some embodiments of the present application is provided.
[0022] Table 1 Bare metal resource information table
[0023] For example, assume that the node ID of the bare metal node R1 is 001, and the node ID of the bare metal node R2 is 002. The bare metal node R1 includes 10 CPU cores and 5 GPUs, and the resource IDs of the CPU cores are 00001-00010, and the resource IDs of the GPUs are 00011-00015. The bare metal node R2 includes 12 CPU cores and 6 GPUs, and the resource IDs of the CPU cores are 00016-00027, and the resource IDs of the GPUs are 00028-00033. Then, the bare metal resource information table can include records as shown in Table 2.
[0024] Table 2: Record example in bare metal resource information table
[0025] Referring to Table 3, a data structure of the bare metal information table provided by some embodiments of the present application is provided.
[0026] Table 3: Bare metal information table
[0027] For example, assume that the node ID of the bare metal node R1 is 001, the node name is R1, the current node state is running, and the bare metal node R1 can be used to deploy an artificial intelligence model, the node ID of the bare metal node R2 is 002, the node name is R2, the current node state is running, and the bare metal node R2 can be used to deploy a non-artificial intelligence service. Then, the bare metal information table can include records as shown in Table 4.
[0028] Table 4: Record example in bare metal information table
[0029] Referring to Table 5, a data structure of the bare metal resource usage table provided by some embodiments of the present application is provided.
[0030] Table 5: Bare metal resource usage table
[0031] Based on the cloud platform shown in FIG. 1, the present application provides a resource allocation method, which can solve the problems of hardware resource waste or hardware resource shortage. The resource allocation method can be applied to the bare metal service shown in FIG. 1. Referring to FIG. 2, a flowchart of the resource allocation method provided by some embodiments of the present application is provided. Figure 1 Based on the cloud platform shown in FIG. 1, the present application provides a resource allocation method, which can solve the problems of hardware resource waste or hardware resource shortage. The resource allocation method can be applied to the bare metal service shown in FIG. 1. Referring to FIG. 2, a flowchart of the resource allocation method provided by some embodiments of the present application is provided. Figure 1 Based on the cloud platform shown in FIG. 1, the present application provides a resource allocation method, which can solve the problems of hardware resource waste or hardware resource shortage. The resource allocation method can be applied to the bare metal service shown in FIG. 1. Referring to FIG. 2, a flowchart of the resource allocation method provided by some embodiments of the present application is provided. Figure 2 Based on the cloud platform shown in FIG. 1, the present application provides a resource allocation method, which can solve the problems of hardware resource waste or hardware resource shortage. The resource allocation method can be applied to the bare metal service shown in FIG. 1. Referring to FIG. 2, a flowchart of the resource allocation method provided by some embodiments of the present application is provided. Figure 2 In some embodiments of the present application, the resource allocation method includes the following steps: Step S201: receiving a resource allocation request.
[0032] Specifically,Figure 1 The bare metal service shown can provide an external interface. When a user needs to deploy an artificial intelligence model or a non-artificial intelligence service using the hardware resources of the server, the user can call the external interface provided by the bare metal through a client to request the cloud platform to allocate hardware resources for deploying the artificial intelligence model or the non-artificial intelligence service.
[0033] Further, in the resource allocation request, the user can specify a resource allocation type. For example, when the resource allocation type is 0, it indicates that the hardware resources are requested to be allocated to the artificial intelligence model; when the resource allocation type is 1, it indicates that the hardware resources are requested to be allocated to the non-artificial intelligence service. Based on the resource allocation type in the resource allocation request, the bare metal service can determine the type of service that needs to use the hardware resources, and then allocate the hardware resources according to different resource allocation logics, wherein the resource allocation logics can be referred to steps S202-S204 and subsequent related descriptions, which are not described here.
[0034] In step S202, if the resource allocation request represents that the hardware resources are allocated to the to-be-deployed artificial intelligence model, the model information and performance target of the artificial intelligence model are obtained.
[0035] Specifically, the model information can include but is not limited to model type, model framework, model calculation mode, model parameter quantity, model network layer number, model input and output dimension, etc. Among them, the model type refers to the category to which the model belongs, such as large language model, recommendation model, etc. The model framework refers to the deep learning framework or inference engine that the model depends on, such as PyTorch 2.3, TensorRT 8.6, etc. The calculation mode of the model can be divided into two modes of model training and model inference. Each artificial intelligence model has its own corresponding model information, and the model information of different artificial intelligence models can be different. For example, the artificial intelligence model R1 is a to-be-trained large language model, its model framework is PyTorch 2.3, the model parameter quantity is 30,000 and has 10 network layers. The artificial intelligence model R2 is a trained large language model for inference, its model framework is TensorRT 8.6, the model parameter quantity is 25,000 and has 8 network layers.
[0036] The performance target refers to the performance index that the artificial intelligence model needs to achieve. The performance index can include but is not limited to the throughput, delay, initialization time of the artificial intelligence model, etc. Each artificial intelligence model has its own corresponding performance target, and the performance targets of different artificial intelligence models can be different. For example, the performance target of the artificial intelligence model R1 is: the throughput needs to be higher than 30 tokens per second, the delay needs to be less than 1 second, and the initialization time needs to be less than 2 seconds. The performance target of the artificial intelligence model R2 is: the throughput needs to be higher than 35 tokens per second, the delay needs to be less than 1.5 seconds, and the initialization time needs to be less than 2 seconds.
[0037] In this embodiment, the resource allocation request can include a request parameter. When the user sends the resource allocation request through the client, the user can fill in the model information and the performance target of the artificial intelligence model in the request parameter according to the actual situation of the artificial intelligence model to be deployed. After the bare metal service receives the resource allocation request, the model information and the performance target of the artificial intelligence model can be extracted from the request parameter of the resource allocation request.
[0038] In this embodiment, by specifying the model information and the performance target in the request parameter of the resource allocation request, the bare metal service does not need to obtain the model information and the performance target through other ways after receiving the resource allocation request, and thus the resource allocation logic can be simplified and the resource allocation efficiency can be improved.
[0039] In other embodiments, before the user sends the resource allocation request through the client, the user can submit the model identifier, the model information and the performance target of the artificial intelligence model to the cloud platform in advance, and the cloud platform can save the model identifier, the model information and the performance target submitted by the user in a database. When the user sends the resource allocation request through the client, the user can fill in the model identifier in the request parameter. After the bare metal service receives the resource allocation request, the model identifier of the artificial intelligence model can be extracted from the request parameter of the resource allocation request, and the model information and the performance target of the artificial intelligence model can be obtained from the database based on the model identifier. In these embodiments, by submitting the model information and the performance target to the cloud platform in advance, the parameter amount of the resource allocation request can be reduced.
[0040] In still other embodiments, after the bare metal service receives the resource allocation request, an information obtaining interface can be displayed, and the model information and the performance target of the artificial intelligence model can be received through the information obtaining interface. Specifically, in these embodiments, the user can not need to fill in the model identifier, the model information and the performance target of the artificial intelligence model in the request parameter of the resource allocation request, but can fill in the model identifier, the model information and the performance target in the information obtaining interface displayed by the cloud platform after sending the resource allocation request to the cloud platform. In this way, on the one hand, the parameter amount of the resource allocation request can be further reduced, and on the other hand, the model identifier, the model information and the performance target can also not need to be submitted to the cloud platform in advance, and thus the operation process of resource allocation can be simplified.
[0041] In step S203, the hardware resource configuration required by the artificial intelligence model is determined based on the model information and the performance target.
[0042] Specifically, the hardware resource configuration refers to the type of hardware resource required by the artificial intelligence model and the number of each type of hardware resource under the condition of achieving the performance target. The type of hardware resource can include but is not limited to CPU core, memory, hard disk, GPU, etc.
[0043] It can be understood that for different two artificial intelligence models, if the model information or performance target of the two artificial intelligence models is not the same, the hardware resource configuration required by the two artificial intelligence models can be different. For example, if the artificial intelligence model R1 has 30,000 model parameters, and the throughput of the artificial intelligence model R1 needs to be higher than 30 tokens per second, and the delay needs to be less than 1 second, the artificial intelligence model R1 can require 16 CPU cores, 50G memory, 150G hard disk and 8 GPUs. If the artificial intelligence model R2 has 20,000 model parameters, and the throughput of the artificial intelligence model R2 needs to be higher than 20 tokens per second, and the delay needs to be less than 1.5 seconds, the artificial intelligence model R2 can require 10 CPU cores, 30G memory, 100G hard disk and 5 GPUs.
[0044] Similarly, for the same artificial intelligence model, the hardware resource configuration required by the artificial intelligence model can also be different under different performance targets. For example, assuming that the artificial intelligence model R1 has 30,000 model parameters. If the throughput of the artificial intelligence model R1 needs to be higher than 30 tokens per second, and the delay needs to be less than 1 second, the artificial intelligence model R1 can require 16 CPU cores, 50G memory, 150G hard disk and 8 GPUs. If the throughput of the artificial intelligence model R1 needs to be higher than 35 tokens per second, and the delay needs to be less than 0.8 seconds, the artificial intelligence model R1 can require 18 CPU cores, 65G memory, 1155G hard disk and 9 GPUs.
[0045] In this embodiment, the model information and the performance target can be input into the trained evaluation model, and the hardware resource configuration required by the artificial intelligence model is output by the evaluation model. The hardware resource configuration output by the evaluation model can include the type of hardware resource required by the artificial intelligence model and the number of each type of hardware resource. The hardware resource types include but are not limited to CPU cores, memory, hard disk, GPU, etc.
[0046] Compared with determining the hardware resource configuration required by the artificial intelligence model based on artificial experience, the hardware resource configuration required by the artificial intelligence model is determined by the trained evaluation model, and the result is more accurate, thereby reducing the problem of waste or insufficient hardware resources under the premise of achieving the performance target of the artificial intelligence model.
[0047] In other embodiments, the cloud platform can include a pre-established mapping relationship between model information, performance targets and hardware resource configurations. After obtaining the model information and performance target of the artificial intelligence model, the hardware resource configuration required by the artificial intelligence model can be determined based on the mapping relationship. The application does not limit the way of obtaining the hardware resource configuration.
[0048] In step S204, in the bare metal node pool, according to the idle hardware resources of each bare metal node, a bare metal node satisfying the hardware resource configuration is searched, and at least one bare metal node searched is allocated to the artificial intelligence model, wherein the bare metal node pool includes a plurality of bare metal nodes, and each bare metal node has a respective corresponding hardware resource.
[0049] Specifically, in the bare metal node pool, if the idle hardware resources of the plurality of bare metal nodes all satisfy the hardware resource configuration, one of the plurality of bare metal nodes with the least idle hardware resources is allocated to the artificial intelligence model.
[0050] For example, assuming that the hardware resource configuration of the artificial intelligence model M5 is 5 CPU cores, 30G memory, 50G hard disk and 3 GPUs. The bare metal node pool includes three bare metal nodes R1-R3. Among them, the bare metal node R1 includes 6 idle GPU cores, 35G idle memory, 60G idle hard disk and 4 idle GPUs; the bare metal node R2 includes 2 idle GPU cores, 15G idle memory, 20G idle hard disk and 2 idle GPUs; the bare metal node R3 includes 8 idle GPU cores, 50G idle memory, 65G idle hard disk and 5 idle GPUs. Based on the idle hardware resources of the bare metal nodes R1-R3, it can be known that the idle hardware resources of the bare metal node R2 do not satisfy the hardware resource configuration required by the artificial intelligence model M5, and the idle hardware resources of the bare metal nodes R1 and R3 both satisfy the hardware resource configuration of the artificial intelligence model M5. Since the idle hardware resources of the bare metal node R1 are the least among the bare metal nodes R1 and R3, 5 CPU cores, 30G memory, 50G hard disk and 3 GPUs can be allocated to the artificial intelligence model M5 in the idle hardware resources of the bare metal node R1.
[0051] In this embodiment, when the idle hardware resources of the plurality of bare metal nodes all satisfy the hardware resource configuration, one of the bare metal nodes with the least idle hardware resources is allocated to the artificial intelligence model, so that the bare metal node with more idle hardware resources can be reserved, and then when other artificial intelligence models require more hardware resources, the hardware resources can be allocated to other artificial intelligence models in the same bare metal node, avoiding cross bare metal node allocation of hardware resources, so that cross bare metal node communication can be avoided when other artificial intelligence models run, thereby improving the performance of other artificial intelligence models.
[0052] For example, assuming that the idle hardware resources of the bare metal node R1 are allocated to the artificial intelligence model M5, if a request for allocating hardware resources to the artificial intelligence model M6 is subsequently received, and the hardware resource configuration required by the artificial intelligence model M6 is 7 CPU cores, 45G memory, 62G hard disk and 4 GPUs. Since the number of hardware resources required by the artificial intelligence model M6 is greater than the number of hardware resources required by the artificial intelligence model M5, the artificial intelligence model M6 can be allocated hardware resources in the idle hardware resources of the bare metal node R3. Since the idle hardware resources of the bare metal node R3 can meet the number of hardware resources required by the artificial intelligence model M6, there is no need to allocate hardware resources across bare metal nodes, thereby improving the performance of the artificial intelligence model M6.
[0053] On the contrary, if the idle hardware resources of the bare metal node R3 are preferentially allocated to the artificial intelligence model M5 when allocating hardware resources to the artificial intelligence model M5, the idle hardware resources of the bare metal nodes R1-R3 cannot meet the hardware resource configuration required by the artificial intelligence model M5 when allocating hardware resources to the artificial intelligence model M6. At this time, it is necessary to allocate hardware resources across bare metal nodes, thereby affecting the performance of the artificial intelligence model M6.
[0054] Of course, in the bare metal node pool, if there is no single bare metal node that meets the hardware resource configuration, multiple bare metal nodes whose sum of idle hardware resources meets the hardware resource configuration are searched for, and the searched multiple bare metal nodes are allocated to the artificial intelligence model. By allocating hardware resources to the artificial intelligence model across bare metal nodes, it can be ensured that sufficient hardware resources are allocated to the artificial intelligence model, thereby avoiding the problem of insufficient hardware resources of the artificial intelligence model.
[0055] In summary, in the technical solutions of some embodiments of the present application, the hardware resource configuration required by the artificial intelligence model is determined according to the model information and performance target of the artificial intelligence model, and in the bare metal node pool, a bare metal node that meets the hardware resource configuration is searched for according to the idle hardware resources of each bare metal node, and at least one searched bare metal node is allocated to the artificial intelligence model. In this way, on the one hand, the hardware resources allocated to the artificial intelligence model can match the performance target, actual parameter quantity, etc. of the artificial intelligence model, thereby avoiding the problem of insufficient hardware resources. On the other hand, under the premise of ensuring that the operation of the artificial intelligence model can reach the performance target, it can be avoided to allocate too many hardware resources to the artificial intelligence model, thereby preventing the problem of resource waste. In summary, when the resource allocation method based on the present application allocates hardware resources to the artificial intelligence model, it can solve the problems of hardware resource waste or insufficient hardware resources in some technologies.
[0056] Further, considering that different manufacturers or different models of the same type of hardware resource can have different performance, therefore, when deploying the same artificial intelligence model (i.e., the model information and performance indicators are the same), if the manufacturer or model of the same type of hardware resource is different, the required number of hardware resources can be different. For example, assuming that the GPU produced by manufacturer A has higher performance than the GPU produced by manufacturer B, then when deploying the same artificial intelligence model, if the GPU produced by manufacturer A is used, only 8 GPUs can be needed, but if the GPU produced by manufacturer B is used, 10 GPUs can be needed. For another example, assuming that the GPU of model C has higher performance than the GPU of model D, then when deploying the same artificial intelligence model, if the GPU of model C is used, only 7 GPUs can be needed, but if the GPU of model D is used, 9 GPUs can be needed.
[0057] In view of this, in some embodiments, the hardware resource configuration in step S203 can further include auxiliary information of each type of hardware resource in addition to the type of hardware resource and the number of each type of hardware resource. That is, in step S203, a plurality of different hardware resource configurations can be output according to the auxiliary information. For example: Hardware resource configuration C11: 5 CPU cores, 30G memory, 100G hard disk, 6 GPUs, wherein the CPU is produced by manufacturer H1, the memory is produced by manufacturer H2, the hard disk is produced by manufacturer H3, and the GPU is produced by manufacturer H4.
[0058] Hardware resource configuration C12: 6 CPU cores, 35G memory, 80G hard disk, 5 GPUs, wherein the CPU is produced by manufacturer H5, the memory is produced by manufacturer H6, the hard disk is produced by manufacturer H7, and the GPU is produced by manufacturer H8.
[0059] In this way, the accuracy of the hardware resource configuration can be improved, and the accuracy of the resource allocation can be improved.
[0060] In some embodiments, before allocating hardware resources to the artificial intelligence model, the auxiliary information (i.e., the manufacturer and the model) of the hardware resources represented by each bare metal node in the bare metal node pool can be collected in advance, and then the model information, the performance target, and the auxiliary information of the hardware resources are input into the trained evaluation model to determine the required hardware resource configuration of the artificial intelligence model by the evaluation model. In this way, the evaluation model can generate the hardware resource configuration based on the hardware resource information in the bare metal node pool, preventing the evaluation model from outputting too many invalid hardware resource configurations. For example, in the hardware resources of the bare metal node pool, there can be no GPU produced by manufacturer H6, but in the hardware resource configuration output by the evaluation model, there is a GPU produced by manufacturer H6, in which case the evaluation model outputs too many invalid hardware resource configurations.
[0061] Further, in combination with referring to Table 3 and Table 4, since the bare metals in the bare metal pool can be divided into two categories by the Tag field, and one category of the bare metals is used to deploy the artificial intelligence model, and the other category of the bare metals is used to deploy the non-artificial intelligence service, the step S204 of searching for the bare metal nodes satisfying the hardware resource configuration according to the idle hardware resources of the bare metal nodes in the bare metal node pool and assigning the found at least one bare metal node to the artificial intelligence model can include: searching for the first type of bare metal nodes in the bare metal node pool, wherein the first type of bare metal nodes refer to the pre-planned candidate bare metal nodes for being assigned to the artificial intelligence model; searching for the bare metal nodes in the running state in the first type of bare metal nodes; searching for the bare metal nodes satisfying the hardware resource configuration according to the idle hardware resources of the bare metal nodes in the running state.
[0062] In this way, the hardware resources can be assigned to the artificial intelligence model according to the pre-planned use of the bare metal nodes. Meanwhile, the bare metal nodes in the running state and satisfying the hardware resource configuration are assigned to the artificial intelligence model in the first type of bare metal nodes, so that the problem of the deployment failure of the artificial intelligence model caused by the abnormal bare metal nodes can be prevented.
[0063] In some embodiments, the searching for the bare metal nodes satisfying the hardware resource configuration according to the idle hardware resources of the bare metal nodes in the running state includes: if the idle hardware resources of multiple bare metal nodes in the running state all satisfy the hardware resource configuration, assigning one of the multiple bare metal nodes with the least idle hardware resources to the artificial intelligence model.
[0064] if there is no single bare metal node satisfying the hardware resource configuration in the running state, searching for multiple bare metal nodes with the sum of the idle hardware resources satisfying the hardware resource configuration, and assigning the found multiple bare metal nodes to the artificial intelligence model.
[0065] Specifically, the relevant principles and the step S204 are similar, and are not described herein.
[0066] In some embodiments, after assigning the bare metal nodes to the artificial intelligence model, the method of the present application further includes: deploying a tool image matched with the artificial intelligence model in the assigned bare metal nodes, wherein the tool image is used to provide a running environment for the artificial intelligence model in the bare metal nodes, and the running environment includes an inference framework, a driver and a dependent library of the artificial intelligence model.
[0067] In this way, normal deployment of subsequent artificial intelligence models can be ensured.
[0068] In some embodiments, the method of the present application further comprises: If the resource allocation request represents allocation of hardware resources to non-artificial intelligence services, a second type of bare metal node is searched in the bare metal node pool, wherein the second type of bare metal node refers to a pre-planned alternative bare metal node for allocation to non-artificial intelligence services.
[0069] Specifically, if the resource allocation request represents allocation of hardware resources to non-artificial intelligence services, the service type can be included in the resource allocation request. The bare metal service can determine the hardware resource configuration required by the non-artificial intelligence service based on the service type, and then allocate hardware resources to the non-artificial intelligence service according to the determined hardware resource configuration.
[0070] By dividing the bare metal nodes in the bare metal node pool into the first type of bare metal node and the second type of bare metal node, deployment of artificial intelligence models in bare metal nodes that do not meet the model deployment requirements can be avoided.
[0071] In some embodiments, before receiving the resource allocation request, the method of the present application further comprises: receiving a bare metal registration request, the bare metal registration request including a target bare metal to be registered and hardware resource information of the target bare metal; adding the target bare metal and the hardware resource information of the target bare metal in the bare metal node pool to complete registration of the target bare metal.
[0072] In the technical solution of some embodiments of the present application, the hardware resource configuration required by the artificial intelligence model is determined according to the model information and performance target of the artificial intelligence model, and in the bare metal node pool, a bare metal node that meets the hardware resource configuration is searched according to the idle hardware resources of each bare metal node, and at least one bare metal node found is allocated to the artificial intelligence model. In this way, on the one hand, the hardware resources allocated to the artificial intelligence model can match the performance target, actual parameter quantity, etc. of the artificial intelligence model, thereby avoiding the problem of insufficient hardware resources. On the other hand, under the premise of ensuring that the operation of the artificial intelligence model can reach the performance target, allocation of excessive hardware resources to the artificial intelligence model can be avoided, thereby preventing the problem of resource waste. In summary, based on the resource allocation method of the present application, when hardware resources are allocated to artificial intelligence models, the problems of hardware resource waste or insufficient hardware resources in some technologies can be solved.
[0073] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software on a general hardware platform as necessary, and of course can also be realized by hardware, but in many cases the former is a better embodiment.
[0074] With reference to Figure 3 A module schematic diagram of a resource allocation apparatus provided for some embodiments of the present application is shown. Figure 3 In some embodiments, the resource allocation apparatus comprises: The request receiving module 301 is configured to receive a resource allocation request. The information obtaining module 302 is configured to, if the resource allocation request represents that a hardware resource is to be allocated to a to-be-deployed artificial intelligence model, obtain model information and performance targets of the artificial intelligence model. The configuration determining module 303 is configured to determine a hardware resource configuration required by the artificial intelligence model based on the model information and the performance targets. The node allocating module 304 is configured to, in a bare metal node pool, find a bare metal node that meets the hardware resource configuration according to idle hardware resources of each bare metal node in the bare metal node pool, and allocate at least one bare metal node found to the artificial intelligence model, wherein the bare metal node pool comprises a plurality of bare metal nodes, and each bare metal node has a respective corresponding hardware resource.
[0075] In some embodiments, the node allocating module 304 is configured to: In the bare metal node pool, if idle hardware resources of a plurality of bare metal nodes all meet the hardware resource configuration, one of the plurality of bare metal nodes with the least idle hardware resources is allocated to the artificial intelligence model.
[0076] In some embodiments, the node allocating module 304 is configured to: In the bare metal node pool, if there is no single bare metal node that meets the hardware resource configuration, a plurality of bare metal nodes whose sum of idle hardware resources meets the hardware resource configuration are found, and the plurality of bare metal nodes found are allocated to the artificial intelligence model.
[0077] In some embodiments, the node allocating module 304 is configured to: In the bare metal node pool, a first type of bare metal node is found, wherein the first type of bare metal node refers to a pre-planned candidate bare metal node for allocation to the artificial intelligence model. In the first type of bare metal node, a bare metal node in a running state is found. In the bare metal node in the running state, a bare metal node that meets the hardware resource configuration is found according to idle hardware resources of each bare metal node.
[0078] In some embodiments, the node allocation module 304 is configured to: In the running state of the bare metal node, if the idle hardware resources of the plurality of bare metal nodes all meet the hardware resource configuration, one of the bare metal nodes with the least idle hardware resources is allocated to the artificial intelligence model.
[0079] In some embodiments, the node allocation module 304 is configured to: In the running state of the bare metal node, if there is no single bare metal node that meets the hardware resource configuration, a plurality of bare metal nodes whose sum of idle hardware resources meets the hardware resource configuration are found, and the plurality of bare metal nodes are allocated to the artificial intelligence model.
[0080] In some embodiments, the node allocation module 304 is further configured to: If the resource allocation request represents the allocation of hardware resources to non-artificial intelligence services, in the pool of bare metal nodes, a second type of bare metal node is found, wherein the second type of bare metal node refers to a pre-planned candidate bare metal node for allocation to non-artificial intelligence services.
[0081] In some embodiments, after allocating the bare metal node to the artificial intelligence model, the node allocation module 304 is further configured to: In the bare metal node allocated to the artificial intelligence model, a tool image matched with the artificial intelligence model is deployed, and the tool image is used to provide a running environment for the artificial intelligence model in the bare metal node, and the running environment includes an inference framework, a driver and a dependent library of the artificial intelligence model.
[0082] In some embodiments, the information acquisition module 302 is configured to: In the request parameter of the resource allocation request, the model information and the performance target of the artificial intelligence model are extracted; Or, in the request parameter of the resource allocation request, the model identifier of the artificial intelligence model is extracted, and based on the model identifier, the model information and the performance target of the artificial intelligence model are acquired in the database; Or, after receiving the resource allocation request, an information acquisition interface is displayed, and the model information and the performance target of the artificial intelligence model are received through the information acquisition interface.
[0083] In some embodiments, the configuration determination module 303 is configured to: The model information and the performance target are input into the trained evaluation model, and the hardware resource configuration required by the artificial intelligence model is output by the evaluation model.
[0084] In some embodiments, the hardware resource configuration output by the evaluation model includes the following information: The hardware resource type required by the artificial intelligence model; a number of each type of hardware resource required by the artificial intelligence model.
[0085] In some embodiments, before receiving the resource allocation request, the request receiving module 301 is further configured to: receive a bare metal registration request, the bare metal registration request including a target bare metal to be registered and hardware resource information of the target bare metal; add the target bare metal and the hardware resource information of the target bare metal in the bare metal node pool to complete registration of the target bare metal.
[0086] In combination with the above Figure 4 Embodiments of the present application also provide an electronic device, including a memory 10 and a processor 20, the memory 10 storing a computer program, and the processor 20 being configured to execute the computer program to perform the steps in any of the above resource allocation method embodiments.
[0087] Embodiments of the present application also provide a computer readable storage medium storing a computer program, wherein the computer program is configured to perform the steps in any of the above resource allocation method embodiments when executed.
[0088] In an example embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0089] Embodiments of the present application also provide a computer program product, the above computer program product including a computer program, and the computer program being executed by a processor to implement the steps in any of the above resource allocation method embodiments.
[0090] Embodiments of the present application also provide another computer program product, including a non-volatile computer readable storage medium, the non-volatile computer readable storage medium storing a computer program, and the computer program being executed by a processor to implement the steps in any of the above resource allocation method embodiments.
[0091] Those skilled in the art will further realize that the mere concepts, teachings, and embodiments described herein are merely meant to provide an enabling description of the applications and are not intended to limit the scope of the applications. Therefore, embodiments or examples described herein are not meant to be limiting, but merely to aid in the understanding of the overall more complete disclosure of the applications. Accordingly, the disclosure of various examples and embodiments is meant to be illustrative and not limiting of the scope of the applications, as claimed.
[0092] The resource allocation method, device, apparatus and storage medium provided by the present application are described in detail above. The principles and implementation modes of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method and its core idea of the present application. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application. These improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A resource allocation method, characterized in that, The method includes: Receive resource allocation requests; If the resource allocation request represents allocating hardware resources to the AI model to be deployed, then obtain the model information and performance target of the AI model; Based on the model information and the performance target, determine the hardware resource configuration required for the artificial intelligence model; In the bare metal node pool, based on the idle hardware resources of each bare metal node, a bare metal node that meets the hardware resource configuration is searched, and at least one bare metal node is assigned to the artificial intelligence model. The bare metal node pool includes multiple bare metal nodes, and each bare metal node has its own corresponding hardware resources.
2. The method according to claim 1, characterized in that, The step of searching for bare metal nodes that meet the hardware resource configuration in the bare metal node pool based on the idle hardware resources of each bare metal node, and allocating at least one of the found bare metal nodes to the artificial intelligence model, includes: If multiple bare metal nodes in the bare metal node pool have idle hardware resources that meet the hardware resource configuration, then the bare metal node with the fewest idle hardware resources will be allocated to the artificial intelligence model.
3. The method according to claim 1, characterized in that, The step of searching for bare metal nodes that meet the hardware resource configuration in the bare metal node pool based on the idle hardware resources of each bare metal node, and allocating at least one of the found bare metal nodes to the artificial intelligence model, includes: If no single bare metal node in the bare metal node pool satisfies the hardware resource configuration, then multiple bare metal nodes whose sum of idle hardware resources satisfies the hardware resource configuration are searched, and the multiple bare metal nodes found are allocated to the artificial intelligence model.
4. The method according to claim 1, characterized in that, The step of searching for bare metal nodes that meet the hardware resource configuration in the bare metal node pool based on the idle hardware resources of each bare metal node, and allocating at least one of the found bare metal nodes to the artificial intelligence model, includes: In the bare metal node pool, search for a first type of bare metal node, wherein the first type of bare metal node refers to a pre-planned candidate bare metal node for allocation to the artificial intelligence model; Among the bare metal nodes of the first type, locate the bare metal nodes that are in a running state; Among the bare metal nodes that are in operation, a bare metal node that meets the hardware resource configuration is found based on the idle hardware resources of each bare metal node.
5. The method according to claim 4, characterized in that, The step of finding a bare metal node that meets the hardware resource configuration based on the available hardware resources of each bare metal node in the running state includes: If multiple bare metal nodes in operation have idle hardware resources that satisfy the hardware resource configuration, then the bare metal node with the fewest idle hardware resources will be allocated to the artificial intelligence model.
6. The method according to claim 4, characterized in that, The step of finding a bare metal node that meets the hardware resource configuration based on the available hardware resources of each bare metal node in the running state includes: If there is no single bare metal node that satisfies the hardware resource configuration among the bare metal nodes in operation, then multiple bare metal nodes whose sum of idle hardware resources satisfies the hardware resource configuration are searched, and the multiple bare metal nodes found are assigned to the artificial intelligence model.
7. The method according to claim 4, characterized in that, The method further includes: If the resource allocation request represents allocating hardware resources to non-AI services, then in the bare metal node pool, a second type of bare metal node is searched, wherein the second type of bare metal node refers to a pre-planned candidate bare metal node for allocation to non-AI services.
8. The method according to any one of claims 1 to 6, characterized in that, After assigning bare metal nodes to the artificial intelligence model, the method further includes: In the bare metal node allocated to the artificial intelligence model, a tool image matching the artificial intelligence model is deployed. The tool image is used to provide a runtime environment for the artificial intelligence model in the bare metal node. The runtime environment includes the inference framework, drivers, and dependency libraries of the artificial intelligence model.
9. The method according to claim 1, characterized in that, The acquisition of model information and performance targets of the artificial intelligence model includes: Extract the model information and performance target of the artificial intelligence model from the request parameters of the resource allocation request; Alternatively, extract the model identifier of the artificial intelligence model from the request parameters of the resource allocation request, and obtain the model information and performance target of the artificial intelligence model from the database based on the model identifier; Alternatively, upon receiving the resource allocation request, an information acquisition interface is displayed, and the model information and performance target of the artificial intelligence model are received through the information acquisition interface.
10. The method according to claim 1, characterized in that, The step of determining the hardware resource configuration required for the artificial intelligence model based on the model information and the performance target includes: The model information and the performance target are input into the trained evaluation model, and the evaluation model outputs the hardware resource configuration required by the artificial intelligence model.
11. The method according to claim 10, characterized in that, The hardware resource configuration output by the evaluation model includes the following information: The types of hardware resources required by the artificial intelligence model; The number of various types of hardware resources required by the artificial intelligence model.
12. The method according to claim 1, characterized in that, Before receiving the resource allocation request, the method further includes: Receive a bare metal registration request, the bare metal registration request including the target bare metal to be registered and the hardware resource information of the target bare metal; The target bare metal and its hardware resource information are added to the bare metal node pool to complete the registration of the target bare metal.
13. A resource allocation device, characterized in that, The device includes: The request receiving module is used to receive resource allocation requests; The information acquisition module is used to acquire the model information and performance target of the artificial intelligence model if the resource allocation request indicates that hardware resources are allocated to the artificial intelligence model to be deployed. A configuration determination module is used to determine the hardware resource configuration required by the artificial intelligence model based on the model information and the performance target; The node allocation module is used to find bare metal nodes that meet the hardware resource configuration in the bare metal node pool according to the idle hardware resources of each bare metal node, and allocate at least one bare metal node found to the artificial intelligence model. The bare metal node pool includes multiple bare metal nodes, and each bare metal node has its own corresponding hardware resources.
14. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the resource allocation method as described in any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the resource allocation method as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Artificial intelligence model training method, device and equipment and storage medium
CN111768006A
Resource allocation method, computing device and storage medium
CN113296926A
Resource allocation method for parallel training of task planning model
CN114610501A
Bare metal server control method and device, and medium
CN114995954A
Bare metal deployment method and device and medium thereof
CN116483382A