Resource allocation method, device, equipment and storage medium
By optimizing the GPU resource allocation method of model services, efficient utilization of GPU resources is achieved based on the number of model instances and resource utilization relationship, reducing the cost of the image processing system and improving processing capabilities.
Patent Information
- Application Number
- CN202011108222.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-16
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2040-10-16
AI Technical Summary
The GPU resource utilization rate of existing model services is low, resulting in high cost of image processing systems and requires the deployment of a large number of GPU cards.
By obtaining the preset relationship between the number of model instances and the computing resource utilization rate in the triplet of the model service, the number of target model instances is determined, and GPU resource allocation is performed based on the target computing resource utilization rate and memory resource utilization rate.
Significantly improve the utilization rate of GPU computing resources and video memory resources, reduce the number of GPUs required by the image processing system, reduce system construction costs, and increase the throughput of the image processing system.
Smart Images

Figure CN112035266B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of resource allocation, and in particular to a resource allocation method, apparatus, device, and storage medium. Background Art
[0002] Existing model services generally have at least two instances, each of which occupies a GPU (Graphics Processing Unit) card. This GPU resource allocation method makes the GPU resource utilization rate very low. Furthermore, for an image processing system that includes a large number of model services, it is also necessary to deploy a large number of GPU cards to allocate them to the model services, resulting in a very high cost to build the image processing system. For example, an image SaaS (Software as a Service) system that mainly provides image processing for advertising delivery, advertising retrieval, and advertising playback, includes a large number of model services for image processing, which requires the deployment of a large number of GPU cards, resulting in high construction costs and low GPU resource utilization. Summary of the Invention
[0003] In view of the above-mentioned technical problems, the present application proposes a resource allocation method, apparatus, device and storage medium.
[0004] According to one aspect of the present application, a resource allocation method is provided, comprising:
[0005] Obtaining a preset relationship between the number of model instances in a triple of multiple model services and the computing resource utilization corresponding to each model instance, and a computing resource utilization threshold;
[0006] Determine the target number of model instances for each model service based on a preset relationship between the number of model instances in the triplet of each model service and the computing resource utilization corresponding to each model instance, and the computing resource utilization threshold; wherein the preset relationship is an inverse relationship, and the computing resource utilization threshold is an upper limit of the computing resource utilization corresponding to each model instance;
[0007] The value of the computing resource utilization corresponding to the number of target model instances of each model service is used as the target computing resource utilization corresponding to each target model instance of each model service;
[0008] Get the target memory resource utilization corresponding to each target model instance of each model service;
[0009] Based on the target computing resource utilization corresponding to each target model instance and the target memory resource utilization corresponding to each target model instance, the target model instance of each model service is allocated to the corresponding GPU.
[0010] According to another aspect of the present application, a resource allocation device is provided, comprising:
[0011] An acquisition module, configured to obtain a preset relationship between the number of model instances in a triple of multiple model services and the computing resource utilization corresponding to each model instance, and a computing resource utilization threshold;
[0012] a target model instance quantity determination module, configured to determine the target model instance quantity for each model service based on a preset relationship between the number of model instances in the triplet of each model service and the computing resource utilization corresponding to each model instance, and the computing resource utilization threshold; wherein the preset relationship is an inverse relationship, and the computing resource utilization threshold is an upper limit of the computing resource utilization corresponding to each model instance;
[0013] a target computing resource utilization determination module, configured to use the computing resource utilization value corresponding to the number of target model instances of each model service as the target computing resource utilization corresponding to each target model instance of each model service;
[0014] The target video memory resource utilization acquisition module is used to obtain the target video memory resource utilization corresponding to each target model instance of each model service;
[0015] The GPU resource allocation module is used to allocate the target model instance of each model service to the corresponding GPU based on the target computing resource utilization corresponding to each target model instance and the target memory resource utilization corresponding to each target model instance.
[0016] According to another aspect of the present application, a resource allocation device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the above method.
[0017] According to another aspect of the present application, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0018] By setting the triples of model services, the preset relationship between the number of model instances in the triples and the computing resource utilization corresponding to each model instance, and the computing resource utilization threshold, the target number of model instances for each model service, the target computing resource utilization corresponding to each target model instance, and the target video memory resource utilization are determined according to the preset relationship between the number of model instances in the triples of each model service and the computing resource utilization corresponding to each model instance, as well as the computing resource utilization threshold, so as to achieve compression of the computing resource utilization of the model instances, so that the GPU resource allocation based on the target computing resource utilization corresponding to each target model instance and the target video memory resource utilization corresponding to each target model instance can significantly improve the utilization of GPU computing resources and video memory resources, and the computing resource utilization can be improved by 93%; while keeping the service capacity of the image processing system unchanged, the number of GPUs required by the image processing system can be effectively reduced, thereby reducing the cost of setting up the image processing system; when using the same number of GPUs, the throughput of the image processing system can be effectively improved, that is, the image processing capacity of the image processing system can be effectively improved.
[0019] Other features and aspects of the present application will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the application and, together with the description, serve to explain the principles of the application.
[0021] Figure 1 A schematic diagram of an application system provided according to an embodiment of the present application is shown.
[0022] Figure 2 A schematic diagram shows the relationship between the computing resource utilization g corresponding to each model instance and the throughput q of each model instance, and the relationship between the memory resource utilization m corresponding to each model instance and the throughput q of each model instance when a model service according to an embodiment of the present application has not reached the maximum daily throughput.
[0023] Figure 3 A schematic diagram showing the relationship between the computing resource utilization corresponding to each model instance of a model service and the number of model instances of the model service according to an embodiment of the present application.
[0024] Figure 4 A flow chart of a resource allocation method according to an embodiment of the present application is shown.
[0025] Figure 5A flowchart is shown for allocating the target model instance of each model service to the corresponding GPU based on the target computing resource utilization corresponding to each target model instance and the target memory resource utilization corresponding to each target model instance according to an embodiment of the present application.
[0026] Figure 6 A flowchart of a method for filtering target allocation model instances from a first set to form a second set based on target computing resource utilization and target memory resource utilization corresponding to the to-be-allocated model instances in the first set according to an embodiment of the present application is shown.
[0027] Figure 7 A flowchart is shown for filtering target allocation model instances from the first set to form a second set based on the target computing resource utilization and target memory resource utilization corresponding to the model instances to be allocated in the first set according to an embodiment of the present application.
[0028] Figure 8 A block diagram of a resource allocation device according to an embodiment of the present application is shown.
[0029] Figure 9 It is a block diagram showing a device 900 for resource allocation according to an exemplary embodiment. DETAILED DESCRIPTION
[0030] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0031] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0032] In addition, numerous specific details are provided in the detailed description below to better illustrate the present application. Those skilled in the art will appreciate that the present application can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present application.
[0033] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0034] In recent years, with the research and advancement of artificial intelligence technology, artificial intelligence technology has been widely used in many fields. The solution provided in the embodiments of this application involves model services, such as image processing model services, which can include image processing models and business communication functions. This is specifically illustrated by the following embodiments:
[0035] See also Figure 1 , Figure 1 FIG. 1 is a schematic diagram of an application system provided according to an embodiment of the present application. The application system can be used in the resource allocation method of the present application. Figure 1 As shown, the application system may at least include a server 01 and a terminal 02.
[0036] In an embodiment of the present application, the server 01 may include an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0037] In the embodiment of the present application, the terminal 02 may include a physical device such as a smartphone, a desktop computer, a tablet computer, a laptop computer, a smart speaker, a digital assistant, an augmented reality (AR) / virtual reality (VR) device, or a smart wearable device. A physical device may also include software running on the physical device, such as an application. The operating system running on the terminal 02 in the embodiment of the present application may include, but is not limited to, Android, iOS, Linux, Windows, and the like.
[0038] In the embodiments of this specification, the terminal 02 and the server 01 may be connected directly or indirectly via wired or wireless communication, which is not limited in this application.
[0039] Terminal 02 can be used to provide user-oriented resource allocation processing. Users can upload model instances to be allocated on Terminal 02. Terminal 02 can receive and display resource allocation results. Terminal 02 can provide user-oriented resource allocation processing in a manner including, but not limited to, an application, a webpage, etc.
[0040] It should be noted that, in the embodiment of the present application, the resource allocation method may be executed by the server 01. Preferably, the resource allocation method is implemented in the server 01, so as to reduce the data processing pressure of the terminal and improve the device performance of the user-oriented terminal.
[0041] Before introducing the resource allocation method, we first introduce the abstraction of the resource requirements of each model service in this application. The abstraction can be a triplet of each model service. The triplet of each model service can include<n,m,g> , the meaning of the triples can be as follows:
[0042]
[0043] The memory resource utilization corresponding to each model instance may refer to the proportion of GPU memory resources required by each model instance. The computing resource utilization corresponding to each model instance may refer to the proportion of GPU computing resources required by each model instance. Each model service can be deployed on at least two terminals, and the model service in each terminal can be referred to as a model instance of the model service. m and g corresponding to each model instance of the same model service can be the same.
[0044] Through the triplet of each model service, resources are divided into two types: compute resources and memory resources. Accordingly, GPU card resources can also include compute resources and memory resources (storage resources). This provides a foundation for deploying model instances with different resource requirements on the same GPU card, enabling a hybrid resource allocation solution and improving overall GPU card resource utilization.
[0045] Furthermore, in order to allocate GPU resources based on the triplet of each model service, it is necessary to first determine the values of n, m, and g. Therefore, it is necessary to first determine the relationship between n and m, and between n and g in the triplet of each model service. In an image processing system, generally, the upstream services accessed do not change frequently, and the maximum daily request volume (maximum daily throughput) of each model service in the image processing system will not change significantly. Therefore, the maximum daily throughput of each model service can be selected as the maximum resource demand Q of each model service. Since the maximum daily throughput of each model service can be regarded as unchanged, Q can be defined as a constant, so that the throughput of each model instance of each model service can be defined.
[0046] In the embodiments of this specification, the relationship between n and m, and between n and g in the triplet of each model service can be measured experimentally. For example, a packet sender can be used. When Q is not reached, the relationship between q and m, and between q and g can be measured by changing the request rate of the packet sender, such as Figure 2As shown, 0% to 80% can refer to computing resource utilization or video memory resource utilization. m is independent of q and is a constant, meaning that m is independent of n. g is proportional to q and can be described as: g = k × q. k here can be obtained based on this measurement.
[0047] Further, combined It can be deduced That is, when Q is a constant, the computing resource utilization rate corresponding to each model instance of a model service is inversely proportional to the number of model instances of the model service. Figure 3 It should be noted that k and Q here are for a certain model service, and k and Q for different model services can be different.
[0048] In the embodiment of this specification, the setting of computing resource utilization threshold is based on computing resource compression technology and The purpose is to minimize the resource competition conflicts between different model instances on the same GPU card. It can be concluded that by increasing n, g can be reduced, that is, by increasing n, the g corresponding to each model instance can be compressed, thereby reducing the proportion of GPU computing resources required for each model instance and reducing computing resource competition conflicts.
[0049] For example, assuming there is an image material fingerprint model, when n=1, a model instance of the image material fingerprint model may include: computing resource utilization rate of 91%, memory resource utilization rate of 7%; the computing resource compression of the image material fingerprint model is performed, for example, let n=2, the image material fingerprint model includes two model instances, based on the formula It can be found that each model instance has a computing resource utilization of 45.5% and a memory resource utilization of 7%. In other words, a model instance with a computing resource utilization of 91% and a memory resource utilization of 7% provides the same image processing capabilities as two model instances with a computing resource utilization of 45.5% and a memory resource utilization of 7%. Therefore, Compress the computing resource utilization of model instances, and convert a model instance with high computing resource utilization into at least two model instances with low computing resource utilization, while the service capacity provided by the model service remains unchanged.
[0050] Suppose there is another logo detection model with two corresponding model instances, each with a compute resource utilization of 46% and a memory resource utilization of 42%. Using the existing method of dedicating one model instance to a GPU, a model instance with a compute resource utilization of 91% and a memory resource utilization of 7%, and two model instances with a compute resource utilization of 46% and a memory resource utilization of 42% would require three GPU cards. However, if the model instance with a compute resource utilization of 91% and a memory resource utilization of 7% is compressed into two model instances with a compute resource utilization of 45.5% and a memory resource utilization of 7%, there are now four model instances: two with a compute resource utilization of 45.5% and a memory resource utilization of 7%, and two with a compute resource utilization of 46% and a memory resource utilization of 42%. In this way, after converting a model instance with high computing resource utilization into at least two model instances with lower computing resource utilization, a model instance with a computing resource utilization of 45.5% and a memory resource utilization of 7% and a model instance with a computing resource utilization of 46% and a memory resource utilization of 42% no longer compete for computing resources on the same GPU card, and can now share the same GPU card. This means that these four model instances can be allocated to two GPU cards. This shows that compressing computing resource utilization can improve the resource utilization of each GPU card. While maintaining the service capabilities of the image processing system, it can reduce the number of GPU cards required by the image SaaS system and lower system costs.
[0051] In actual applications, for the setting of the computing resource utilization threshold, actual tests have found that: the memory resource utilization of the model instances of more than 70% of the model services exceeds 33%. From the perspective of memory resource utilization, this means that two model instances can be accommodated on one GPU card. Therefore, you can choose to compress the computing resource utilization to less than or equal to 50%, so as to ensure that from the perspective of computing resource utilization, two model instances can be accommodated on one GPU card. Therefore, in one example, the computing resource utilization threshold can be set to 50%. Optionally, in order to increase the number of model instances accommodated on the GPU card, the computing resource utilization threshold can also be set to a lower value, such as 33%, which is not limited in this application.
[0052] Based on the above derivation That is, the preset relationship between the number of model instances in each model service triple and the computing resource utilization corresponding to each model instance, as well as the computing resource utilization threshold, can implement the resource allocation method of this application. It should be noted that the following is a possible sequence of steps and is not actually limited to this order. Some steps can be executed in parallel without relying on each other.
[0053] like Figure 4As shown, Figure 4 A flow chart of a resource allocation method according to an embodiment of the present application is shown. The method may include:
[0054] S401, obtaining a preset relationship between the number of model instances in a triple of multiple model services and the computing resource utilization corresponding to each model instance, and a computing resource utilization threshold.
[0055] In an embodiment of the present specification, when allocating model instances of multiple model services to a GPU, a preset relationship between the number of model instances in the triples of the multiple model services and the computing resource utilization corresponding to each model instance and a computing resource utilization threshold can be obtained. The preset relationship can be an inverse relationship, and the computing resource utilization threshold can be an upper limit of the computing resource utilization corresponding to each model instance. The multiple model services can refer to at least two model services, and the number of model services can be set according to the actual needs of the image processing system, which is not limited in this application.
[0056] S403 , determining the target number of model instances for each model service according to a preset relationship between the number of model instances in the triplet of each model service and the computing resource utilization corresponding to each model instance and the computing resource utilization threshold.
[0057] In the embodiment of this specification, for each model service, n can be substituted into the preset relationship The value of n that makes g less than or equal to the computing resource utilization threshold is selected as the target number of model instances for each model service. If there are multiple values of n that make g less than or equal to the computing resource utilization threshold, one of the values of n can be selected as the target number of model instances for each model service.
[0058] In a possible implementation, S403 may be implemented by the following steps:
[0059] The number of model instances in the preset relationship of each model service is traversed from small to large positive integers to obtain the initial number of model instances corresponding to each model service; wherein the value of the computing resource utilization corresponding to the initial number of model instances is less than or equal to the computing resource utilization threshold;
[0060] The minimum number of initial model instances corresponding to each model service is used as the target number of model instances for each model service.
[0061] By traversing n, the minimum number of initial model instances is selected as the target number of model instances for each model service. This ensures that when there is no contention between any two video memory resource utilizations, a GPU can be shared with the maximum target computing resource utilization, thereby ensuring the computing resource utilization of each GPU.
[0062] S405 : Using the value of the computing resource utilization corresponding to the number of target model instances of each model service as the target computing resource utilization corresponding to each target model instance of each model service.
[0063] In the embodiments of this specification, after the number of target model instances for each model service is determined, the number of target model instances can be substituted into the preset relationship of each model service, so that the target computing resource utilization corresponding to each target model instance of each model service can be determined.
[0064] S407: Obtain the target graphics memory resource utilization corresponding to each target model instance of each model service.
[0065] In the embodiment of this specification, the target video memory resource utilization corresponding to each target model instance of each model service, that is, the value of m corresponding to each target model instance of each model service, can be measured through the above-mentioned experimental method.
[0066] S409 , allocating the target model instance of each model service to the corresponding GPU based on the target computing resource utilization corresponding to each target model instance and the target memory resource utilization corresponding to each target model instance.
[0067] In the embodiments of this specification, target model instances of different model services can be assigned to the same GPU to achieve the purpose of sharing one GPU by target model instances of different model services. For example, a brute force search method can be used to assign the target model instance of each model service to the corresponding GPU. Alternatively, the target model instances of different model services can be combined in pairs, and then the target model instances of the two-by-two different model services can be assigned to one GPU, where the total computing resource utilization and total memory resource utilization of the target model instances of the two-by-two different model services are both less than 1. This application does not limit the specific method of assigning the target model instance of each model service to the corresponding GPU, as long as the GPU resource utilization can be improved.
[0068] By setting the triples of model services, the preset relationship between the number of model instances in the triples and the computing resource utilization corresponding to each model instance, and the computing resource utilization threshold, the target number of model instances for each model service, the target computing resource utilization corresponding to each target model instance, and the target video memory resource utilization are determined according to the preset relationship between the number of model instances in the triples of each model service and the computing resource utilization corresponding to each model instance, as well as the computing resource utilization threshold, so as to achieve compression of the computing resource utilization of the model instances, so that the GPU resource allocation based on the target computing resource utilization corresponding to each target model instance and the target video memory resource utilization corresponding to each target model instance can significantly improve the utilization of GPU computing resources and video memory resources, and the computing resource utilization can be improved by 93%; while keeping the service capacity of the image processing system unchanged, the number of GPUs required by the image processing system can be effectively reduced, thereby reducing the cost of setting up the image processing system; when using the same number of GPUs, the throughput of the image processing system can be effectively improved, that is, the image processing capacity of the image processing system can be effectively improved.
[0069] Figure 5 The flowchart of allocating the target model instance of each model service to the corresponding GPU based on the target computing resource utilization rate corresponding to each target model instance and the target memory resource utilization rate corresponding to each target model instance according to an embodiment of the present application is shown. Figure 5 As shown, in a possible implementation, S409 may include:
[0070] S501 : Taking the target model instance of each model service as the set of model instances to be allocated for each model service.
[0071] In the embodiments of this specification, the target model instance of each model service can be placed in a set of model instances to be allocated for use in subsequent GPU resource allocation.
[0072] S503 , extracting one model instance to be allocated from the set of model instances to be allocated of each model service to form a first set.
[0073] In the embodiments of this specification, in order to ensure that model instances of different model services share the same GPU, and also to avoid model instances of the same model service being allocated to the same GPU, one model instance to be allocated can be extracted from the set of model instances to be allocated of each model service to form a first set, so that the model instances in the first set can come from different model services.
[0074] S505, an empty GPU is used as the target GPU;
[0075] S507: Based on the target computing resource utilization and target memory resource utilization corresponding to the to-be-allocated model instances in the first set, filter target allocation model instances from the first set to form a second set, and clear the first set.
[0076] S509: Allocate the target allocation model instance in the second set to the target GPU.
[0077] In the embodiments of this specification, S503 to S509 can be regarded as a resource allocation process for a target GPU. The model instances to be allocated in the first set can be traversed, and at least one model instance to be allocated can be screened out from these model instances to be allocated as the target allocation model instance, so that the total computing resource utilization rate of the screened target allocation model instance and / or the total video memory resource utilization rate of the screened target allocation model instance is maximized. The target allocation model instances screened out from the first set can be combined into a second set to allocate the target allocation model instances in the second set to the target GPU. The first set can also be cleared to facilitate the use of the next target GPU for resource allocation.
[0078] S511 , removing the target allocation model instances in the second set from the set of model instances to be allocated for each model service, to obtain an updated set of model instances to be allocated for each model service.
[0079] In an embodiment of the present specification, the allocated target allocation model instances can be removed from the set of model instances to be allocated for each model service, that is, the target allocation model instances in the second set can be removed from the set of model instances to be allocated for each model service to obtain an updated set of model instances to be allocated for each model service, so as to ensure that the updated set of model instances to be allocated for each model service contains all model instances to be allocated.
[0080] S513: Based on the updated set of model instances to be allocated for each model service, repeat S503 to S511 above until the updated set of model instances to be allocated for each model service is empty. Specifically, a model instance to be allocated is extracted from the updated set of model instances to be allocated for each model service to form a first set. The process then proceeds to S505 to allocate resources to the next target GPU until the updated set of model instances to be allocated for each model service is empty.
[0081] In one example of the present specification, Figure 5 As shown, after S511, it can be determined whether the updated set of model instances to be allocated for each model service is empty. If so, the process can be terminated, indicating that the target model instances of each model service are allocated to the corresponding GPU. Here, whether the updated set of model instances to be allocated for each model service is empty can mean that the updated set of model instances to be allocated for each model service is empty.
[0082] If it is not empty, a model instance to be allocated can be extracted from the updated set of model instances to be allocated for each model service to form a first set, and then the process goes to S505 to allocate resources to the next target GPU until the updated set of model instances to be allocated for each model service is empty.
[0083] Figure 6 The flowchart of the method for filtering target allocation model instances from the first set to form a second set based on the target computing resource utilization and target memory resource utilization corresponding to the model instances to be allocated in the first set according to an embodiment of the present application is shown. Figure 6 As shown, in a possible implementation, S507 may include:
[0084] S601 : Taking the minimum value of the target computing resource utilization and the target video memory resource utilization corresponding to the to-be-allocated model instance in the first set as the target unit.
[0085] In the embodiments of this specification, for example, the first set of model instances to be allocated includes three: a, b, and c, with corresponding target computing resource utilizations of 0.5, 0.4, and 0.4; and corresponding target memory resource utilizations of 0.6, 0.5, and 0.4. A minimum of 0.4 can be determined as the target unit for partitioning the total computing resource utilization and total memory resource utilization of the target GPU. In this case, the target unit can optionally be directly selected as 0.1 to increase the granularity of partitioning the total computing resource utilization and total memory resource utilization of the target GPU.
[0086] It should be noted that, assuming that the minimum value of the target computing resource utilization and the target memory resource utilization corresponding to the model instance to be allocated in the first set is 0.35, then the target unit can be determined to be 0.01, that is, the minimum unit value corresponding to the minimum value can be used as the target unit.
[0087] S603 , dividing the total computing resource utilization of the target GPU based on the target unit to obtain a first resource state set of the computing resource utilization; the first resource state set includes at least one first resource state.
[0088] In an embodiment of the present specification, the total computing resource utilization of the target GPU may be 1. Based on the target unit, the total computing resource utilization of the target GPU is divided to obtain a first resource state set of computing resource utilization. The first resource state may include a first resource state in which the computing resource utilization is the target unit, a first resource state in which the computing resource utilization is 1, and may also include a first resource state in which the computing resource utilization is an integer multiple of the target unit. For example, assuming that the above-mentioned target unit is 0.4, the first resource state set includes at least one first resource state that may include 0.4, 0.8, and 1. Assuming that the above-mentioned target unit is 0.1, the first resource state set includes at least one first resource state that may include 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, and 1.
[0089] S605 , dividing the total memory resource utilization of the target GPU based on the target unit to obtain a second resource state set of the memory resource utilization; the second resource state set includes at least one second resource state.
[0090] In the embodiment of this specification, the total memory resource utilization of the target GPU may be 1. This step can be implemented with reference to the above S603 and will not be described in detail here; and S603 and S605 may be executed in parallel, which is not limited in this application.
[0091] S607 : Combine at least one first resource state and at least one second resource state in pairs to obtain a target resource state set of the target GPU. The target resource state set may include multiple target resource states.
[0092] In one embodiment of this specification, assuming a: 0.5, 0.6; b: 0.4, 0.5; and c: 0.4, 0.4, and taking the example that a first resource state set may include 0.4, 0.8, and 1, and a second resource state set may include 0.4, 0.8, and 1, the first resource state and the second resource state are combined to obtain a target resource state set for the target GPU. The target resource state set may include the following target resource states: (0.4, 0.4), (0.4, 0.8), (0.4, 1), (0.8, 0.4), (0.8, 0.8), (0.8, 1), (1, 0.4), (1, 0.8), and (1, 1).
[0093] S609: Determine a selection combination of model instances to be allocated in the first set.
[0094] In the embodiments of this specification, the first selection combination can be set as a model instance to be allocated according to the number of choices of the model instance to be allocated, and one choice can be added successively until the final selection combination includes the entire set of model instances to be allocated in the first set. For example, the model instances to be allocated in the first set include a, b, and c. The selection combinations for determining the model instances to be allocated in the first set may include: (b), (b, c), and (b, c, a).
[0095] S611, arrange multiple target resource states from small to large as columns of the resource state table, and use the selection combinations as rows of the resource state table; wherein the selection combination in the first row includes one model instance to be allocated, and the selection combination in each row has one more model instance to be allocated than the selection combination in the previous row.
[0096] In one example, the rows and columns of the resource status table may be as shown in Table 1:
[0097] Table 1
[0098]
[0099] The resource status table initially has only rows and columns. The selection combination of each row can represent the model instances to be allocated that can be selected in each row; the target resource status of each column can represent the upper limit of computing resource utilization and the upper limit of video memory resource utilization that can be accommodated by each column. Each column converts the total computing resource utilization and total video memory resource utilization of the target GPU into smaller computing resource utilization and video memory resource utilization.
[0100] S613 , starting from the first row of the resource status table, traverse the target resource status of each row, and determine the initial allocation model instance corresponding to each target resource status of each row from the selection combination of each row.
[0101] Taking (0.4, 0.8) in the second row (b, c) of Table 1 as an example, the sum of the target memory resource utilization of b and c is greater than 0.8, so b and c cannot be placed in (0.4, 0.8) at the same time. It can be determined that the initial allocation model instance corresponding to (0.4, 0.8) in the second row (b, c) can include: b or c.
[0102] S615: Determine the target computing resource utilization and target video memory resource utilization corresponding to the initial allocation model instance.
[0103] In the embodiment of this specification, if the initial allocation model instance includes a model instance to be allocated, the target computing resource utilization and target memory resource utilization corresponding to the model instance to be allocated may be determined as the target computing resource utilization and target memory resource utilization corresponding to the initial allocation model instance;
[0104] If the initial allocation model instance includes at least two model instances to be allocated, the sum of the target computing resource utilization rates corresponding to the two model instances to be allocated can be used as the target computing resource utilization rate corresponding to the initial allocation model instance; the sum of the target video memory resource utilization rates corresponding to the two model instances to be allocated can be used as the target video memory resource utilization rate corresponding to the initial allocation model instance.
[0105] S617 : Determine the dominant resource utilization of the initial allocation model instance based on the target computing resource utilization and the target video memory resource utilization corresponding to the initial allocation model instance.
[0106] In the embodiments of this specification, the dominant resource utilization rate can be the larger or smaller of the target computing resource utilization rate and the target memory resource utilization rate, the sum of the squares of the target computing resource utilization rate and the target memory resource utilization rate, or the sum of the target computing resource utilization rate and the target memory resource utilization rate, etc. This application is not limited to this. In this way, the two resource utilization rates: computing resource utilization rate and memory resource utilization rate, can be unified into a single dominant resource utilization rate, which is convenient for use in resource allocation.
[0107] The target computing resource utilization and the target memory resource utilization correspond to each other. Furthermore, the target computing resource utilization and the target memory resource utilization can be the target computing resource utilization and the target memory resource utilization corresponding to the initial allocation model instance, the target computing resource utilization and the target memory resource utilization corresponding to the to-be-allocated model instance, or the target computing resource utilization and the target memory resource utilization corresponding to the target combination.
[0108] Taking the example that the dominant resource utilization can be the smaller value of the target computing resource utilization and the target memory resource utilization, for example, the target computing resource utilization and the target memory resource utilization corresponding to the initial allocation model instance are 0.4 and 0.5 respectively, it can be determined that the dominant resource utilization of the initial allocation model instance is min(0.4, 0.5)=0.4.
[0109] S619, determining the initial allocation model instance corresponding to the maximum dominant resource utilization of each target resource state in each row as the target allocation model instance corresponding to each target resource state in each row, until the target allocation model instance corresponding to the last target resource state in the last row is determined.
[0110] For example, let's say the initial allocation model instance corresponding to the mth target resource state in row n is a+c;b+c. The dominant resource utilization of a+c is 0.9, and the dominant resource utilization of b+c is 0.8. Therefore, the target allocation model instance corresponding to the mth target resource state in row n is determined to be a+c.
[0111] As an example, by traversing Table 1, the target allocation model instance corresponding to the last resource status of the last row can be obtained. Assume a: 0.5, 0.6; b: 0.4, 0.5; c: 0.4, 0.4. Here, for each target resource status in each row, the space includes the target allocation model instance and the dominant resource utilization rate of this target allocation model instance.
[0112] As an example of traversing Table 1, it can be taken the target allocation model instance and the dominant resource utilization rate corresponding to the h-th row and the target resource status (i, j). As shown in Table 1, 1 ≤ h ≤ 3; 0.4 ≤ i ≤ 1; 0.4 ≤ j ≤ 1; h is a positive integer, and i and j are integer multiples of 0.4. If S(h, i, j) represents the target allocation model instance and the dominant resource utilization rate corresponding to the h-th row and the target resource status (i, j). To determine the target allocation model instance and the dominant resource utilization rate corresponding to S(h, i, j), it is necessary to compare the dominant resource utilization rate Z1 of S(h - 1, i, j) with the dominant resource utilization rate Z2 of <g h , j - m h ) + <g h , m h . If Z1 > Z2, S(h, i, j) = S(h - 1, i, j); if Z1 < Z2, take <g h , m h > + S(h - 1, i - g h , j - m h ) as S(h, i, j); if Z1 = Z2, it can be selected that S(h, i, j) = S(h - 1, i, j), or, select to take <g h , m h > + S(h - 1, i - g h , j - m h ) as S(h, i, j).
[0113] Taking S(2, 0.8, 0.8) as an example, to determine the target allocation model instance and the dominant resource utilization rate corresponding to this S(2, 0.8, 0.8), Z1 = 0.4, that is, the dominant resource utilization rate of S(1, 0.8, 0.8); <g h , m h > are the target computing resource utilization rate and the target video memory resource utilization rate of c: <0.4, 0.4>; S(h - 1, i - g h , j - m h ) + <g h , m h> is S(1, 0.4, 0.4)+<0.4, 0.4>, and S(1, 0.4, 0.4) corresponds to 0, as shown in Table 2. Therefore, Z2=0.4=min(0+0.4, 0+0.4), Z1=Z2, and the target allocation model instance corresponding to S(2, 0.8, 0.8) can be b or c; that is, the target allocation model instance corresponding to S(2, 0.8, 0.8) and the dominant resource utilization rate are b / 0.4 or c / 0.4.
[0114] After traversing Table 1 based on the above example, we can obtain the following Table 2.
[0115] Table 2
[0116]
[0117] It is determined that the target allocation model instances corresponding to the last target resource state in the last row include c and a, and the corresponding dominant resource utilization is the largest.
[0118] It should be noted that the granularity of the target resource state (i, j) is not fine enough, and there will be S(h-1,ig h , jm h ) is not in Table 2, then we can first determine S(h-1,ig h , jm h ) corresponding target allocation model instance, dominant resource utilization. Then calculate S(h-1,ig h , jm h )+ <g h , m h >. Based on this situation, in actual application, the target unit can be selected as 0.1 to divide the first resource state set and the second resource state set, so that the target resource state (i, j) can cover all state conditions. In this way, when traversing Table 1 to obtain Table 2, S(h-1,ig h , jm h ) will be present in both Table 1 and Table 2.
[0119] S621: Put the target allocation model instance corresponding to the last target resource state in the last row into the second set as the target allocation model instance.
[0120] In this embodiment of the present specification, the target allocation model instance corresponding to the last target resource state in the last row can be placed in the second set, and the target allocation model instances in the second set are allocated to the target GPU. As can be seen from Table 2, the target allocation model instance corresponding to the last target resource state in the last row has the highest resource utilization.
[0121] Figure 7The flowchart of the method for filtering target allocation model instances from the first set to form the second set based on the target computing resource utilization and target memory resource utilization corresponding to the model instances to be allocated in the first set according to an embodiment of the present application is shown. Figure 7 As shown, in a possible implementation, S507 may further include:
[0122] S701, putting the to-be-allocated model instances in the first set into the third set;
[0123] S703: Determine the dominant resource utilization corresponding to the model instances to be allocated in the third set based on the target computing resource utilization and target memory resource utilization corresponding to the model instances to be allocated in the third set. Specific implementation methods can be found in S617 and will not be described here in detail.
[0124] S705 : Transfer the to-be-allocated model instance with the largest dominant resource utilization in the third set to the empty second set, to obtain the current second set and the current third set.
[0125] In one example, assuming that the dominant resource utilization is the smaller of the target computing resource utilization and the target memory resource utilization corresponding to the model instance to be allocated, the dominant resource utilization corresponding to the model instance to be allocated in the third set can be obtained first, and then the model instance to be allocated with the largest dominant resource utilization in the third set can be transferred to the second set, resulting in the current second set and the current third set.
[0126] S707: Obtain a target combination consisting of each to-be-allocated model instance in the current third set and the current second set.
[0127] In the embodiments of this specification, each model instance to be assigned in the current third set can be combined with the current second set to form a target combination. For example, the current third set includes model instance 1 to be assigned and model instance 4 to be assigned; the current second set includes model instance 2 to be assigned and model instance 3 to be assigned; the target combinations formed by each model instance to be assigned in the current third set and the current second set are: (model instance 1 to be assigned, model instance 2 to be assigned, and model instance 3 to be assigned), (model instance 4 to be assigned, model instance 2 to be assigned, and model instance 3 to be assigned).
[0128] S709: Determine the target computing resource utilization and target video memory resource utilization corresponding to each target combination.
[0129] In the embodiments of this specification, the sum of the target computing resource utilizations corresponding to the to-be-allocated model instances in each combination can be calculated as the target computing resource utilization corresponding to each target combination; and the sum of the target video memory resource utilizations corresponding to the to-be-allocated model instances in each target combination can be calculated as the target video memory resource utilization corresponding to each target combination. For example, if the target combination is: (to-be-allocated model instance 1, to-be-allocated model instance 2, and to-be-allocated model instance 3), the target video memory resource utilization corresponding to this target combination = the target video memory resource utilization corresponding to to-be-allocated model instance 1 + the target video memory resource utilization corresponding to to-be-allocated model instance 2 + the target video memory resource utilization corresponding to to-be-allocated model instance 3.
[0130] S711: Delete from the current third set the model instances to be allocated in the target combination corresponding to the target computing resource utilization being greater than the total computing resource utilization of the target GPU, or the target video memory resource utilization being greater than the total video memory resource utilization of the target GPU. That is, based on the current second set, if the model instances to be allocated in the current third set would result in insufficient target GPU resources if placed on the target GPU, these model instances to be allocated that cannot be placed on the target GPU may be deleted from the current third set. Insufficient target GPU resources may include insufficient computing resources and / or insufficient video memory resources of the target GPU.
[0131] S713, determining the dominant resource utilization of the target combination corresponding to the target computing resource utilization being less than or equal to the total computing resource utilization of the target GPU and the target video memory resource utilization being less than or equal to the total video memory resource utilization of the target GPU;
[0132] S715, transfer the model instance to be allocated in the target combination corresponding to the maximum dominant resource utilization from the current third set to the current second set; go to S705, until the current third set is an empty set, and use the model instance to be allocated in the current second set as the target allocation model instance to form the second set.
[0133] In one example of the present specification, Figure 7 As shown, the to-be-allocated model instance in the target combination corresponding to the maximum dominant resource utilization is transferred from the current third set to the current second set. It can be determined whether the current third set is an empty set. If so, the process can be terminated, indicating that resource allocation for the target GPU has ended. If not, the process can proceed to S705 to continue resource allocation for the target GPU until the current third set is empty.
[0134] Figure 8 FIG. 1 shows a block diagram of a resource allocation device according to an embodiment of the present application. Figure 8 As shown, the device may include:
[0135] An acquisition module 801 is configured to acquire a preset relationship between the number of model instances in a triple of multiple model services and the computing resource utilization corresponding to each model instance, and a computing resource utilization threshold;
[0136] The target model instance number determination module 803 is used to determine the target model instance number for each model service based on a preset relationship between the number of model instances in the triples of each model service and the computing resource utilization corresponding to each model instance, and the computing resource utilization threshold; wherein the preset relationship is an inverse relationship, and the computing resource utilization threshold is an upper limit of the computing resource utilization corresponding to each model instance;
[0137] The target computing resource utilization determination module 805 is configured to use the computing resource utilization corresponding to the number of target model instances of each model service as the target computing resource utilization corresponding to each target model instance of each model service;
[0138] The target video memory resource utilization acquisition module 807 is used to acquire the target video memory resource utilization corresponding to each target model instance of each model service;
[0139] The GPU resource allocation module 809 is used to allocate the target model instance of each model service to the corresponding GPU based on the target computing resource utilization corresponding to each target model instance and the target memory resource utilization corresponding to each target model instance.
[0140] By setting the triples of model services, the preset relationship between the number of model instances in the triples and the computing resource utilization corresponding to each model instance, and the computing resource utilization threshold, the target number of model instances for each model service, the target computing resource utilization corresponding to each target model instance, and the target video memory resource utilization are determined according to the preset relationship between the number of model instances in the triples of each model service and the computing resource utilization corresponding to each model instance, as well as the computing resource utilization threshold, so as to achieve compression of the computing resource utilization of the model instances, so that the GPU resource allocation based on the target computing resource utilization corresponding to each target model instance and the target video memory resource utilization corresponding to each target model instance can significantly improve the utilization of GPU computing resources and video memory resources, and the computing resource utilization can be improved by 93%; while keeping the service capacity of the image processing system unchanged, the number of GPUs required by the image processing system can be effectively reduced, thereby reducing the cost of setting up the image processing system; when using the same number of GPUs, the throughput of the image processing system can be effectively improved, that is, the image processing capacity of the image processing system can be effectively improved.
[0141] In a possible implementation, the target model instance quantity determination module 803 may include:
[0142] an initial model instance quantity determination unit, configured to traverse the number of model instances in the preset relationship of each model service from small to large positive integers to obtain the initial model instance quantity corresponding to each model service; wherein the value of the computing resource utilization corresponding to the initial model instance quantity is less than or equal to the computing resource utilization threshold;
[0143] The target model instance quantity determination unit is used to use the minimum initial model instance quantity corresponding to each model service as the target model instance quantity of each model service.
[0144] In one possible implementation, the preset relationship between the number of model instances in the triples of multiple model services and the computing resource utilization corresponding to each model instance may include:
[0145] Among them, g is the computing resource utilization corresponding to each model instance of each model service; n is the number of model instances in the triplet of each model service; k is the constant corresponding to each model service; Q is the maximum resource demand of each model service.
[0146] In one possible implementation, the GPU resource allocation module 809 may include:
[0147] a to-be-allocated model instance set determining unit, configured to use the target model instance of each model service as the to-be-allocated model instance set of each model service;
[0148] A first set determining unit is configured to extract a to-be-allocated model instance from the to-be-allocated model instance set of each model service to form a first set;
[0149] A target GPU determination unit is used to set an empty GPU as a target GPU;
[0150] a second set generating unit, configured to filter target allocation model instances from the first set based on target computing resource utilization and target memory resource utilization corresponding to the to-be-allocated model instances in the first set, to form a second set, and clear the first set;
[0151] A target GPU resource allocation unit, configured to allocate the target allocation model instances in the second set to the target GPU;
[0152] a to-be-allocated model instance set updating unit, configured to remove the target allocation model instance in the second set from the to-be-allocated model instance set of each model service, to obtain an updated to-be-allocated model instance set of each model service;
[0153] The iterative unit is used to repeat the above steps of forming the first set to eliminating the target allocation model instance in the second set based on the updated set of model instances to be allocated for each model service, until the updated set of model instances to be allocated for each model service is empty.
[0154] In a possible implementation, the second set generating unit may include:
[0155] a target unit determination subunit, configured to use the minimum value of the target computing resource utilization and the target video memory resource utilization corresponding to the to-be-allocated model instance in the first set as the target unit;
[0156] A first resource state set acquisition subunit is configured to divide the total computing resource utilization of the target GPU based on the target unit to obtain a first resource state set of computing resource utilization; the first resource state set includes at least one first resource state;
[0157] A second resource status set acquisition subunit is configured to divide the total memory resource utilization of the target GPU based on a target unit to obtain a second resource status set of the memory resource utilization; the second resource status set includes at least one second resource state;
[0158] a target resource state set determination subunit, configured to combine at least one first resource state and at least one second resource state in pairs to obtain a target resource state set of a target GPU, the target resource state set including a plurality of target resource states;
[0159] A selection combination determination subunit, configured to determine a selection combination of model instances to be allocated in the first set;
[0160] A resource status table row and column generation subunit is configured to arrange a plurality of target resource states from smallest to largest as columns of the resource status table and to arrange selection combinations as rows of the resource status table; wherein the selection combination in the first row includes one model instance to be allocated, and the selection combinations in each row include one more model instance to be allocated than the selection combination in the previous row;
[0161] The resource status table traversal subunit is used to traverse the target resource status of each row starting from the first row of the resource status table, and determine at least one to-be-allocated model instance with the largest dominant resource utilization in the selection combination of each row as the target allocation model instance corresponding to each target resource status in each row, until the target allocation model instance corresponding to the last target resource status in the last row is determined;
[0162] The second set generating subunit is configured to put the target allocation model instance corresponding to the last target resource state in the last row into the second set as the target allocation model instance.
[0163] In a possible implementation, the second set generating unit may include:
[0164] A third set obtaining subunit, configured to put the to-be-allocated model instances in the first set into a third set;
[0165] The to-be-allocated model instance transfer subunit is used to transfer the to-be-allocated model instance with the largest dominant resource utilization in the third set to the empty second set, thereby obtaining the current second set and the current third set;
[0166] A target combination acquisition subunit is used to acquire a target combination consisting of each to-be-allocated model instance in the current third set and the current second set;
[0167] A target resource determination subunit is used to determine the target computing resource utilization and target video memory resource utilization corresponding to each target combination;
[0168] The to-be-allocated model instance deletion subunit is used to delete, from the current third set, the to-be-allocated model instance in the target combination corresponding to the target computing resource utilization being greater than the total computing resource utilization of the target GPU or the target video memory resource utilization being greater than the total video memory resource utilization of the target GPU;
[0169] The dominant resource utilization determination subunit of the target combination is used to determine the dominant resource utilization of the target combination corresponding to the target computing resource utilization being less than or equal to the total computing resource utilization of the target GPU and the target video memory resource utilization being less than or equal to the total video memory resource utilization of the target GPU;
[0170] The second set is composed of a sub-unit, which is used to transfer the model instance to be allocated in the target combination corresponding to the maximum dominant resource utilization from the current third set to the current second set, and to proceed to the step of obtaining the target combination until the current third set is an empty set, and the model instance to be allocated in the current second set is used as the target allocation model instance to form the above-mentioned second set.
[0171] Regarding the apparatus in the above embodiment, the specific manner in which each module and unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0172] In another aspect, the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the resource allocation method provided in the various optional implementations described above.
[0173] Figure 9 FIG. 9 is a block diagram showing a resource allocation apparatus 900 according to an exemplary embodiment. For example, the apparatus 900 may be provided as a server. Figure 9 The apparatus 900 includes a processing component 922, which further includes one or more processors, and memory resources represented by a memory 932 for storing instructions, such as applications, that can be executed by the processing component 922. The application stored in the memory 932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 922 is configured to execute the instructions to perform the above-described method.
[0174] The device 900 may also include a power supply component 926 configured to perform power management of the device 900, a wired or wireless network interface 950 configured to connect the device 900 to a network, and an input / output (I / O) interface 958. The device 900 may operate based on an operating system stored in the memory 932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.
[0175] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 932 including computer program instructions that can be executed by the processing component 922 of the apparatus 900 to perform the above-described method.
[0176] The present application may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present application.
[0177] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.
[0178] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0179] The computer program instructions for performing the operation of the present application can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or object code written in any combination of one or more programming languages, wherein the programming language includes object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or executed completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (such as by using an Internet service provider to connect to the Internet). In certain embodiments, by utilizing the state information of computer-readable program instructions to personalize electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLAs), the electronic circuits can execute computer-readable program instructions, thereby realizing various aspects of the present application.
[0180] Various aspects of the present application are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0181] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0182] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0183] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the system, method and computer program product according to multiple embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a special hardware-based system that performs the function or action of the specification, or can be implemented by a combination of special hardware and computer instructions.
[0184] The embodiments of the present application have been described above. The above description is illustrative and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.
Claims
1. A resource allocation method, characterized in that: include: Obtaining a preset relationship between the number of model instances in a triple of multiple model services and the computing resource utilization corresponding to each model instance, and a computing resource utilization threshold; The number of model instances in the preset relationship of each model service is traversed from small to large positive integers to obtain the initial number of model instances corresponding to each model service; wherein the value of the computing resource utilization corresponding to the initial number of model instances is less than or equal to the computing resource utilization threshold; the preset relationship of each model service is a preset relationship between the number of model instances in the triple of each model service and the computing resource utilization corresponding to each model instance, and the preset relationship is an inverse proportional relationship; Using one of the initial model instance numbers corresponding to each model service as the target model instance number for each model service; The value of the computing resource utilization corresponding to the number of target model instances of each model service is used as the target computing resource utilization corresponding to each target model instance of each model service; Get the target memory resource utilization corresponding to each target model instance of each model service; Based on the target computing resource utilization corresponding to each target model instance and the target memory resource utilization corresponding to each target model instance, the target model instance of each model service is allocated to the corresponding GPU.
2. The method according to claim 1, characterized in that The method of using one of the initial model instance quantities corresponding to each model service as the target model instance quantity for each model service includes: The minimum number of initial model instances among the initial number of model instances corresponding to each model service is used as the target number of model instances for each model service.
3. The method according to claim 1, characterized in that The preset relationship between the number of model instances in the triples of the multiple model services and the computing resource utilization corresponding to each model instance includes: ; Among them, g is the computing resource utilization corresponding to each model instance of each model service; n is the number of model instances in the triplet of each model service; k is the constant corresponding to each model service; Q is the maximum resource demand of each model service.
4. The method according to claim 1, wherein The method of allocating the target model instance of each model service to the corresponding GPU based on the target computing resource utilization corresponding to each target model instance and the target memory resource utilization corresponding to each target model instance includes: The target model instance of each model service is used as the set of model instances to be assigned to each model service; Extract one model instance to be assigned from the set of model instances to be assigned of each model service to form a first set; Use an empty GPU as the target GPU; Based on the target computing resource utilization and the target memory resource utilization corresponding to the to-be-allocated model instances in the first set, filter target allocation model instances from the first set to form a second set, and clear the first set; Allocate the target allocation model instances in the second set to the target GPU; Eliminate the target allocation model instances in the second set from the set of model instances to be allocated for each model service, to obtain an updated set of model instances to be allocated for each model service; Based on the updated set of model instances to be allocated for each model service, the steps of forming the first set to eliminating the target allocation model instances in the second set are repeated until the updated set of model instances to be allocated for each model service is empty.
5. The method according to claim 4, characterized in that The target allocation model instances are screened out from the first set based on the target computing resource utilization and the target memory resource utilization corresponding to the to-be-allocated model instances in the first set to form a second set, including: The minimum value of the target computing resource utilization and the target video memory resource utilization corresponding to the to-be-allocated model instance in the first set is used as the target unit; Dividing the total computing resource utilization of the target GPU based on the target unit to obtain a first resource state set of computing resource utilization; the first resource state set includes at least one first resource state; Dividing the total memory resource utilization of the target GPU based on the target unit to obtain a second resource state set of the memory resource utilization; the second resource state set includes at least one second resource state; Combining the at least one first resource state and the at least one second resource state in pairs to obtain a target resource state set of the target GPU, where the target resource state set includes a plurality of target resource states; determining a selection combination of model instances to be allocated in the first set; Arrange the multiple target resource states from smallest to largest as columns of a resource state table, and use the selection combinations as rows of the resource state table; wherein the selection combinations in the first row include one model instance to be allocated, and the selection combinations in each row include one more model instance to be allocated than the selection combinations in the previous row; Starting from the first row of the resource status table, traverse the target resource status of each row, and determine the initial allocation model instance corresponding to each target resource status of each row from the selection combination of each row; Determining a target computing resource utilization and a target video memory resource utilization corresponding to the initial allocation model instance; Determining a dominant resource utilization of the initial allocation model instance based on a target computing resource utilization and a target video memory resource utilization corresponding to the initial allocation model instance; Determine the initial allocation model instance corresponding to the maximum dominant resource utilization of each target resource state in each row as the target allocation model instance corresponding to each target resource state in each row, until the target allocation model instance corresponding to the last target resource state in the last row is determined; The target allocation model instance corresponding to the last target resource state in the last row is placed into the second set.
6. The method according to claim 4, characterized in that The target allocation model instances are screened out from the first set based on the target computing resource utilization and the target memory resource utilization corresponding to the to-be-allocated model instances in the first set to form a second set, including: Putting the to-be-allocated model instances in the first set into a third set; Determine the dominant resource utilization corresponding to the model instance to be allocated in the third set based on the target computing resource utilization and the target memory resource utilization corresponding to the model instance to be allocated in the third set; Transferring the to-be-allocated model instance with the largest dominant resource utilization in the third set to the empty second set, thereby obtaining a current second set and a current third set; Obtain a target combination consisting of each to-be-assigned model instance in the current third set and the current second set; Determine the target computing resource utilization and target memory resource utilization corresponding to each target combination; Delete the to-be-allocated model instances in the target combinations corresponding to target computing resource utilization being greater than the total computing resource utilization of the target GPU or target video memory resource utilization being greater than the total video memory resource utilization of the target GPU from the current third set; Determine the dominant resource utilization of the target combination corresponding to the target computing resource utilization being less than or equal to the total computing resource utilization of the target GPU and the target video memory resource utilization being less than or equal to the total video memory resource utilization of the target GPU; The model instance to be allocated in the target combination corresponding to the maximum dominant resource utilization is transferred from the current third set to the current second set, and the step of obtaining the target combination is performed until the current third set is an empty set. The model instance to be allocated in the current second set is used as the target allocation model instance to form the second set.
7. The method according to claim 5 or 6, characterized in that The dominant resource utilization is the smaller value between the target computing resource utilization and the corresponding target video memory resource utilization.
8. A resource allocation device, characterized in that: include: An acquisition module, configured to obtain a preset relationship between the number of model instances in a triple of multiple model services and the computing resource utilization corresponding to each model instance, and a computing resource utilization threshold; a target model instance number determination module, configured to determine the target model instance number for each model service based on a preset relationship between the number of model instances in the triplet of each model service and the computing resource utilization rate corresponding to each model instance and the computing resource utilization rate threshold; a target computing resource utilization determination module, configured to use the computing resource utilization value corresponding to the number of target model instances of each model service as the target computing resource utilization corresponding to each target model instance of each model service; The target video memory resource utilization acquisition module is used to obtain the target video memory resource utilization corresponding to each target model instance of each model service; A GPU resource allocation module is used to allocate the target model instance of each model service to the corresponding GPU based on the target computing resource utilization corresponding to each target model instance and the target memory resource utilization corresponding to each target model instance; The target model instance quantity determination module includes: an initial model instance quantity determination unit, configured to traverse the number of model instances in the preset relationship of each model service from small to large positive integers to obtain the initial number of model instances corresponding to each model service; wherein the value of the computing resource utilization corresponding to the initial number of model instances is less than or equal to the computing resource utilization threshold; the preset relationship of each model service is a preset relationship between the number of model instances in the triple of each model service and the computing resource utilization corresponding to each model instance, and the preset relationship is an inverse proportional relationship; The target model instance quantity determination unit is used to use an initial model instance quantity among the initial model instance quantities corresponding to each model service as the target model instance quantity for each model service.
9. A resource allocation device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions to implement the method according to any one of claims 1 to 7.
10. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, cause a computer to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for service resource allocation and electronic equipment
CN110033148A
System resource scheduling method and device, machine readable medium and system
CN111158879A