Graphic processor resource management system, method and server
By introducing an identification mechanism into the graphics processor resource management system and rationally selecting sharable or non-sharable GPUs, the problem of low multi-instance GPU resource utilization is solved, achieving more efficient GPU resource allocation and cost optimization.
Patent Information
- Application Number
- CN202511271925.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-08
AI Technical Summary
In the existing technology, multi-instance GPU partitioning makes it difficult to fully utilize GPU resources, especially in small model inference services, which are difficult to flexibly partition and expand on demand.
By introducing the first identifier, the second identifier, and the third identifier into the graphics processor resource management system, the inference engine module reasonably selects a sharable or non-sharable GPU based on these identifiers and the GPU occupancy status information, ensuring the video memory requirements of the inference instance and optimizing resource allocation.
This enables multiple inference instances to more reasonably share the same GPU resources, improves GPU utilization and resource efficiency, reduces costs, and ensures service stability and reliability.
Smart Images

Figure CN120765447A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer processors, in particular to a graphics processor resource management system, method and server. BACKGROUND
[0002] With the rapid development of artificial intelligence (AI) and big data, using machine learning technology to train business models and using trained business models to realize intelligent processing of big data business has gradually become a general means in the big data industry. For a specific business model, after training the business model in a specified training environment, the business model is deployed as an online inference service, which can be used when a user uses the same running environment as the training environment.
[0003] The execution of the inference service depends on graphics processor (GPU) resources. At present, in order to improve the resource utilization of the GPU, multi-instance GPU (MIG) division is performed, and the plurality of "virtual GPUs" (GPU instances) obtained by the division are provided to a plurality of inference services. However, due to the mechanism of MIG, the MIG division is usually equal division, which makes it difficult to fully utilize the GPU resources. SUMMARY
[0004] Therefore, the present application provides a graphics processor resource management system, method and server to solve the problem that a plurality of inference services cannot fully utilize GPU resources.
[0005] In a first aspect, the present application provides a graphics processor resource management system, which includes an inference engine module and an inference service module. The inference engine module and the inference service module are in communication connection with a graphics processor resource pool. The graphics processor resource pool includes a plurality of graphics processors. The graphics processors are configured with a first identifier, which is used to indicate whether the graphics processors can be shared by a plurality of inference instances. The inference service module is used to generate a message request indicating the start of a target inference instance, and is used to send the message request to the inference engine module. The message request includes a second identifier and a third identifier. The second identifier is used to indicate whether the target inference instance occupies a shareable graphics processor. The third identifier is used to indicate the target memory value required by the target inference instance after starting. The inference engine module is used to determine a target graphics processor from the plurality of graphics processors according to the first identifier, the second identifier, the third identifier and the occupancy state information of the plurality of graphics processors, and is used to send a response message to the inference service module. The response message includes indication information of the target graphics processor. The inference service module is further used to mount the target inference instance to the target graphics processor according to the response message.
[0006] The inference engine module configures a first identifier for the GPU, and the inference service module adds a second identifier and a third identifier in a message request sent to the inference engine module, so that the inference engine module can select a more suitable GPU for the inference instance, multiple inference instances can more reasonably occupy the same GPU, and GPU resources are used in a refined manner to reduce costs. At the same time, the inference engine module and the inference service module cooperate with each other, which can also improve the GPU resource utilization efficiency and optimize the resource energy consumption.
[0007] In an optional implementation, the inference engine module is configured to, when the second identifier indicates that the target inference instance occupies a shareable GPU, select at least one first GPU from the plurality of GPUs according to the first identifier, the first GPU being a shareable GPU; the inference engine module is further configured to determine whether a remaining memory value of the at least one first GPU is greater than a target memory value according to the occupancy state information of the at least one first GPU; and the inference engine module is further configured to determine one of the at least one second GPU as the target GPU, the second GPU being the first GPU whose remaining memory value is greater than the target memory value.
[0008] In an optional implementation, the inference engine module is configured to determine the second GPU whose difference between the remaining memory value and the target memory value is greater than a preset memory value as the target GPU.
[0009] In this embodiment, the second GPU whose difference between the remaining memory value and the target memory value is greater than a preset memory value is determined as the target GPU, so that the target GPU still has a resource space after mounting the target inference instance, preventing memory overflow and GPU crash, and ensuring service stability and reliability.
[0010] In an optional implementation, in the case where there is no first GPU or there is no second GPU, the inference engine module is further configured to select at least one third GPU from the plurality of GPUs according to the occupancy state information of the plurality of GPUs and the first identifier, the third GPU being an unshareable and unoccupied GPU; and the inference engine module is further configured to determine one of the at least one third GPU as the target GPU.
[0011] In an optional implementation, the inference service module is further configured to update the occupancy state information of the target GPU after the target inference instance is successfully started, and to synchronize the updated occupancy state information of the target GPU and information generated by the target inference instance to the inference engine module.
[0012] In an optional implementation, the inference engine module is further configured to update the first identifier corresponding to the target graphics processor to change the target graphics processor to the sharable graphics processor after determining the target graphics processor.
[0013] In an optional implementation, the inference engine module is further configured to, when the second identifier indicates that the target inference instance occupies the unshareable graphics processor, select at least one third graphics processor from the plurality of graphics processors according to the occupation state information of the plurality of graphics processors and the first identifier; and determine, as the target graphics processor, the third graphics processor with a display memory value greater than the target display memory value.
[0014] In an optional implementation, the plurality of graphics processors are deployed on different servers.
[0015] In a second aspect, the present application provides a graphics processor resource management method, which is applied to the inference engine module of the graphics processor resource management system in the first aspect or any of the corresponding implementations, and includes: receiving a packet request from an inference service module, the packet request including a second identifier and a third identifier, the second identifier being used to indicate whether a target inference instance occupies a sharable graphics processor, and the third identifier being used to indicate a target display memory value required by the target inference instance after being started; determining a target graphics processor from a plurality of graphics processors according to a first identifier, the second identifier, the third identifier and occupation state information of the plurality of graphics processors, the first identifier being used to indicate whether a graphics processor can be shared by a plurality of inference instances; and sending a response message to the inference service module, the response message including indication information of the target graphics processor, so as to enable the target inference instance to be mounted to the target graphics processor.
[0016] The graphics processor resource management method provided in this embodiment increases three identifier fields for configuration when starting an inference instance, and the combination facilitates the new creation of inference instances and the business, so that a plurality of inference instances can share the same GPU, and the GPU resources are used more efficiently.
[0017] In a third aspect, the present application provides a server, which includes the graphics processor resource management system in the first aspect or any of the corresponding implementations. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the specific embodiments or related art, the following will briefly introduce the drawings needed to be used in the specific embodiments or related art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0019] Figure 1 is a structural schematic diagram of a graphics processor resource management system and a graphics processor resource pool according to an embodiment of the present application; Figure 2 is a schematic diagram of an interactive process of a graphics processor resource management system according to an embodiment of the present application; Figure 3 is a flowchart of a graphics processor resource management method according to an embodiment of the present application; Figure 4 is a flowchart of another graphics processor resource management method according to an embodiment of the present application; Figure 5 is a structural schematic diagram of a server according to an embodiment of the present application. DETAILED DESCRIPTION
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative work based on the embodiments in the present application fall within the scope of protection of the present application.
[0021] To facilitate understanding of the present application, the related terms involved in the present application are introduced before the technical solutions of the present application are described.
[0022] (1) GPU GPU is a hardware device specially used for processing graphics and image calculations, and is widely used in graphics rendering, game development, video processing, scientific calculation, machine learning, etc. A GPU card usually contains an independent graphics processing chip, as well as related video memory and external interfaces. Due to the powerful parallel processing capability of GPU, compared with the central processing unit (CPU), the GPU card can process a large amount of data at the same time and perform high-speed parallel calculation, so it is used to accelerate various computationally intensive tasks.
[0023] (2) Inference service The process of generating information (Response) that the user wants to obtain by using a trained AI large model and according to the prompt (Prompt) input by the user is called "model inference" (Model Inference). That is, the inference service is the process of deploying a trained model online for real-time prediction.
[0024] With the rapid development of AI technology, the demand for compute-intensive tasks such as deep learning and neural networks has increased significantly. To meet these demands, specialized accelerators such as GPUs have been introduced into computing hardware design. However, efficiently utilizing a single GPU for inference services for small models has become a challenge. Small models refer to AI models with fewer parameters and lower computational requirements. Small models typically occupy less than 10GB of GPU memory.
[0025] As mentioned in the background, GPUs can be partitioned using MIG (Multi-Integrated Graphics) (MIG) to enable a single GPU to support inference services for multiple small models. However, this approach suffers from issues such as difficulty flexibly partitioning on demand, limited support for NVIDIA cards, and inconvenient expansion of MIG partitioning before models start inference services.
[0026] In view of this, the present invention provides a graphics processor resource management system, method and server, which manages multiple inference instances and GPUs by identification, so that multiple inference instances can more reasonably occupy the resources of the same GPU, thereby improving GPU utilization.
[0027] The graphics processor resource management system provided by the present invention is described in detail below with reference to the accompanying drawings.
[0028] like Figure 1 As shown, the graphics processor resource management system 100 includes an inference engine module 110 and an inference service module 120. The inference engine module 110 and the inference service module 120 are both communicatively connected to the graphics processor resource pool 200, and the inference engine module 110 and the inference service module 120 are also communicatively connected to each other.
[0029] The GPU resource pool 200 includes multiple GPUs, which can be deployed on different servers or on the same server. Specifically, the GPU resource pool 200 can include a server with multiple GPUs deployed on it, or it can include multiple servers, each of which has one or more GPUs deployed on it.
[0030] Figure 1 Take the example of a graphics processor resource pool 200 including two servers (referred to as a first server 1 and a second server 2), but the present invention is not limited thereto. Figure 1 In the example, each server includes four GPUs, the first server 1 includes GPUs 11 to 14 , and the second server 2 includes GPUs 21 to 24 .
[0031] Specifically, each GPU is configured with a first identifier, the first identifier is used to indicate whether the GPU can be shared by multiple inference instances. For example, the first identifier can be a field Gpu Is Share. When the field Gpu Is Share is 1, it indicates that the GPU can be shared by multiple inference instances (inference models). When the field Gpu Is Share is 0, it indicates that the GPU cannot be shared by multiple inference instances, and is generally exclusively used by a single inference instance.
[0032] The inference engine module 110 can configure the first identifier when the server initializes the GPU. The inference engine module 110 can determine whether the GPU can be shared according to the GPU memory value and / or business requirements. For example, the inference engine module 110 can determine that a GPU with a memory value greater than a preset value (e.g., 10G) is a shareable GPU. The inference engine module 110 can determine that all GPUs except for specific business designated GPUs are shareable GPUs.
[0033] The inference service module 120 is configured to generate a message request indicating the start of a target inference instance, and to send the message request to the inference engine module 110. Specifically, the inference service module 120 generates the message request in response to a user operation. The inference service module 120 includes multiple inference instances. The target inference instance can be one of the inference instances selected by the user from the multiple inference instances.
[0034] The message request includes a second identifier and a third identifier. The second identifier is used to indicate whether the target inference instance occupies a shareable GPU. The third identifier is used to indicate the memory size (target memory value) required by the target inference instance after being started.
[0035] For example, the second identifier can be a field GPU Utlization Share Enabled. The third identifier can be a field Gpu Memory Use. When the field GPU Utlization Share Enabled is 1, it indicates that a shareable GPU is preferred for starting the target inference instance. When the field GPU Utlization Share Enabled is 0, it indicates that an exclusive GPU is selected for starting the target inference instance. The value (INT value) corresponding to the field Gpu Memory Use is the target memory value. For example, the target memory value can be 5G, 10G, or 20G, etc.
[0036] The inference engine module 110 is configured to determine a target graphics processor from the plurality of GPUs according to the first identifier, the second identifier, the third identifier and the occupation state information of the plurality of GPUs after receiving the message request. The occupation state information includes information about whether each GPU is occupied and a GPU occupied memory value. The inference engine module 110 can interact with the servers in the graphics processor resource pool 200 to collect the occupation state information of the plurality of GPUs. The target graphics processor is a GPU selected by the inference engine module 110 from the plurality of GPUs to meet the running requirements of the target inference instance.
[0037] For example, after obtaining the occupation state information of the plurality of GPUs, the inference engine module 110 can synchronize the occupation state information of the plurality of GPUs to the inference service module 120.
[0038] After determining the target graphics processor, the inference engine module 110 is further configured to send a response message to the inference service module 120, wherein the response message includes indication information of the target graphics processor. After receiving the response message, the inference service module 120 is further configured to mount the target inference instance to the target graphics processor according to the response message.
[0039] The graphics processor resource management system provided in the embodiment can configure the first identifier for the GPU by the inference engine module 110, and add the second identifier and the third identifier in the message request sent by the inference service module 120 to the inference engine module 110, so that the inference engine module 110 can select more suitable GPUs for the plurality of inference instances. The plurality of inference instances can more reasonably occupy the same GPU, finely use the GPU resources, and reduce the cost. At the same time, the inference engine module 110 and the inference service module 120 can also cooperate to improve the GPU resource utilization process and optimize the resource energy consumption.
[0040] Specifically, when the second identifier indicates that the target inference instance occupies a shareable GPU, i.e., GPU_Utlization_Share_Enabled = 1, the inference engine module 110 is configured to first screen at least one first graphics processor from the plurality of GPUs according to the first identifier, then determine whether the remaining memory value of the at least one first graphics processor is greater than the target memory value according to the occupation state information of the at least one first graphics processor, and finally determine one of the at least one second graphics processor as the target graphics processor. The first graphics processor is a shareable graphics processor, and the second graphics processor is a first graphics processor with a remaining memory value greater than the target memory value.
[0041] For example, Figure 1GPU 11, GPU 13 and GPU 21 in the GPU 11, GPU 12, GPU 13, GPU 21, GPU 22, GPU 23 and GPU 24 are shareable GPUs, and GPU 12, GPU 14, GPU 22, GPU 23 and GPU 24 are non-shareable GPUs. On this basis, if the inference engine module 110 receives a message request and GPU_Utlization_Share_Enabled = 1, the remaining memory values of GPU 11, GPU 13 and GPU 21 are determined. If the remaining memory values of GPU 11 and GPU 21 are greater than the target memory value, GPU 11 or GPU 21 can be directly determined as the target GPU.
[0042] In one example, when the number of determined second GPUs is multiple, the target GPU is preferentially selected from the second GPUs in the occupied state.
[0043] In this embodiment, preferentially selecting the occupied second GPU as the target GPU can keep the occupied GPU at a high utilization rate and fully exert the computing performance of the occupied GPU.
[0044] Further, the inference engine module 110 is configured to determine the second GPU with a difference between the remaining memory value and the target memory value greater than a preset memory value as the target GPU. The preset memory value can be determined based on the total memory value of the GPU. For example, the preset memory value can be 5% of the total memory value, that is, the sum of the GPU memory occupancy rates of each inference instance needs to be less than 95%.
[0045] For example, after determining that the remaining memory values of GPU 11 and GPU 21 are greater than the target memory value, the difference between the remaining memory value and the target memory value of GPU 11 and the difference between the remaining memory value and the target memory value of GPU 21 are determined. If the difference between the remaining memory value and the target memory value of GPU 21 is greater than the preset memory value, GPU 21 is determined as the target GPU.
[0046] In this embodiment, the second GPU with a difference between the remaining memory value and the target memory value greater than the preset memory value is determined as the target GPU, which can leave a resource space for the target GPU after mounting the target inference instance, prevent memory overflow and GPU crash, and ensure service stability and reliability.
[0047] In other embodiments, when the second identifier indicates that the target inference instance occupies a shareable GPU, the inference engine module 110 can also be used to first select a GPU whose remaining video memory value is greater than the target video memory value from multiple GPUs based on the occupancy status information of the multiple GPUs, and then be used to select a shareable GPU from the GPUs whose remaining video memory value is greater than the target video memory value based on the first identifier, and finally determine the shareable GPU as the target graphics processor.
[0048] Specifically, when the first graphics processor or the second graphics processor does not exist, the inference engine module 110 is further configured to select at least one third graphics processor from the multiple graphics processors based on the occupation status information of the multiple GPUs and the first identifier, and to determine one of the at least one third graphics processor as the target graphics processor. The third graphics processor is a non-shareable and unoccupied GPU.
[0049] That is, if the inference engine module 110 cannot determine the first graphics processor or the second graphics processor, an unoccupied and non-shareable GPU among the multiple GPUs is determined as the target graphics processor.
[0050] Furthermore, after the inference engine module 110 cannot determine the first graphics processor or the second graphics processor and determines a third graphics processor as the target graphics processor, the inference engine module 110 is also used to update the first identifier corresponding to the target graphics processor to change the target graphics processor to a shareable GPU.
[0051] For example, Figure 1 GPUs 11 to 14 and GPUs 22 to 24 are non-shareable GPUs, and GPU 21 is a shareable GPU but the remaining video memory value does not meet the conditions (is less than the target video memory value). At this time, the unoccupied GPU 11 can be determined as the target graphics processor, and after determination, the first identifier in GPU 11 is updated, such as updating the value 0 in the first identifier to the value 1.
[0052] Exemplarily, the inference service module is further used to update the occupancy status information of the target graphics processor after the target inference instance is successfully started, and to synchronize the updated occupancy status information of the target graphics processor and the information generated by the target inference instance to the inference engine module.
[0053] Specifically, when the second identifier indicates that the target inference instance occupies a non-shareable GPU, the inference engine module 110 is used to filter out at least one third graphics processor from multiple GPUs based on the occupancy status information of multiple GPUs and the first identifier, and to determine the third graphics processor whose video memory value is greater than the target video memory value as the target graphics processor.
[0054] For example, Figure 1 GPU 11, GPU 13 and GPU 21 are shareable GPUs, and GPU 12, GPU 14, GPU 22, GPU 23 and GPU 24 are non-shareable GPUs. On this basis, if the inference engine module 110 receives a message request and GPU_Utlization_Share_Enabled=0, the occupancy status information of GPU 12, GPU 14, GPU 22, GPU 23 and GPU 24 is determined; if GPU 12, GPU 14 and GPU 22 are not occupied, it is determined whether the video memory values of GPU 12, GPU 14 and GPU 22 are greater than the target video memory value; if the video memory values of GPU 12 and GPU 22 are greater than the target video memory value, GPU 12 or GPU 22 can be directly determined as the target graphics processor.
[0055] Specifically, the inference engine module 110 is further configured to update the occupancy status information and inference service status of the target GPU after the target inference instance releases the target GPU.
[0056] The following is combined with Figure 2 In the scenario where the first inference instance instance1 is mounted on the non-shareable GPU 12 and the second inference instance instance2 is mounted on the shareable GPU 21, the processing flow of the graphics processor resource management system after the third inference instance instance3 is triggered is described.
[0057] Specifically, when the third inference instance instance3 of a single small model is started, it can support the use of shared GPU card resources, and the graphics card of this model requires the size of AG video memory resources.
[0058] When inference service module 120 starts the third inference instance, instance3, it configures the parameters GPU_Utlization_Share_Enabled = 1 and Gpu_Memory_Use = A. Inference service module 120 is also configured to generate a message request that carries the GPU_Utlization_Share_Enabled and Gpu_Memory_Use field information. Inference service module 120 can effectively manage inference services by using a stateful service (StatefuleSet) to identify an inference instance service, ensuring the stability and availability of the inference service.
[0059] At the same time, inference engine module 110 interacts with GPU resource pool 200 to monitor the activated inference instances (e.g., first inference instance instance1 and second inference instance instance2) and the occupancy status of multiple GPUs. Upon receiving the message request, inference engine module 110 processes it differently in different scenarios.
[0060] Scenario 1: There are currently no supported shared GPU card resources.
[0061] Specifically, after receiving the message request for starting the third inference instance 3, the inference engine module 110 queries whether there is a GPU card that meets the requirements in the current environment. If there is no GPU card resource with Gpu_Is_Share=1 in the current environment, the inference engine module 110 synchronizes the currently unshared GPU card resources to the third inference instance 3 and actively synchronizes a GPU card that can support exclusive use to supply the third inference instance 3. After receiving the reply information from the inference engine module 110, the inference service module 120 actively mounts the third inference instance 3 to the exclusive GPU resource provided by the inference engine module 110 and starts the third inference instance 3. At the same time, the inference engine module 110 changes the status of the GPU card to a shareable state. After the third inference instance 3 is successfully started, the inference service module 120 synchronizes the information of the third inference instance 3 and the corresponding occupancy status information of the GPU card to the inference engine module 110.
[0062] Scenario 2: There are GPU card resources that support sharing, but the computing power of the GPU card does not meet the requirements.
[0063] Specifically, after receiving the message request of starting the third inference instance instance3, the inference engine module 110 queries whether there is a GPU card meeting the requirement in the current environment. The current environment has a GPU card resource with Gpu Is Share = 1. At this time, the inference engine module 110 judges whether the existing shared GPU resource is greater than the GPU Memory USE field (i.e. AG). If neither meets the requirement, the inference engine module 110 synchronizes to the third inference instance instance3 that although there is a shareable GPU card resource at present, the computing power cannot meet the requirement. Subsequently, the inference engine module 110 actively synchronizes to the inference service module 120 a GPU resource that can support exclusive supply for the third inference instance instance3. After receiving the reply information of the inference engine module 110, the inference service module 120 actively mounts the third inference instance instance3 to the exclusive GPU resource, and starts the third inference instance instance3. After the third inference instance instance3 is successfully started, the inference instance information and the corresponding occupation state information of the GPU card are synchronized to the inference engine module 110. Moreover, the inference engine module 110 changes the state of the GPU card to a shareable state.
[0064] Scenario 3: There is a shareable GPU card resource at present, and the computing power meets the starting requirement.
[0065] Specifically, after receiving the message request of starting the third inference instance instance3, the inference engine module 110 queries whether there is a GPU card meeting the requirement in the current environment. The current environment has a GPU card resource with Gpu Is Share = 1. At this time, the inference engine module 110 compares whether the existing shared GPU resource is greater than the GPU Memory USE field. If the current computing power meets the starting requirement of the third inference instance instance3, the inference engine module 110 synchronizes to the third inference instance instance3 that there is a shareable GPU card resource and computing power at present. After receiving the reply information of the inference engine module 110, the inference service module 120 actively mounts the third inference instance instance3 to the shareable GPU selected by the inference engine module 110, and starts the third inference instance instance3, Figure 2 Taking mounting the third inference instance instance3 to the GPU 21 as an example. After the third inference instance instance3 is successfully started, the inference service module 120 synchronizes the inference instance information and the corresponding occupation state information of the GPU card to the inference engine module 110.
[0066] It should be noted that the same strategy is used to process the subsequent inference instance startup process. At the same time, when multiple inference instances share the same GPU, the sum of the GPU memory occupancy rates of each inference service is less than 95%, and some resources are reserved to avoid other problems. The inference engine module 110 needs to count the state information and specific information of multiple inference instances sharing a single GPU card.
[0067] After the inference instance releases the GPU card and the computing power of the GPU card, the inference engine module 110 refreshes the state and computing power of the current GPU card and the inference service at the same time.
[0068] In this embodiment, a graphics processor resource management method is also provided, which can be used in the inference engine module, Figure 3 is a flowchart of a graphics processor resource management method according to an embodiment of the application, as Figure 3 shown, the method comprises the following steps: Step S301, receiving a packet request for indicating the startup of a target inference instance from an inference service module.
[0069] The packet request includes a second identifier and a third identifier. The second identifier is used to indicate whether the target inference instance occupies a shareable graphics processor. The third identifier is used to indicate the target memory value required by the target inference instance after startup.
[0070] Step S302, determining a target graphics processor from the plurality of graphics processors according to the first identifier, the second identifier, the third identifier, and the occupancy state information of the plurality of graphics processors.
[0071] The first identifier is used to indicate whether the graphics processor can be shared by multiple inference instances.
[0072] Step S303, sending a response message to the inference service module.
[0073] The response message includes indication information of the target graphics processor, so that the target inference instance is mounted to the target graphics processor.
[0074] Correspondingly, the inference service module receives the response message and mounts the target inference instance to the target graphics processor according to the response message.
[0075] The graphics processor resource management method provided in this embodiment increases three identification fields for configuration when starting an inference instance, which facilitates the new creation of inference instances and business processing, so that multiple inference instances can share the same GPU, and the GPU resources can be used more efficiently.
[0076] In this embodiment, another graphics processor resource management method is also provided, which can be used in the inference engine module,Figure 4 FIG. 1 is a flow chart of another method for managing graphics processor resources according to an embodiment of the present invention. Figure 4 As shown, the method includes the following steps: Step S401: receiving a message request from an inference service module for instructing the start of a target inference instance.
[0077] Step S402 : determining a target graphics processor from the plurality of graphics processors according to the first identifier, the second identifier, the third identifier, and the occupation status information of the plurality of graphics processors.
[0078] Specifically, the above step S402 includes: Step S4021 : When the second identifier indicates that the target inference instance occupies a shareable graphics processor, at least one first graphics processor is selected from the plurality of graphics processors according to the first identifier.
[0079] The first graphics processor is a shareable graphics processor.
[0080] Step S4022: Determine whether the remaining video memory value of the at least one first graphics processor is greater than the target video memory value based on the occupancy status information of the at least one first graphics processor.
[0081] Step S4023: Determine one of the at least one second graphics processor as a target graphics processor.
[0082] The second graphics processor is the first graphics processor whose remaining video memory value is greater than the target video memory value.
[0083] In some embodiments, step S4023 may specifically include: determining the second graphics processor, whose difference between the remaining video memory value and the target video memory value is greater than the preset video memory value, as the target graphics processor.
[0084] The preset video memory value may be determined based on the total video memory value of the GPU. For example, the preset video memory value may be 5% of the total video memory value.
[0085] Exemplarily, when the first graphics processor or the second graphics processor does not exist, step S402 further includes step a1 and step a2: Step a1: Filter out at least one third graphics processor from the multiple graphics processors based on the occupation status information of the multiple graphics processors and the first identifier.
[0086] The third graphics processor is a non-shareable and unoccupied graphics processor.
[0087] Step a2: Determine one of the at least one third graphics processor as a target graphics processor.
[0088] For example, when the second identifier indicates that the target inference instance occupies a non-shareable graphics processor, the step S402 further includes steps b1 and b2: In step b1, at least one third graphics processor is selected from the plurality of graphics processors according to the occupation state information of the plurality of graphics processors and the first identifier.
[0089] In step b1, the same as step a1.
[0090] In step b2, the third graphics processor with a display memory value greater than the target display memory value is determined as the target graphics processor.
[0091] In step S403, a response message is sent to the inference service module.
[0092] In this embodiment, the three added identifiers are used in combination, so that the inference engine module can more reasonably control multiple inference instances to share a single GPU card according to the requirements of the inference instance and the occupation state information of the GPU, and optimize the utilization rate of the GPU.
[0093] In this embodiment, a server is also provided, which includes the graphics processor resource management system provided in any of the above embodiments. In one example, as shown in Figure 5 The graphics processor resource pool 200 can be deployed in the same server 500 as the graphics processor resource management system 100.
[0094] The server can also include one or more processors, memories, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components communicate with each other using different buses, and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the server, including instructions stored in the memory or on the memory to display a GUI on an external input / output device such as a display device coupled to the interface. In some optional embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple storage devices, if necessary. Similarly, multiple servers can be connected, each providing part of the necessary operations.
[0095] The processor can be a central processing unit, a network processing unit, or a combination thereof. The processor can further include a hardware chip. The hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device can be a complex programmable logic device, a field programmable logic gate array, a generic array logic, or any combination thereof.
[0096] The memory can include a program storage area and a data storage area. The program storage area can store an operating system and application programs required for at least one function. The data storage area can store data created according to the use of the server, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid state memory device.
[0097] The memory can include a volatile memory such as a random access memory, and can also include a non-volatile memory such as a flash memory, a hard disk, or a solid state disk. The memory can also include a combination of the above-mentioned types of memory.
[0098] It should be understood that various aspects of the application can be implemented in hardware, software, firmware or a combination thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies known in the art or their combinations can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits (ASICs) having appropriate combinational logic gates, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0099] In the description of the present specification, the description of the terms "the present embodiment", "one embodiment", "some embodiments", "example", "specific example" or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.
[0100] In addition, the terms "first", "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise explicitly specified.
[0101] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations shall all fall within the scope defined by the present invention.
Claims
1. A graphics processor resource management system, characterized in that: The system includes an inference engine module and an inference service module, wherein the inference engine module and the inference service module are both communicatively connected to a graphics processor resource pool, wherein the graphics processor resource pool includes multiple graphics processors, each of which is configured with a first identifier, and the first identifier is used to indicate whether the graphics processor can be shared by multiple inference instances; The inference service module is used to generate a message request indicating the start of a target inference instance, and is used to send the message request to the inference engine module, wherein the message request includes a second identifier and a third identifier, wherein the second identifier is used to indicate whether the target inference instance occupies a shareable graphics processor, and the third identifier is used to indicate a target video memory value required to be occupied after the target inference instance is started; The inference engine module is configured to determine a target graphics processor from the multiple graphics processors based on the first identifier, the second identifier, the third identifier, and the occupancy status information of the multiple graphics processors, and to send a response message to the inference service module, the response message including indication information of the target graphics processor; The inference service module is further configured to mount the target inference instance to the target graphics processor according to the response message.
2. The system according to claim 1, wherein: The inference engine module is configured to, when the second identifier indicates that the target inference instance occupies a shareable graphics processor, filter out at least one first graphics processor from the multiple graphics processors according to the first identifier, where the first graphics processor is a shareable graphics processor; The inference engine module is further configured to determine, based on the occupancy status information of the at least one first graphics processor, whether the remaining video memory value of the at least one first graphics processor is greater than the target video memory value; The inference engine module is further configured to determine one of at least one second graphics processor as the target graphics processor, where the second graphics processor is a first graphics processor whose remaining graphics memory value is greater than the target graphics memory value.
3. The system according to claim 2, characterized in that The inference engine module is configured to determine a second graphics processor, whose difference between the remaining video memory value and the target video memory value is greater than a preset video memory value, as the target graphics processor.
4. The system according to claim 2, wherein: In a case where the first graphics processor or the second graphics processor does not exist, the inference engine module is further configured to select at least one third graphics processor from the multiple graphics processors based on the occupation status information of the multiple graphics processors and the first identifier, where the third graphics processor is a non-sharable and unoccupied graphics processor; The inference engine module is further configured to determine one of at least one third graphics processor as the target graphics processor.
5. The system according to claim 4, characterized in that The inference service module is further used to update the occupancy status information of the target graphics processor after the target inference instance is successfully started, and to synchronize the updated occupancy status information of the target graphics processor and the information generated by the target inference instance to the inference engine module.
6. The system according to claim 4 or 5, characterized in that The inference engine module is further configured to, after determining the target GPU, update a first identifier corresponding to the target GPU to change the target GPU to a shareable GPU.
7. The system according to claim 4 or 5, characterized in that The inference engine module is further configured to, when the second identifier indicates that the target inference instance occupies a non-shareable graphics processor, filter out at least one third graphics processor from the multiple graphics processors based on the occupancy status information of the multiple graphics processors and the first identifier; The inference engine module is further configured to determine a third graphics processor having a video memory value greater than the target video memory value as the target graphics processor.
8. The system according to any one of claims 1 to 5, characterized in that The multiple graphics processors are deployed on different servers.
9. A graphics processor resource management method, characterized in that: The method is applied to the inference engine module of the graphics processor resource management system according to any one of claims 1 to 8, and the method comprises: Receiving a message request from the inference service module for instructing the start of a target inference instance, the message request including a second identifier and a third identifier, the second identifier being used to indicate whether the target inference instance occupies a shareable graphics processor, and the third identifier being used to indicate a target video memory value required to be occupied after the target inference instance is started; Determining a target graphics processor from the multiple graphics processors according to the first identifier, the second identifier, the third identifier, and occupation status information of the multiple graphics processors, wherein the first identifier is used to indicate whether the graphics processor can be shared by multiple inference instances; A response message is sent to the inference service module, where the response message includes indication information of the target graphics processor, so that the target inference instance is mounted to the target graphics processor.
10. A server, characterized in that: The server includes the graphics processor resource management system according to any one of claims 1 to 8.
Citation Information
Patent Citations
Shared GPU (Graphics Processing Unit) scheduling method based on Kubernetes
CN111506404A
Cluster GPU resource management scheduling system and method and computer readable storage medium
CN111538586A
Resource isolation method, distributed platform, computer equipment and storage medium
CN112114958A
Resource scheduling method and device, storage medium and electronic equipment
CN112631780A
GPU resource management method and system
CN115437781A
Cited By
Resource scheduling method and device
CN121092327A