A GPU resource isolation method, system, medium and product
By formatting the GPU attribute information into bit charts and using atomic operations for resource allocation and recycling, the problem of GPU resource isolation in heterogeneous computing scenarios is solved, and the maximum utilization of GPU resources and the overall utilization rate are improved.
Patent Information
- Application Number
- CN202510144261.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The prior art is difficult to effectively isolate GPU hardware resources in heterogeneous computing scenarios, resulting in low utilization of a single computing unit and the inability to maximize the computing resources utilizing GPUs.
By obtaining the attribute information of each GPU, formatting it into a bitmap containing the computing resource allocation status, and using atomic operations to allocate and recycle computing resources for GPU resource requests, realizing the isolation and dynamic scheduling of GPU resources.
It realizes effective isolation and maximization of GPU hardware resources in heterogeneous computing scenarios, improves the overall utilization rate of GPUs, and avoids the situation of starving a single application.
Smart Images

Figure CN119597491B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud computing technology, and in particular to a GPU resource isolation method, system, medium and product. Background Art
[0002] With the sudden popularity of artificial intelligence applications such as ChatGPT, artificial intelligence technology has been mentioned by more and more people. Nvidia provides CUDA, a heterogeneous computing development platform. Since the driver is closed source, heterogeneous computing based on CUDA needs to limit the hardware to use Nvidia's GPU. Therefore, resource isolation in cloud computing is mostly based on the hardware isolation solution provided by hardware manufacturers. This solution can divide a single GPU into multiple independent vGPU units through hardware, and execute computing tasks concurrently without interfering with each other. However, this solution is limited by hardware, and the maximum supported isolation units are generally only a few, and the utilization rate of a single computing unit may not be fully utilized, so that the computing resources of the GPU cannot be maximized. How to solve the isolation of GPU hardware resources in heterogeneous computing scenarios and maximize the utilization of GPU computing resources has become a key technical problem that needs to be solved urgently. Summary of the invention
[0003] Technical problem to be solved by the present invention: In view of the above-mentioned problems in the prior art, a GPU resource isolation method, system, medium and product are provided. The present invention aims to solve the isolation of GPU hardware resources in heterogeneous computing scenarios and maximize the utilization of GPU computing resources.
[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0005] A GPU resource isolation method comprises the following steps:
[0006] S1, obtain the attribute information of each GPU;
[0007] S2, formatting the attribute information of each GPU into a bitmap table in the form of a bitmap containing a computing resource allocation state and saving it;
[0008] S3, waiting for a GPU resource request from an application in the virtual machine, and after receiving the GPU resource request, accessing the bitmap table in the form of atomic operations to allocate available computing resources to the GPU resource request, submitting the GPU resource request to the corresponding GPU and updating the corresponding computing resource allocation state value;
[0009] S4, waiting for the GPU execution to be completed, and after the GPU execution is completed, recovering the computing resources allocated to the application in the form of atomic operations, and updating the corresponding computing resource allocation state value.
[0010] Optionally, the attribute information of the GPU in step S1 includes the GPU number, the number of computing units included in the GPU, the number of workgroups supported by a single computing unit, and the maximum number of threads supported by a single workgroup.
[0011] Optionally, when the attribute information of each GPU is formatted into a bitmap table in the form of a bitmap containing the computing resource allocation status in step S2, the bitmap table is saved in an array and the attribute information of each GPU is recorded using the GPU number as an index, and the bitmap composed of the attribute information of each GPU includes the maximum number of threads supported by a single workgroup, the number of single instruction multiple data supported by each computing unit and the computing unit status, each bit in the computing unit status is mapped to the allocation status of a workgroup, and a value of 1 indicates that the workgroup has been allocated, and a value of 0 indicates that the workgroup has not yet been allocated; updating the corresponding computing resource allocation status value in step S3 refers to setting the allocation status of the allocated workgroup to 1; updating the corresponding computing resource allocation status value in step S4 refers to setting the allocation status of the recycled workgroup to 0.
[0012] Optionally, the array is a 64-bit array, the index number range [0, 7] in the bitmap is the maximum number of threads supported by a single work group, the index number range [8, 10] is the number of single instruction multiple data supported by each computing unit, and the index number range [11, 63] is the computing unit status.
[0013] Optionally, in step S3, accessing the bitmap table in the form of atomic operations to allocate available computing resources to the GPU resource request and submitting it to the corresponding GPU includes: first, querying the bitmap table with a specified scheduling strategy and assigning the found unallocated work group to the GPU resource request; then submitting the GPU resource request to the queue of the GPU corresponding to the allocated work group, so that the corresponding GPU obtains the GPU resource request from the queue to execute and return the calculation result of the GPU resource request.
[0014] Optionally, it also includes creating an instance of the application for the application in the virtual machine when requesting the GPU resources of the system, registering a write timestamp event for the instance of the application, so that the application instructions execute the write timestamp event to write the timestamp to the memory after submission and execution; in step S4, after the GPU execution is completed, it also includes querying the timestamp of the application write timestamp event written into the memory, converting the GPU clock cycle between the timestamp at submission and the timestamp at completion into the usage time of this GPU resource request and recording it in the sampling information.
[0015] Optionally, when querying the bitmap table with the specified scheduling strategy and assigning the unallocated workgroup found to the GPU resource request, it also includes reading the usage time of the application's last GPU resource request and the average usage time avg_time of the GPU resource request recorded in the sampling information. If the reading fails, it means that the application is issuing a GPU resource request for the first time, and a default number of workgroups are found from the bitmap table and assigned to the GPU resource request; if the reading is successful, it means that the application is issuing a GPU resource request again. If the usage time of the last GPU resource request exceeds the average usage time avg_time of the GPU resource request, the number of workgroups assigned to the GPU resource request is increased based on the number of workgroups assigned last time; otherwise, the number of workgroups assigned to the GPU resource request is reduced based on the number of workgroups assigned last time to achieve elastic scaling for the application according to the number of workgroups assigned last time.
[0016] In addition, the present invention also provides a GPU resource isolation system, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the GPU resource isolation method.
[0017] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored. The computer program or instruction is programmed or configured to execute the GPU resource isolation method through a processor.
[0018] In addition, the present invention also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the GPU resource isolation method through a processor.
[0019] Compared with the prior art, the present invention has the following advantages: the method of the present invention comprises obtaining the attribute information of each GPU; formatting the attribute information of each GPU into a bitmap table in the form of a bitmap containing the computing resource allocation status and saving it; waiting for a GPU resource request from an application in a virtual machine, and after receiving the GPU resource request, accessing the bitmap table in the form of an atomic operation to allocate available computing resources to the GPU resource request, submitting the GPU resource request to the corresponding GPU and updating the corresponding computing resource allocation status value; waiting for the GPU to be executed, and after the GPU is executed, recovering the computing resources allocated to the application in the form of an atomic operation, and updating the corresponding computing resource allocation status value. The present invention schedules and allocates GPU resources of the GPU by setting a virtual mask composed of a bitmap containing the computing resource allocation status for the GPU resources, which can solve the isolation of GPU hardware resources in heterogeneous computing scenarios and realize the maximum utilization of GPU computing resources. Compared with the prior art, the advantages are: (1) the bitmap is used to manage the computing resources of all GPUs, and resource allocation and recovery can be performed with O(n) time consumption. (2) the bitmap is used as a virtual mask, and there is no need to invade the specific scheduling algorithm of the hardware. Logically manage resources through software to achieve balanced use of GPU resources and avoid starvation of a single application. (3) Obtain the specific execution time of the application through sampling, dynamically adjust the allocation of resources used by the application, and improve the overall utilization of the GPU. (4) The present invention can be applied to scenarios such as machine learning and deep learning that require the use of GPU resources, including but not limited to hardware driver solutions based on Vulkan. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Schematic diagram of the basic flow of the method of the embodiment of the present invention.
[0021] Figure 2 Schematic diagram of the system topology structure in an embodiment of the present invention.
[0022] Figure 3 Schematic diagram of the structure of a bitmap table in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be further described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0024] like Figure 1 As shown, the GPU resource isolation method of this embodiment includes the following steps:
[0025] S1, obtain the attribute information of each GPU;
[0026] S2, formatting the attribute information of each GPU into a bitmap table in the form of a bitmap containing a computing resource allocation state and saving it;
[0027] S3, waiting for a GPU resource request from an application in the virtual machine, and after receiving the GPU resource request, accessing the bitmap table in the form of atomic operations to allocate available computing resources to the GPU resource request, submitting the GPU resource request to the corresponding GPU and updating the corresponding computing resource allocation state value;
[0028] S4, waiting for the GPU execution to be completed, and after the GPU execution is completed, recovering the computing resources allocated to the application in the form of atomic operations, and updating the corresponding computing resource allocation state value.
[0029] like Figure 2 As shown, in this embodiment, steps S1-S4 are performed based on the daemon kGPU (kGPU server) as the execution subject. In a GPU cluster environment that supports Vulkan drivers, the virtual machine is configured with a vGPU to access hardware GPU resources. The kGPU server is started on the host to receive vGPU requests from the virtual machine. Applications in the virtual machine, such as llama.cpp, call the Vulkan API to request the backend physical GPU.
[0030] In step S1 of this embodiment, the kGPU server is specifically notified to call the Vulkan API to obtain the GPU resources on the system, and the attribute information of a single GPU is polled. In this embodiment, the attribute information of the GPU includes the GPU number, the number of compute units (Compute Unit) contained in the GPU, the number of work groups (Work Group) supported by a single compute unit, and the maximum number of threads (Invocations) supported by a single work group. It should be noted that the names of the compute units in different CPUs may be different. For example, for NVIDIA devices, the name is Steam Multiprocessors. A compute unit is usually divided into smaller scheduling unit work groups, and a work group can share the cache resources of the compute unit.
[0031] In step S2 of this embodiment, the attribute information of each GPU is formatted into bitmaps in the form of a bitmap containing the computing resource allocation status and saved, such as Figure 3As shown, in step S2 of this embodiment, when the attribute information of each GPU is formatted into a bitmap table in the form of a bitmap containing the computing resource allocation status and saved, the bitmap table is saved in an array manner and the attribute information of each GPU is recorded using the GPU number as an index, and the bitmap composed of the attribute information of each GPU includes the maximum number of threads supported by a single work group, the number of single instruction multiple data supported by each computing unit, and the computing unit status, each bit in the computing unit status is mapped to the allocation status of a work group, and a value of 1 indicates that the work group has been allocated, and a value of 0 indicates that the work group has not been allocated; updating the corresponding computing resource allocation status value in step S3 refers to setting the allocation status (bitmap value) of the allocated work group to 1; updating the corresponding computing resource allocation status value in step S4 refers to setting the allocation status (bitmap value) of the recycled work group to 0.
[0032] See also Figure 3 In this embodiment, the bitmap table uses an array of uint64_t to store the attribute information of each GPU. Each bitmap is 64 bits. The index number interval [0, 7] in the bitmap is the maximum number of threads supported by a single work group, the index number interval [8, 10] is the number of SIMD supported by each computing unit, and the index number interval [11, 63] is the computing unit status. When requesting GPU resources in the virtual machine, the request will first be sent to the kGPU server of the host. The kGPU server accesses the corresponding bitmap in the bitmap table through atomic operations, determines the allocation of each work group on all current GPUs, and allocates work group resources according to the request.
[0033] In step S3 of this embodiment, accessing the bitmap table in the form of atomic operations to allocate available computing resources for the GPU resource request and submitting it to the corresponding GPU includes: first querying the bitmap table with a specified scheduling strategy (such as polling in a Round-Robin manner, etc.) and assigning the unallocated work group found to the GPU resource request; then submitting the GPU resource request to the queue of the GPU corresponding to the allocated work group, so that the corresponding GPU obtains the GPU resource request from the queue to execute and return the calculation result of the GPU resource request. When the GPU obtains the GPU resource request from the queue to execute, it will extract the command to the computing unit in the order in which the command is submitted, but the order in which the command is executed is unknown.
[0034] This embodiment also includes creating an instance of the application when the application in the virtual machine requests the GPU resources of the system, registering a write timestamp event (writeTimestamp event) for the instance of the application, so that the application instruction executes the write timestamp event to write the timestamp (Timestamp) into the memory after submission and execution; in step S4, after the GPU execution is completed, it also includes querying the timestamp written into the memory by the write timestamp event of the application, converting the GPU clock cycle (tick) between the timestamp at the time of submission and the timestamp at the time of completion into the usage time of this GPU resource request and recording it in the sampling information (sample info), so as to obtain the actual execution time of the computing task by sampling the computing instance, and realize accurate monitoring of GPU resources, so as to provide basic data for dynamically allocating GPU computing resources to computing tasks.
[0035] Further, in this embodiment, for a single application instance, the callback function is registered to sample the actual GPU time consumed by the instance execution, and the number of workgroups allocated can be dynamically adjusted according to the consumed time. Specifically, in this embodiment, when the bitmap table is queried with the specified scheduling strategy and the unallocated workgroups found are allocated to the GPU resource request, it also includes reading the usage time of the last GPU resource request of the application recorded in the sampling information and the average value avg_time of the usage time of the GPU resource request. In the case of a reading failure, it means that the application is issuing a GPU resource request for the first time, and a default number of workgroups are found from the bitmap table and allocated to the GPU resource request. For example, as an optional implementation, the default number of workgroups allocated in this embodiment is 8*8, occupying one computing unit; in the case of a reading success, it means that the application is issuing a GPU resource request again. If the usage time of the last GPU resource request exceeds the average value avg_time of the usage time of the GPU resource request, the number of workgroups allocated to the GPU resource request is increased on the basis of the number of workgroups allocated last time, otherwise the number of workgroups allocated to the GPU resource request is reduced on the basis of the number of workgroups allocated last time to achieve elastic scaling for the application according to the number of workgroups allocated last time, and to achieve dynamic allocation of GPU computing resources. Among them, the method of increase and decrease can be adjusted according to actual needs, for example, increase and decrease can be achieved by adding and subtracting a fixed or dynamically variable adjustment amount (step size), or by multiplying by a fixed or dynamically variable coefficient greater than or less than 1, etc.
[0036] In summary, the GPU resource isolation method of this embodiment schedules and allocates computing resources by setting a virtual mask of GPU computing resources, and dynamically allocates computing resources by sampling the GPU execution time of the application. By avoiding intrusion into the GPU driver and managing the isolated allocation of GPU computing resources in user mode, it can effectively balance the GPU usage between applications and improve the overall GPU utilization.
[0037] In addition, this embodiment also provides a GPU resource isolation system, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the GPU resource isolation method.
[0038] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored. The computer program or instruction is programmed or configured to execute the GPU resource isolation method through a processor.
[0039] In addition, this embodiment also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the GPU resource isolation method through a processor.
[0040] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present application may be in the form of methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the process Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0041] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A GPU resource isolation method, characterized in that: The steps include: S1, obtain the attribute information of each GPU; S2, formatting the attribute information of each GPU into a bitmap table in the form of a bitmap containing a computing resource allocation state and saving it; S3, waiting for a GPU resource request from an application in the virtual machine, and after receiving the GPU resource request, accessing the bitmap table in the form of atomic operations to allocate available computing resources to the GPU resource request, submitting the GPU resource request to the corresponding GPU and updating the corresponding computing resource allocation state value; S4, waiting for the GPU execution to be completed, and after the GPU execution is completed, reclaiming the computing resources allocated to the application in the form of atomic operations, and updating the corresponding computing resource allocation state value; The attribute information of the GPU in step S1 includes the GPU number, the number of computing units included in the GPU, the number of workgroups supported by a single computing unit, and the maximum number of threads supported by a single workgroup; When the attribute information of each GPU is formatted into a bitmap table in the form of a bitmap containing the computing resource allocation status in step S2 and saved, the bitmap table is saved in an array manner and the attribute information of each GPU is recorded using the GPU number as an index, and the bitmap composed of the attribute information of each GPU includes the maximum number of threads supported by a single work group, the number of single instruction multiple data supported by each computing unit and the computing unit status, each bit in the computing unit status is mapped to the allocation status of a work group, and a value of 1 indicates that the work group has been allocated, and a value of 0 indicates that the work group has not yet been allocated; updating the corresponding computing resource allocation status value in step S3 refers to setting the allocation status of the allocated work group to 1; updating the corresponding computing resource allocation status value in step S4 refers to setting the allocation status of the recycled work group to 0.
2. The GPU resource isolation method according to claim 1, characterized in that: The array is a 64-bit array, the index number interval [0, 7] in the bitmap is the maximum number of threads supported by a single work group, the index number interval [8, 10] is the number of single instruction multiple data supported by each computing unit, and the index number interval [11, 63] is the computing unit status.
3. The GPU resource isolation method according to claim 1, characterized in that: In step S3, accessing the bitmap table in the form of atomic operations to allocate available computing resources to the GPU resource request and submitting it to the corresponding GPU includes: first, querying the bitmap table with a specified scheduling strategy and assigning the unallocated work group found to the GPU resource request; then submitting the GPU resource request to the queue of the GPU corresponding to the allocated work group, so that the corresponding GPU obtains the GPU resource request from the queue to execute and return the calculation result of the GPU resource request.
4. The GPU resource isolation method according to claim 3, characterized in that: It also includes creating an instance of the application when requesting GPU resources of the system for the application in the virtual machine, registering a write timestamp event for the instance of the application, so that the application instruction executes the write timestamp event to write the timestamp into the memory after submission and execution; After the GPU execution is completed in step S4, it also includes querying the timestamp of the application's write timestamp event written into the memory, converting the GPU clock cycle between the timestamp at submission and the timestamp at completion into the usage time of this GPU resource request and recording it in the sampling information.
5. The GPU resource isolation method according to claim 4, characterized in that: When querying the bitmap table with the specified scheduling strategy and allocating the unallocated workgroup to the GPU resource request, the method further includes reading the usage time of the last GPU resource request of the application and the average usage time avg_time of the GPU resource request recorded in the sampling information. If the reading fails, it means that the application has issued a GPU resource request for the first time, and a default number of workgroups are found from the bitmap table and allocated to the GPU resource request. If the reading is successful, it means that the application has issued a GPU resource request again. If the usage time of the last GPU resource request exceeds the average usage time avg_time of the GPU resource request, the number of workgroups allocated to the GPU resource request is increased based on the number of workgroups allocated last time. Otherwise, the number of workgroups allocated to the GPU resource request is reduced based on the number of workgroups allocated last time to achieve elastic scaling for the application according to the number of workgroups allocated last time.
6. A GPU resource isolation system, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the GPU resource isolation method according to any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the GPU resource isolation method described in any one of claims 1 to 5 through a processor.
8. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the GPU resource isolation method described in any one of claims 1 to 5 through a processor.
Citation Information
Patent Citations
GPU resource management allocation method and apparatus
CN108241532A
Load prediction method and device, electronic equipment and storage medium
CN115454620A