A method for dynamic allocation of GPU resources under multi-tasking concurrency

By realizing the configurability of resource usage and dynamic resource regulation of C/S architecture in the GPU multi-task environment, the problem of unreasonable allocation of GPU resources in multi-task situations is solved, and resource utilization efficiency and system throughput are improved.

CN114048026BActive Publication Date: 2025-05-23BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111258248.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-27
Publication Date
2025-05-23
Estimated Expiration
2041-10-27

AI Technical Summary

Technical Problem

In the case of multitasking, the prior art cannot effectively allocate GPU resources, resulting in insufficient resource utilization, decreased system throughput, and unreasonable resource allocation.

Method used

By realizing the configurability of resource usage during GPU program runtime without modifying hardware or driver details, the C/S architecture is adopted, the client and the server are separated, and the server dynamically regulates resource usage, and the concurrent execution of complementary tasks of resource requirements is achieved.

Benefits of technology

It improves the efficiency of GPU resources, enhances multi-task processing capabilities, and solves the problems of resource idleness and system throughput reduction caused by static resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114048026B_ABST
    Figure CN114048026B_ABST
Patent Text Reader

Abstract

The present invention provides a method for dynamically allocating GPU resources under multi-tasking conditions to solve the problems of a large number of idle resources, reduced system throughput, and unreasonable resource allocation caused by the static resource allocation method when NVIDIA GPU has multiple tasks concurrently. The method has three obvious characteristics: (1) configurability. The native GPU environment cannot autonomously configure the amount of resources occupied when the program is running. The system proposes a software method to achieve the configurability of resource usage when the GPU program is running without modifying any hardware or driver details; (2) high efficiency. The method considers the affinity of tasks to different types of resources, and concurrently executes tasks with complementary resource requirements, thereby improving the efficiency of GPU resource utilization and accelerating multi-tasking; (3) ease of use. The method provides a simple program conversion mode. Developers only need to use fixed operating steps to migrate native programs to run under the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field:

[0001] The invention discloses a method for dynamically allocating GPU resources in a multi-tasking situation, relates to the challenges faced by high-performance computing, and belongs to the field of computer technology. Background technology:

[0002] As a device with a large number of computing cores and capable of providing high-speed parallel computing, GPU has long been used beyond the scope of graphics computing and rendering and is widely used in large-scale parallel computing such as high-performance computing, massive data processing and artificial intelligence. As an emerging heterogeneous computing platform, many programming frameworks provide support for it, including the CUDA programming model proposed by NVIDIA and the OpenCL programming language supported by many manufacturers. Although there are efficient development programming models for GPU devices, the leftover resource allocation strategy adopted by the GPU driver at runtime greatly damages the parallelism of GPU devices in multi-task situations. Even if developers write reasonable and high-performance GPU processing programs, under multi-task concurrency, the insufficient scheduling strategy of the driver will directly lead to insufficient utilization of GPU resources, which in turn leads to the GPU acceleration effect being not obvious or even ineffective. How to design and implement an intelligent multi-task concurrent platform that can make scheduling decisions based on load characteristics is the focus of the current use of GPUs for high-performance computing.

[0003] In order to solve the problem of unreasonable GPU resource allocation in multi-tasking, many solutions have emerged in recent years:

[0004] 1) Hyper-Q technology: In the Nvidia Fermi architecture, the GPU already supports basic spatial reuse. Users can specify which kernels are unrelated and can run simultaneously. The GPU will try to allocate the remaining resources to other parallel kernels when scheduling tasks. The Kepler architecture further improves the GPU's support for multiple kernels in parallel. It sets up multiple kernel queues that are not dependent on each other. When the tasks in one queue do not occupy all GPU resources, the GPU will call the tasks in another queue into the GPU. This kernel queue technology is called Hyper-Q technology. Each core on the CPU can submit a GPU task queue, and when the GPU is running, it selects parallel kernels from multiple parallel queues for calculation. Unlike the Fermi architecture, the Kepler architecture truly supports concurrent scheduling of multiple queues, while the Fermi architecture merges multiple queues into one execution queue during compilation, providing a pseudo-concurrency effect.

[0005] 2) MPS technology: MPS is a GPU multi-task concurrent operation framework developed by NVIDIA. Hyper-Q solves the concurrency problem on GPU hardware, while MPS solves the concurrency problem of multiple processes on software. MPS maps the CUDA contexts of multiple processes to the same CUDA context through the CS architecture, but the task scheduling mode is closed inside the GPU driver, and developers are not allowed to modify it according to their needs.

[0006] 3) GPU-Sim: GPU-Sim is an open source GPU simulator that can simulate the working process of a GPU. Researchers can improve GPU hardware or scheduling strategies by modifying the source code of GPU-Sim. Many researchers have studied hardware-based GPU scheduling methods based on GPU-Sim, such as modifying the amount of resources occupied by two kernels by adjusting the throughput of SM when trying different resource ratios during operation; using shared memory and L1 cache for fast GPU context switching, etc. The experimental verification of these methods all uses GPU-Sim to obtain simulation results. However, there is no corresponding hardware function in real GPUs, and the resource allocation and scheduling strategies demonstrated in the experiments cannot be used in real production environments.

[0007] The existing GPU multi-tasking concurrency support and resource allocation strategies do not solve the GPU resource scheduling in the case of multi-tasking concurrency from the perspectives of efficiency and feasibility. MPS and Hyper-Q have serious resource waste, and the hardware transformation based on GPU-Sim does not have real GPU device support, and cannot truly solve the problem of GPU resource waste in the case of multi-tasking concurrency.

[0008] Specifically, the problem with existing multi-tasking GPU concurrency is that there is no more effective GPU-controllable GPU resource allocation method, and it is impossible to make intelligent resource allocation choices based on task load characteristics during runtime, resulting in serious waste of GPU resources and inability to maximize the capabilities of parallel computing. Summary of the invention:

[0009] The main purpose of the present invention is to provide a method for dynamically allocating GPU resources under multi-tasking conditions, so as to solve the problems of a large number of idle resources, reduced system throughput and unreasonable resource allocation caused by the static resource allocation method when NVIDIA GPU multi-tasking is performed concurrently. The present invention has three obvious characteristics: (1) Configurability. The native GPU environment cannot autonomously configure the amount of resources occupied when the program is running. The present system proposes a software method to achieve the configurability of resource usage when the GPU program is running without modifying any hardware or driver details; (2) Efficiency. The present method considers the affinity of tasks to different types of resources, and concurrently executes tasks with complementary resource requirements, thereby improving the efficiency of GPU resource utilization and accelerating multi-tasking processing; (3) Ease of use. The present method provides a simple program conversion mode. Developers only need to use fixed operating steps to migrate native programs to run under the present system.

[0010] The technical solution of the present invention is:

[0011] A method for dynamically allocating GPU resources in a multi-tasking situation. First, a control code segment is inserted into a CUDA program through a fixed source code modification mode, so that the amount of resources occupied by the program during operation can be controlled. Then the operation mode of the CUDA program is modified to a C / S architecture, and the client is a set of self-implemented CUDA APIs, which replace the CUDA API in the original host program. The server uses a backend process to receive CUDA calls from different clients and queues the calls for processing. Finally, the server dynamically regulates the resource usage. The server traverses the CUDA request queue and puts the computing tasks on the GPU for execution. When the task at the head of the queue complements the resource requirements of the running task, the server dynamically reduces the amount of resources occupied by the running task and starts the waiting task to achieve concurrent processing.

[0012] The following steps are included:

[0013] 1) Modify the GPU program, insert a code segment that controls the program resource usage, turn the dynamically allocated blocks in the CUDA program into workers, and loop to pull blocks for execution.

[0014] 2) Sample the GPU program and run the GPU program for a period of time to ensure that all GPU resources are occupied and continue for a period of time, obtain the total number of instructions executed by the GPU program during this period of time and the number of all memory instructions.

[0015] 3) Use Json format text to save GPU program sampling data, as well as the number of occupied registers and shared memory;

[0016] 4) When compiling the program, replace the native CUDA API and link the CUDA API library implemented by this system. The CUDA API library implemented by this system will send the CUDA request to the server through the socket.

[0017] 5) The server creates a thread for each active client connection, and the thread saves the CUDA task queue.

[0018] 6) Before starting the kernel, a space is allocated on the GPU memory to save the resource allocation configuration when the program is running. When the program is running, the control code segment controls the amount of resources used by the GPU.

[0019] 7) When there are two or more GPU programs that need to be run, the optimal allocation method is calculated by comparing sample data, setting the configuration data on the GPU video memory, and running the kernel.

[0020] 8) When a GPU task is completed, the server is notified, and the server chooses to pull a new task from the queue and execute it concurrently, or expand the capacity of the running GPU program.

[0021] Wherein, step 1) comprises the following steps:

[0022] Step (1.1) The developer adds a global memory usage declaration in the GPU code segment. This global memory is the configuration saved in the video memory, and in order to complete the workload, the running kernel threads can read it concurrently;

[0023] Step (1.2) inserts a control code segment, which is executed immediately when the program starts, determines the SM number it is in, and sets its own worker number;

[0024] Step (1.3) Copy a copy of the source program and replace the blockIdx and gridDim variables in the source program;

[0025] Step (1.4) The worker loops and pulls blocks to execute until all tasks are completed;

[0026] Wherein, step 2) comprises the following steps:

[0027] Step (2.1) sets the appropriate scale data so that the kernel program can occupy all GPU resources and run continuously for more than 1ms;

[0028] Step (2.2) Use nvprof to run the test program and output the total number of program instructions executed and the number of global memory access instructions;

[0029] Wherein, step 3) comprises the following steps:

[0030] Step (3.1) Use nvvp to run the test program and record the number of registers occupied by a kernel thread and the amount of shared memory occupied by a block;

[0031] Step (3.2) uses json format text to save program sample data, which contains four fields: the number of sampled global memory access instructions; the number of sampled total instructions; the number of registers; and the number of shared memories.

[0032] Step (3.3) The server reads the configuration file when initializing;

[0033] Wherein, step 4) comprises the following steps:

[0034] Step (4.1) Modify the cudaapi library path linked when the GPU program is compiled to point to the cudaapi implementation of this system;

[0035] Step (4.2) replace the cuda header file in the source program and modify it to the cudaapi header file of this system;

[0036] Step (4.3) Change the way the kernel program is called with triple angle brackets in the source program to be called with the cudaLaunchKernel interface;

[0037] Step (4.4) All cudaapi calls in the source program will be transferred to the cudaapi implemented by this system;

[0038] Step (4.5) When the cudaapi library implemented by this system is initialized, it first tries to connect to the cudaapi server. If the server is not enabled, it notifies the user to enable the server before running.

[0039] Step (4.6) After the cudaapi library is successfully connected, the server will create a thread to handle the newly connected api client;

[0040] Wherein, step 5) comprises the following steps:

[0041] Step (5.1) The server listens to a socket file and waits for the client to connect;

[0042] Step (5.2) The client establishes a connection with the server through the socket and transmits the cuda call request using the specified message format;

[0043] Step (5.3) After receiving the client connection request, the server opens a new thread to connect to the client;

[0044] Step (5.4) The server establishes a task queue for each client;

[0045] Wherein, step 6) comprises the following steps:

[0046] Step (6.1) Before the server starts the kernel, a global memory is allocated on the GPU memory to save the SM number that the kernel can occupy, the maximum number of workers that can coexist on each SM, the number of workers currently existing on each SM, the number of completed blocks, the total number of blocks, and the kernel's GridDim;

[0047] Step (6.2) The server starts the kernel, and the blocks assigned to the SM become workers, which cyclically pull blocks for execution;

[0048] Step (6.3) When the worker is initialized, it first checks the sm_id it is in and the maximum number of workers that can exist on the sm. If the worker exceeds the limit, it will exit directly. If it does not exceed the limit, it will continue to exist on the device;

[0049] Step (6.4) The worker that persists on the device contains a loop body, and each loop will pull a block;

[0050] Step (6.5) The worker calculates the blockIdx of the block, passes in the source program copy described in step (1.3), and executes the copy program;

[0051] After the execution of step (6.6) block is completed, the worker checks whether the task has been completed. If it has been completed, the worker will exit the execution, otherwise it will enter the next cycle;

[0052] Wherein, step 7) comprises the following steps:

[0053] Step (7.1) The server calculates the memory instruction ratio of the running kernel, that is, the number of global memory instructions in the sampled data divided by the total number of instructions;

[0054] Step (7.2) calculates the memory-to-instruction ratio of the kernel to be run;

[0055] Step (7.3) derives all resource allocation combinations of the two kernels running in pairs according to the register field and the shared memory field in the sampled data;

[0056] In step (7.4), if there is a situation in the combination where the system memory-to-instruction ratio is close to 0.05 when two kernels are running concurrently, then the concurrent running is performed according to this configuration;

[0057] In step (7.5), if the system memory-to-instruction ratio of all combinations is greater than 0.055 or less than 0.045, no concurrent operation is performed;

[0058] Wherein, step 8) comprises the following steps:

[0059] Step (8.1) The server checks the task queue. If there is a runnable task, proceed to step (7). If the queue is empty, proceed to step (8.2).

[0060] Step (8.2) The server updates the running kernel configuration so that it occupies all resources;

[0061] Step (8.3) reads the kernel's runtime configuration. If the remaining surviving workers meet the resource configuration, continue to run as is; if the remaining surviving workers are less than the configuration, proceed to step (8.4)

[0062] Step (8.4) starts a new instance of the kernel and shares the configuration with the original instance so that the number of active workers of the kernel on the system reaches the configuration requirement;

[0063] Advantages of the present invention include:

[0064] Compared with the prior art, the method for dynamically allocating GPU resources in a multi-tasking situation proposed by the present invention has the following advantages:

[0065] In order to solve the problems of a large number of idle resources, reduced system throughput and unreasonable resource allocation caused by the static resource allocation method when NVIDIA GPU is running multiple tasks concurrently, an efficient GPU multi-tasking runtime system is proposed. The system has three obvious characteristics: (1) Configurability. The native GPU environment cannot independently configure the amount of resources occupied by the program when running. This system proposes a software method to achieve the configurability of resource usage when the GPU program is running without modifying any hardware or driver details. (2) Efficiency. This method considers the affinity of tasks to different types of resources and executes tasks with complementary resource requirements concurrently, thereby improving the efficiency of GPU resource utilization and accelerating multi-tasking processing. (3) Ease of use. This method provides a simple program conversion mode. Developers only need to follow fixed operating steps to migrate native programs to run under this system. Description of the drawings:

[0066] Figure 1 The present invention is a flowchart of a method for dynamically allocating GPU resources under multi-task concurrency.

[0067] Figure 2 It is a diagram of the GPU resource configuration data structure.

[0068] Figure 3 This is a diagram of the worker execution mode.

[0069] Figure 4 This is the structure diagram of the GPU multi-task concurrent execution system.

[0070] Figure 5 It is the flowchart of starting kernel.

[0071] Figure 6 It is a dynamic resource allocation flow chart. Specific implementation method:

[0072] The present invention is further described in detail below with reference to the accompanying drawings.

[0073] like Figure 1 The figure shows the resource configuration data structure diagram of the GPU program runtime. The configuration consists of a 32-bit, 4-byte fixed resource limit segment and an indefinite length running status segment. The first 2 bytes of the fixed resource limit segment indicate the maximum number of workers that can run on each SM, and the third and fourth bytes indicate the upper and lower limits of the SM number that can be used, respectively; the first three fields of the running status segment save the Grid dimension of the kernel, and the first field "Number of completed blocks" records the current number of completed blocks, and the second field is the total number of blocks that need to be completed for the task; starting from the third field, the content saved is the number of active workers on SMs numbered 0, 1, 2, ...n. The number of entries in this field is consistent with the number of SMs of the GPU hardware, recording the number of active workers of the kernel on each SM; when the kernel is initialized, the system first allocates a piece of GPU memory for it to save Figure 1 The runtime configuration shown has a lifecycle equal to the kernel runtime.

[0074] The execution process of Worker is as follows Figure 2 As shown in Figure 1, when a block is assigned to an SM, it becomes a worker. The worker first obtains the current SMid and worker number through PTX assembly code, and increases the number of active workers by one through atomic operations; the worker reads Figure 1The runtime configuration in the configuration is compared, and the SMid is compared with the upper and lower limits in the configuration, and the workerid is compared with the worker upper limit in the configuration to determine whether the worker exceeds the configuration limit. If the SM or worker limit is exceeded, the worker exits and the number of active workers is reduced by one; if it does not exceed the configuration limit, the unfinished block is pulled, and the number of completed blocks is increased by one, and the block number is compared with the total number of blocks. If the number exceeds the total number of blocks, the worker exits and the number of active workers is reduced by one; if the number does not exceed the total number of blocks, the blockIdx is calculated according to the GridDIM in the configuration and the source program copy is called. After the source program copy exits, the worker pulls the next unexecuted block and repeats the above process until the number of completed blocks exceeds the total number of blocks.

[0075] The structure diagram of GPU multi-task concurrent execution system is as follows Figure 3 As shown in the figure, the overall architecture of the system is divided into the client and the server. The client consists of cudaapi and GPU programs. The system supports the concurrent execution of multiple GPU programs. The underlying cudaapi uses the IPC library to communicate with the server. The server supports concurrent connections of multiple clients, where the receiver module establishes a thread for each client connection to process all cuda requests of a client. The client's cuda requests will be placed in a unified Commandlists. The scheduler module on the server will poll the cuda requests in the commandlist. After making a decision, the scheduler will start an APIconductor to execute the real cuda task. The Device module saves the computing power and total resources of the current device, as well as the running kernel and the amount of resources it occupies. When the Scheduler makes a scheduling decision, it calls the Device module to obtain the latest current status of the device.

[0076] A kernel startup flowchart is as follows Figure 4As shown in the figure, before starting the kernel, the scheduler will check whether the current device is idle. When the device is completely idle, the scheduler will set the kernel to completely occupy the GPU resources. If the device is not idle, the scheduler will check whether the current device is running only one kernel or two kernels. If the current device is occupied by two kernels, the kernel to be scheduled will block and wait for the task on the GPU to release resources. If the current device is occupied by a kernel, the scheduler will calculate the memory instruction ratio of the current kernel and the running kernel. If there is a combination close to 0.05, the runtime configuration will be set and saved on the GPU. If there is no such combination, the kernel to be scheduled will enter a blocked state. When the scheduler determines that the kernel to be scheduled can run on the GPU, it will first set the runtime configuration and then start the kernel. When the kernel is completed, it will release the occupied GPU resources and notify the scheduler. The scheduler will choose to wake up the originally blocked kernel or pull a new kernel from the waiting queue for scheduling.

[0077] Figure 5 The dynamic scheduling process of GPU resources is shown. In the initial state, there is a kernel running on the GPU. When the scheduler schedules a new task, it will first calculate the memory instruction ratio of different combinations. If there is a combination with memory instruction in the range of 0.045-0.055, it will set concurrent execution. If there is no such combination, the scheduler will block the newly scheduled task and wait for the existing kernel to release resources. If the scheduler sets concurrent execution, it will first set the runtime configuration of the kernel on the GPU, resulting in a reduction in the resources occupied by the running kernel. Then the scheduler will start a new task, and the two tasks will continue to execute concurrently. If the original kernel is executed first, it will release resources and notify the scheduler. If the later kernel is executed first, it will first check whether there are blocked tasks waiting to be scheduled. If there are blocked tasks, the kernel will immediately release resources and notify the scheduler. If there are no blocked tasks, the scheduler will make the original kernel occupy all GPU resources again. The scheduler first reads the runtime configuration of the kernel saved on the GPU. If the number of active workers is less than the maximum limit, the scheduler will enable the backup kernel of the kernel to increase the number of active workers on the GPU. If the number of active workers already meets the maximum limit, the scheduler does nothing. After the kernel is executed, the resources will be released and the scheduler will be notified. The scheduler will enter a new round of task scheduling.

[0078] Finally, it should be noted that the present invention may also have many other application scenarios. Without departing from the spirit and essence of the present invention, technicians familiar with the field can make various corresponding changes and deformations based on the present invention, but these corresponding changes and deformations should all fall within the scope of protection of the present invention.

Claims

1. A method for dynamically allocating GPU resources in a multi-tasking situation. First, a control code segment is inserted into a CUDA program through a fixed source code modification mode, so that the amount of resources occupied by the program during operation can be controlled. Then, the operation mode of the CUDA program is modified to a C / S architecture. The client is a set of self-implemented CUDA APIs, which replace the CUDA APIs in the original host program. The server uses a backend process to receive CUDA calls from different clients and queues the calls for processing. Finally, the server dynamically regulates the resource usage. The server traverses the CUDA request queue and puts the computing tasks on the GPU for execution. When the resource requirements of the task at the head of the queue complement those of the task being run, the server dynamically reduces the amount of resources occupied by the running task and starts the waiting task to achieve concurrent processing. The following steps are involved: 1) Modify the GPU program, insert a code segment that controls the program resource usage, turn the dynamically allocated blocks in the CUDA program into workers, and cyclically pull blocks for execution; 2) Sample the GPU program and run the GPU program for a period of time to ensure that all GPU resources are occupied and continue for a period of time, and obtain the total number of instructions executed by the GPU program during this period of time and the number of all memory instructions; 3) Use Json format text to save GPU program sampling data, as well as the number of occupied registers and shared memory; 4) When compiling the program, replace the native CUDA API and link the CUDA API library implemented by this system. The CUDA API library implemented by this system will send the CUDA request to the server through the socket; 5) The server creates a thread for each active client connection, and the thread saves the CUDA task queue; 6) Before starting the kernel, a space is allocated on the GPU memory to save the resource allocation configuration when the program is running. When the program is running, the control code segment controls the amount of resources used by the GPU; 7) When there are two or more GPU programs to run, calculate the optimal allocation method by comparing sample data, set the configuration data on the GPU memory, and run the kernel; 8) When a GPU task is completed, the server is notified, and the server chooses to pull a new task from the queue and execute it concurrently, or expand the capacity of the running GPU program.

2. The method according to claim 1, It is characterized in that The step 1) comprises the following steps: Step (1.1) The developer adds a global memory usage declaration in the GPU code segment. This global memory is the configuration saved in the video memory, and in order to complete the workload, the running kernel threads can read it concurrently; Step (1.2) inserts a control code segment, which is executed immediately when the program starts, determines the SM number it is in, and sets its own worker number; Step (1.3) Copy a copy of the source program and replace the blockIdx and gridDim variables in the source program; Step (1.4) The worker loops and pulls blocks to execute until all tasks are completed.

3. The method according to claim 1, It is characterized in that The step 2) comprises the following steps: Step (2.1) sets the appropriate scale data so that the kernel program can occupy all GPU resources and run continuously for more than 1ms; Step (2.2) uses nvprof to run the test program and output the total number of program instructions executed and the number of global memory access instructions.

4. The method according to claim 1, It is characterized in that The step 3) comprises the following steps: Step (3.1) Use nvvp to run the test program and record the number of registers occupied by a kernel thread and the amount of shared memory occupied by a block; Step (3.2) uses json format text to save program sample data, which contains four fields: the number of sampled global memory access instructions; the number of sampled total instructions; the number of registers; and the number of shared memories. Step (3.3) reads the configuration file when the server is initialized.

5. The method according to claim 1, It is characterized in that The step 4) comprises the following steps: Step (4.1) Modify the cudaapi library path linked when the GPU program is compiled to point to the cudaapi implementation of this system; Step (4.2) replace the cuda header file in the source program and modify it to the cudaapi header file of this system; Step (4.3) Change the way the kernel program is called with triple angle brackets in the source program to be called with the cudaLaunchKernel interface; Step (4.4) All cudaapi calls in the source program will be transferred to the cudaapi implemented by this system; Step (4.5) When the cudaapi library implemented by this system is initialized, it first tries to connect to the cudaapi server. If the server is not enabled, it notifies the user to enable the server before running. Step (4.6) After the cudaapi library is successfully connected, the server will create a thread to handle the newly connected api client.

6. The method according to claim 1, It is characterized in that The step 5) comprises the following steps: Step (5.1) The server listens to a socket file and waits for the client to connect; Step (5.2) The client establishes a connection with the server through the socket and transmits the cuda call request using the specified message format; Step (5.3) After receiving the client connection request, the server opens a new thread to connect to the client; Step (5.4) The server creates a task queue for each client.

7. The method according to claim 1, It is characterized in that The step 6) comprises the following steps: Step (6.1) Before the server starts the kernel, a global memory is allocated on the GPU memory to save the SM number that the kernel can occupy, the maximum number of workers that can coexist on each SM, the number of workers currently existing on each SM, the number of completed blocks, the total number of blocks, and the kernel's GridDim; Step (6.2) The server starts the kernel, and the blocks assigned to the SM become workers, which cyclically pull blocks for execution; Step (6.3) When the worker is initialized, it first checks the sm_id it is in and the maximum number of workers that can exist on the sm. If the worker exceeds the limit, it will exit directly. If it does not exceed the limit, it will continue to exist on the device. Step (6.4) The worker that persists on the device contains a loop body, and each loop will pull a block; Step (6.5) The worker calculates the blockIdx of the block, passes in the source program copy described in step (1.3), and executes the copy program; After the execution of step (6.6) block is completed, the worker checks whether the task has been completed. If it has been completed, the worker will exit the execution, otherwise it will enter the next cycle.

8. The method according to claim 1, It is characterized in that The step 7) comprises the following steps: Step (7.1) The server calculates the memory instruction ratio of the running kernel, that is, the number of global memory instructions in the sampled data divided by the total number of instructions; Step (7.2) calculates the memory-to-instruction ratio of the kernel to be run; Step (7.3) derives all resource allocation combinations of the two kernels running in pairs according to the register field and the shared memory field in the sampled data; In step (7.4), if there is a situation in the combination where the system memory-to-instruction ratio is close to 0.05 when two kernels are running concurrently, then the concurrent running is performed according to this configuration; In step (7.5), if the system memory-to-instruction ratio of all combinations is greater than 0.055 or less than 0.045, no concurrent execution is performed.

9. The method according to claim 1, It is characterized in that The step 8) comprises the following steps: Step (8.1) The server checks the task queue. If there is a runnable task, proceed to step (7). If the queue is empty, proceed to step (8.2). Step (8.2) The server updates the running kernel configuration so that it occupies all resources; Step (8.3) reads the kernel's runtime configuration. If the remaining surviving workers meet the resource configuration, the system continues to run as it is. If the remaining surviving workers are less than the configuration, the system proceeds to step (8.4). Step (8.4) starts a new instance of the kernel and shares the configuration with the original instance so that the number of active workers of the kernel on the system reaches the configured requirement.

Citation Information

Patent Citations

  • GPU thread scheduling optimization method

    CN103336718A

  • GPU cluster environment-oriented method for avoiding GPU resource contention

    CN107943592A