An application quick switching method and system applied to GPU computing power scheduling
By dividing the tasks of AI programs on the GPU into multiple execution subtasks and realizing pipelined task context switching, the problem of large overhead of GPU task switching in the prior art is solved, and the utilization rate and switching efficiency of the GPU are improved.
Patent Information
- Application Number
- CN202411626843.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-11-14
Smart Images

Figure CN119149248B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of GPUs, and in particular, to an application fast switching method and system for GPU computing power scheduling. Background Art
[0002] Currently, for the main workloads in AI applications, including throughput-intensive training tasks and latency-sensitive inference tasks of AI programs, the mainstream approach is to provide dedicated GPU clusters for training and inference respectively. However, in large model architectures, it is difficult to use the training cluster to serve both inference tasks and training tasks.
[0003] Refer to Figure 1 , the task execution of an AI program includes three stages: task preparation, task execution, and task cleaning. Since there is a high overhead when the GPU switches between new tasks and old tasks, and the current switching method is at the second level, while most inference tasks have strict service level objectives (SLOs) that need to be controlled within the range of dozens to hundreds of milliseconds, the switching efficiency is low, and it is difficult to pack multiple deep learning (DL) programs onto the same GPU server.
[0004] The existing solution is to let multiple different models share the GPU video memory spatially, including Figure 2 the NVIDIA Multi-Process Service (MPS) shown in Summary of the Invention
[0005] In order to reduce the task switching overhead of AI programs running on the GPU and improve the utilization rate of the GPU, this application provides an application fast switching method and system for GPU computing power scheduling.
[0006] The first invention object of this application is achieved through the following technical solutions:
[0007] An application fast switching method for GPU computing power scheduling includes the steps of:
[0008] Dividing a task of an AI program running on the GPU into several execution subtasks, and based on a preset preparation rule, dividing the task preparation stage into a first preparation stage and a second preparation stage;
[0009] Based on a preset operation rule, during the task preparation stage of the current task, running the first preparation stage of the next task;
[0010] When the preset process is reached during the task execution phase of the current task, run the second preparation phase of the first execution subtask in the next task;
[0011] After the second preparation phase of the first execution subtask in the next task has finished running, run the first execution subtask;
[0012] During the running of the first execution subtask, run the second preparation phase of the next execution subtask, and at the same time run the task cleaning phase to clear the tasks that have completed running;
[0013] When all execution subtasks have finished running, run the task cleaning phase.
[0014] By adopting the above technical solution, in order to effectively time-share the GPU and minimize the overhead of task switching, the tasks of running the AI program on the GPU are divided into multiple execution subtasks, and the task preparation phase is divided. During the preparation process of the current task, that is, start the first preparation phase of the next task. That is, when the old task starts, start the preparation of some new tasks first, and run the second preparation phase of the execution subtask simultaneously with the execution of the previous task, and run the task cleaning phase immediately after the execution subtask runs, keeping enough video memory on the GPU to prepare and run each execution subtask, and at the same time improving the utilization rate; Therefore, through the pipeline-style task context switching setting above, the video memory management and the running of the main and standby working subtasks are unified, greatly reducing the switching overhead and improving the utilization rate of the GPU.
[0015] Optionally, based on the preset preparation rules, the task preparation phase is divided into a first preparation phase and a second preparation phase, including:
[0016] The first preparation phase includes an environment initialization phase, which is used for process startup, the loading of the PyTorch CUDA runtime, and the initialization of the CUDA context;
[0017] The second preparation phase includes a video memory allocation phase and a data transfer phase. The video memory allocation phase is used to apply for matching video memory for the execution subtasks in the AI program, and the AI program is stored in the preset host memory;
[0018] The data transfer phase is used to copy the model data of the AI program model and the tasks that need to be run on the GUP from the host memory to the video memory.
[0019] By adopting the above technical solution, since only the runtime loading and the initialization of the CUDA context are performed in the environment initialization phase, the video memory occupied by the running is small, and the initialization work can be performed on multiple tasks or execution subtasks, improving the GPU utilization rate.
[0020] Optionally, during the task preparation phase of the current task, the first preparation phase of the next task is run based on a preset operation rule; this includes:
[0021] Based on the preset operation rule, when the current task starts the running environment initialization phase, the environment initialization phase of the next task is run synchronously;
[0022] When the video memory occupied by all tasks in the synchronously run environment initialization phase reaches the first preset value, the environment initialization phase of new tasks is paused;
[0023] After the task cleaning phase runs, when the video memory occupied by all tasks in the synchronously run environment initialization phase is less than the first preset value again, the environment initialization phase of the next task is run again.
[0024] By adopting the above technical solution, considering the influence of the video memory size, for the early preparation of the environment initialization phase of each task, the number also needs to be controlled. Therefore, the first preset value is set. When the video memory occupied by the tasks or subtasks performing environment initialization preparation reaches the first preset value, the early preparation of the environment initialization phase is first paused, and after the video memory is released during the task cleaning phase, the early preparation of the environment initialization is carried out again. This is to reduce the switching overhead time while maintaining a high GPU utilization rate.
[0025] Optionally, when the current task reaches a preset process during the task execution phase, the second preparation phase of the first execution subtask in the next task is run; this includes:
[0026] Obtain the number of execution subtasks in the current task and identify the quantity information of the unrun execution subtasks;
[0027] When the quantity information is less than the preset quantity, obtain the video memory application information corresponding to each unrun execution subtask;
[0028] Based on the video memory application information, when the video memory applied for by any unrun execution subtask is greater than the second preset value, mark this unrun execution subtask;
[0029] When the marked execution subtask finishes running, run the second preparation phase of the first execution subtask in the next task.
[0030] By adopting the above technical solution, in order to meet low-overhead task switching while maintaining high GPU utilization, for the operation in the second preparation stage, that is, the operation settings in the video memory allocation stage and the data transmission stage, it is necessary to identify the quantity and video memory of the execution subtasks that are not yet completed. If there are still many execution subtasks that are not yet run, then the second preparation stage of the execution subtasks in the next task will not be run first. After the number of execution subtasks that are not yet run is less than the preset quantity, it is still necessary to identify the video memory of the execution subtasks that are not yet run. If its video memory is greater than the second preset value, then mark this execution subtask and determine that there is still a risk of high switching overhead. Therefore, after waiting for the marked execution subtask to complete its operation, then run the second preparation stage of the next execution subtask, which can reduce the overhead of context switching.
[0031] Optionally, after the second preparation stage of the first execution subtask in the next task ends, run the first execution subtask; it includes:
[0032] When the second preparation stage of the first execution subtask in the next task ends, determine whether the task execution stage of the current task has ended;
[0033] If the task execution stage of the current task ends, then immediately run the first execution subtask;
[0034] If the task execution stage of the current task has not ended, then determine whether the sum of the video memory applied for by the currently running execution subtask and the first execution subtask to be run exceeds the second preset value. If so, wait until the task execution stage of the current task ends, and then run the first execution subtask.
[0035] By adopting the above technical solution, after the second preparation stage of the execution subtask is completed, that is, after the video memory application and data copy are completed, it is still necessary to determine whether the current execution subtask can be run immediately to reduce the overhead of context switching. Therefore, after the second preparation stage is completed, it is necessary to determine whether the sum of the video memory of the remaining incomplete execution subtasks and the execution subtask to be run exceeds the second preset value, and selectively determine whether the next execution subtask is run immediately or delayed, so as to effectively reduce the overhead of context switching in each operation link.
[0036] Optionally, the video memory allocation stage is used to apply for matching video memory for the execution subtasks in the AI program, and it includes:
[0037] Obtain all the execution subtasks in the AI program to be executed, and input the execution subtasks into the pre-completed trained video memory matching model in sequence;
[0038] When the video memory matching model receives an execution subtask, it identifies the type information and task division information of the task to which the execution subtask belongs;
[0039] Based on the type information and task division information, it matches a corresponding video memory application value for the current execution subtask; and when the video memory allocation stage of the current execution subtask is running, it generates and sends a video memory application instruction based on the video memory application value.
[0040] By adopting the above technical solution, in order to improve the efficiency of the video memory application stage of each execution subtask in the AI program, through a preset video memory matching model, the type of the task and the situation of the task division are identified, and the required video memory size can be accurately matched in advance for each upcoming task or execution subtask, thus improving the efficiency of the second preparation stage.
[0041] Optionally, the matching of a corresponding video memory application value for the current execution subtask based on the type information and task division information includes:
[0042] Based on the type information, a video memory range corresponding to the type of the task is screened out from the preset database of the video memory matching model;
[0043] Based on the task division information, the number of task divisions of the task to which the current execution subtask belongs is identified, and the video memory that matches the number of task divisions is screened out from the video memory range as the video memory application value of the current execution subtask.
[0044] By adopting the above technical solution, in order to further improve the efficiency of video memory matching, the general video memory matching range can be quickly determined through the type information of the task, and then it can be identified through the task division information how many execution subtasks the current task is divided into, so that the required video memory application value can be quickly matched within the already screened video memory range.
[0045] The second invention object of the present application is achieved through the following technical solution:
[0046] An application fast switching system applied to GPU computing power scheduling includes:
[0047] A division module, configured to divide a task of running an AI program on a GPU into several execution subtasks, and based on a preset preparation rule, divide the task preparation stage into a first preparation stage and a second preparation stage;
[0048] A first operation module, configured to run the first preparation stage of the next task during the task preparation stage of the current task based on a preset operation rule;
[0049] A second operation module, configured to, when a preset process is reached during the task execution phase of the current task, operate the second preparation phase of the first execution subtask in the next task;
[0050] A third operation module, configured to, after the second preparation phase of the first execution subtask in the next task ends, operate the first execution subtask;
[0051] A fourth operation module, configured to, during the operation of the first execution subtask, operate the second preparation phase of the next execution subtask, and simultaneously operate the task cleaning phase to clear the tasks that have completed their operations;
[0052] A fifth operation module, configured to, after all execution subtasks have completed their operations, operate the task cleaning phase.
[0053] The above-mentioned third objective of the present application is achieved through the following technical solutions:
[0054] A computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the steps of the above-mentioned application fast switching method applied to GPU computing power scheduling are implemented.
[0055] The above-mentioned fourth objective of the present application is achieved through the following technical solutions:
[0056] A computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned application fast switching method applied to GPU computing power scheduling are implemented.
[0057] In summary, the present application includes at least the following beneficial technical effects:
[0058] 1. During the preparation process of the current task, that is, at the beginning of the first preparation phase of the next task, that is, when the old task starts, the preparation of some new tasks is started first, and the second preparation phase of the execution subtask is run simultaneously with the execution of the previous task, and the task cleaning phase is run immediately after the execution subtask runs, so as to keep enough video memory on the GPU to prepare and run each execution subtask, and at the same time improve the utilization rate; Therefore, through the above pipeline-style task context switching setting, the video memory management and the operation of the main and standby working subtasks are unified, greatly reducing the switching overhead and improving the utilization rate of the GPU;
[0059] 2. Since only runtime loading and CUDA context initialization are performed during the environment initialization phase, the video memory occupied by the operation is small, and multiple tasks or execution subtasks can be initialized, improving the GPU utilization rate;
[0060] 3. Considering the impact of the video memory size, for the early preparation in the environment initialization stage of each task, it is also necessary to control its quantity. Therefore, a first preset value is set. When the video memory occupied by the task for environment initialization preparation or the execution of subtasks reaches the first preset value, the early preparation in the environment initialization stage is paused first. After the video memory is released during the task cleaning stage, the early preparation for environment initialization is carried out again, which is the time to reduce the switching overhead while maintaining a high GPU utilization rate currently;
[0061] 4. To improve the efficiency of the video memory application stage for each execution subtask in the AI program, through a preset video memory matching model, the type of the task and the division situation of the task are identified, and the required video memory size can be accurately pre-matched for each task or execution subtask that is about to run, thus improving the efficiency of the second preparation stage. Description of the Drawings
[0062] Figure 1 It is a prior art drawing in the background art of an application fast switching method applied to GPU computing power scheduling in the present application;
[0063] Figure 2 It is a schematic diagram of MPS switching in the background art of an application fast switching method applied to GPU computing power scheduling in the present application;
[0064] Figure 3 It is a flowchart of an implementation in an embodiment of an application fast switching method applied to GPU computing power scheduling in the present application;
[0065] Figure 4 It is a schematic diagram of the task execution stage in an embodiment of an application fast switching method applied to GPU computing power scheduling in the present application;
[0066] Figure 5 It is a schematic diagram of task switching in an embodiment of an application fast switching method applied to GPU computing power scheduling in the present application;
[0067] Figure 6 It is a schematic diagram of comparing switching modes in an embodiment of an application fast switching method applied to GPU computing power scheduling in the present application;
[0068] Figure 7 It is a schematic diagram of comparing throughput in an embodiment of an application fast switching method applied to GPU computing power scheduling in the present application. Detailed Embodiments
[0069] The most mainstream workloads include throughput-intensive training tasks and latency-sensitive inference tasks. The current mainstream approach is to provide dedicated GPU clusters for training and inference respectively. The main reason is that in large model architectures, inference tasks cannot use the training cluster for serving, and when the inference load is low, training tasks cannot utilize the inference cluster either.
[0070] Ideally, multiple DL applications should be able to be packed onto the same GPU server to maximize GPU utilization through time-sharing. This is exactly how the operating system achieves high CPU utilization through task scheduling and context switching. However, this cannot be achieved on GPUs because CPU switching is at the microsecond level, while GPU has a high overhead when switching between tasks, and the current switching method is at the second level. And current inferences have strict SLOs in the range of tens to hundreds of milliseconds, so they cannot be packed onto the same GPU server.
[0071] The existing solution is to let different models share the GPU video memory spatially, such as NVIDIA's Multi-Process Service (MPS). The tasks are all prepared in advance and the switching speed is fast enough. But the problem with this is that multiple models occupy a large amount of video memory space, resulting in a lower overall utilization rate of the GPU.
[0072] The following will Figure 3 - 7 further elaborate on this application in detail.
[0073] In one embodiment, as Figure 3 shown, this application discloses an application fast switching method applied to GPU computing power scheduling, which specifically includes the following steps:
[0074] S10: Divide a task of an AI program running on the GPU into several execution subtasks, and based on a preset preparation rule, divide the task preparation stage into a first preparation stage and a second preparation stage;
[0075] In this embodiment, the AI program model is stored in the host memory. The host memory is larger than the GPU memory and is also more economical. The GPU can quickly switch between contexts. Referring to Figure 4 , it is a schematic diagram of task execution in the AI program, which is a hierarchical execution method, including steps such as input, encoder, decoder, and output, so the task can be split.
[0076] The first preparation stage includes an environment initialization stage. The environment initialization stage is used for process startup, loading of the PyTorch CUDA runtime, and CUDA context initialization; the environment initialization only performs runtime loading and context initialization, occupying very little video memory, so that initialization work can be carried out for multiple tasks or execution subtasks.
[0077] The second preparation stage includes a video memory allocation stage and a data transfer stage. The video memory allocation stage is used to apply for a matching video memory for the execution subtasks in the AI program;
[0078] The data transfer stage is used to copy the model data of the AI program model and the tasks to be run on the GUP from the host memory to the video memory.
[0079] Furthermore, when applying for a matching video memory for the execution subtasks in the AI program, the steps include:
[0080] S11: Obtain all the execution subtasks in the AI program to be executed, and input the execution subtasks into the pre-trained video memory matching model in sequence;
[0081] S12: When the video memory matching model receives an execution subtask, identify the type information and task division information of the task to which the execution subtask belongs;
[0082] S13: Based on the type information and task division information, match a corresponding video memory application value for the current execution subtask; and when the video memory allocation stage of the current execution subtask is running, generate and send a video memory application instruction based on the video memory application value.
[0083] In this embodiment, the video memory matching model is a model that has been pre-trained through a neural network for matching the video memory of different execution subtasks. Different type information represents tasks with different video memory requirements, and the task division information refers to the number information of the execution subtasks into which the task is divided.
[0084] The video memory application value refers to the video memory value sufficient for the execution subtask to run.
[0085] Among them, step S13 includes:
[0086] S131: Based on the type information, screen out the video memory range corresponding to the type of the task from the preset database of the video memory matching model;
[0087] S132: Based on the task division information, identify the number of task divisions to which the current execution subtask belongs, and screen out the video memory that matches the number of task divisions from the video memory range as the video memory application value of the current execution subtask.
[0088] Specifically, based on the type information of the task, screen out the video memory range for the task to run from the preset database of the video memory matching model. This video memory range includes the video memory values corresponding to different division numbers of the task. Then, by identifying the number information of the execution subtasks into which the task is divided, select the video memory value that meets the division number from the video memory range as the video memory application value of the execution subtask.
[0089] Furthermore, there are still differences in the video memory values required between different execution subtasks in the same task. Through the field feature information of the execution subtasks, the video memory application values of different execution subtasks in the same task can be more accurately matched.
[0090] S20: Based on the preset operation rules, during the task preparation stage of the current task, run the first preparation stage of the next task.
[0091] Specifically, it includes the steps:
[0092] S21: Based on the preset operation rules, when initializing the running environment of the current task, synchronously run the environment initialization stage of the next task.
[0093] S22: When the video memory occupied by all tasks in the synchronously run environment initialization stage reaches the first preset value, pause running the environment initialization stage of new tasks.
[0094] S23: After the task cleaning stage runs, when the video memory occupied by all tasks in the synchronously run environment initialization stage is less than the first preset value again, restart running the environment initialization stage of the next task.
[0095] In this embodiment, the first preset value is set customarily. The task cleaning stage is used to delete the environment, delete the data of the executed subtasks that have been run, and release the video memory.
[0096] S30: When the task execution stage of the current task reaches the preset process, run the second preparation stage of the first execution subtask in the next task.
[0097] Specifically, it includes the steps:
[0098] S31: Obtain the number of execution subtasks in the current task and identify the quantity information of the unexecuted execution subtasks.
[0099] S32: When the quantity information is less than the preset quantity, obtain the video memory application information corresponding to each of all unexecuted execution subtasks.
[0100] S33: Based on the video memory application information. When the video memory applied for by any one of the unexecuted execution subtasks is greater than the second preset value, mark this unexecuted execution subtask.
[0101] S34: When the marked execution subtask finishes running, run the second preparation stage of the first execution subtask in the next task.
[0102] In this embodiment, from the video memory application model, it is possible to obtain the number of unexecuted subtasks in the current task and their corresponding video memory application values, that is, the video memory application information; the preset number and the second preset value can be customized.
[0103] Among them, when there are two or more marked execution subtasks, after all the execution subtasks are completed, the second preparation stage of the first execution subtask in the next task is run.
[0104] S40: After the second preparation stage of the first execution subtask in the next task ends, run the first execution subtask;
[0105] Specifically, it includes the steps:
[0106] S41: When the second preparation stage of the first execution subtask in the next task ends, determine whether the task execution stage of the current task has ended;
[0107] S42: If the task execution stage of the current task ends, immediately run the first execution subtask;
[0108] S43: If the task execution stage of the current task has not ended, determine whether the sum of the video memory applied for by the currently running execution subtask and the first execution subtask to be run exceeds the second preset value. If so, wait until the task execution stage of the current task ends and then run the first execution subtask.
[0109] In this embodiment, when the task execution stage of the current task has not ended, there are no marked self-executing subtasks, so only the unmarked execution subtasks are still in the running process. It is necessary to consider whether the sum of the video memory value of the current execution subtask and the video memory value of the first execution subtask to be run exceeds the second preset value. To determine whether the two execution subtasks can be run simultaneously to minimize the context switching overhead.
[0110] S50: During the running of the first execution subtask, run the second preparation stage of the next execution subtask and run the task cleaning stage to clear the tasks that have completed running;
[0111] In this embodiment, when switching the running between different execution subtasks in the same task, the second preset value also needs to be considered.
[0112] S60: When all the execution subtasks are completed, run the task cleaning stage.
[0113] In one embodiment, after all the execution subtasks in the same task have completed running, the task cleaning phase will clean the environmental information of the previous task. If there are differences in the environmental information of the execution subtasks in the same task, the environmental information will be cleared immediately after the completion of the running of one execution subtask.
[0114] In one embodiment, referring to Figure 5 as shown, the task is divided into three execution subtasks, G1, G2, and G3. When the old task is in the preparation phase, the first preparation phase of the new task, i.e., task preparation 1, starts; when the task execution phase of the old task reaches the preset process, the second preparation phase of the first execution subtask G1 in the new task, i.e., task preparation 2, starts; when the execution phase of the old task ends, the execution subtask G1 is run, i.e., task execution 1, and at the same time, the second preparation phase of the execution subtask G2, i.e., task preparation 2, is carried out; further, the execution task G2 is run, and at the same time, the second preparation phase of the execution subtask G3 and the task cleaning phase are carried out.
[0115] Refer to Figure 6 , through the comparison of the time consumption of the switching modes such as NVIDIA MPS switching, hard switching, and the pipeline switching of the present application with the ready models on three general models, and the throughput comparison of the three switching modes is as Figure 7 shown. It can be seen from Figure 7 that after adopting the pipeline switching mode, when the AI program on the GPU is switched, the time is greatly shortened, and the throughput can be maintained at a high level, while the existing methods such as the hard switching mode pauses for more than 6 seconds, and the MPS switching has a long time and occupies too much video memory.
[0116] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0117] In one embodiment, an application fast switching system applied to GPU computing power scheduling is provided. The application fast switching method applied to GPU computing power scheduling corresponds one-to-one with the application fast switching system applied to GPU computing power scheduling in the above embodiments. The application fast switching system applied to GPU computing power scheduling includes:
[0118] A division module, configured to divide a task of an AI program running on the GPU into several execution subtasks, and divide the task preparation phase into a first preparation phase and a second preparation phase based on a preset preparation rule;
[0119] A first running module, configured to run the first preparation phase of the next task during the task preparation phase of the current task based on a preset running rule;
[0120] A second operation module, configured to, when a preset process is reached during the task execution phase of the current task, run the second preparation phase of the first execution subtask in the next task;
[0121] A third operation module, configured to, after the second preparation phase of the first execution subtask in the next task ends, run the first execution subtask;
[0122] A fourth operation module, configured to, during the running of the first execution subtask, run the second preparation phase of the next execution subtask, and at the same time run a task cleaning phase to clear the tasks that have completed running;
[0123] A fifth operation module, configured to, after all execution subtasks have completed running, run the task cleaning phase.
[0124] For the specific limitations of an application fast switching system applied to GPU computing power scheduling, reference can be made to the limitations of an application fast switching system applied to GPU computing power scheduling in the above text, which will not be elaborated here. Each module in the above application fast switching system applied to GPU computing power scheduling can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0125] In one embodiment, a computer device is provided. The computer device can be a server. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an application fast switching method applied to GPU computing power scheduling.
[0126] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements an application fast switching method applied to GPU computing power scheduling.
[0127] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements an application fast switching method applied to GPU computing power scheduling.
[0128] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0129] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0130] The above embodiments are only used to illustrate the technical solutions of the present application, not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for fast application switching for GPU computing power scheduling, characterized in that: Divide a task of an AI program running on a GPU into several execution subtasks, and divide the task preparation phase into a first preparation phase and a second preparation phase based on a preset preparation rule; The first preparation stage includes an environment initialization stage, which is used to start multiple task processes and load the PyTorch CUDA runtime and initialize the CUDA context; The second preparation stage includes a video memory allocation stage and a data transmission stage, wherein the video memory allocation stage is used to apply for matching video memory for the execution subtask in the AI program, and the AI program is stored in a preset host memory; The data transmission stage is used to copy the model data of the AI program model and the tasks that need to be run on the GPU from the host memory to the video memory; Based on the preset running rules, when the current task starts running the environment initialization phase, the environment initialization phase of the next task is run synchronously; When the task execution phase of the current task reaches a preset progress, the second preparation phase of the first execution subtask in the next task is executed; When the second preparation phase of the first execution subtask in the next task is completed, the first execution subtask is executed; During the running of the first execution subtask, the second preparation phase of the next execution subtask is run, and at the same time, the task cleaning phase is run to clear the tasks that have been completed; When all execution subtasks have finished running, the task cleaning phase is run.
2. The method for fast application switching for GPU computing power scheduling according to claim 1, characterized in that: The method based on the preset operation rules, when the current task starts to run the environment initialization phase, synchronously runs the environment initialization phase of the next task; includes: When the video memory occupied by all tasks in the synchronous running environment initialization phase reaches a first preset value, the environment initialization phase of running new tasks is suspended; After the task cleaning phase is executed, when the video memory occupied by all tasks in the synchronously executed environment initialization phase is less than the first preset value again, the environment initialization phase of the next task is executed again.
3. The method for fast application switching for GPU computing power scheduling according to claim 1, characterized in that: When the task execution phase of the current task reaches a preset progress, the second preparation phase of the first execution subtask in the next task is run; comprising: Get the number of subtasks executed in the current task, and identify the number of subtasks that have not been executed; When the quantity information is less than the preset quantity, the corresponding video memory application information of all the execution subtasks that have not been run is obtained; Based on the video memory application information, when the video memory applied by any of the unexecuted execution subtasks is greater than a second preset value, marking the unexecuted execution subtask; When the marked execution subtask is completed, the second preparation phase of the first execution subtask in the next task is executed.
4. The method for fast application switching for GPU computing power scheduling according to claim 3, characterized in that: After the second preparation phase of the first execution subtask in the next task is completed, the first execution subtask is executed; comprising: When the second preparation phase of the first execution subtask in the next task is finished, determining whether the task execution phase of the current task is finished; If the task execution phase of the current task is completed, the first execution subtask is immediately run; If the task execution phase of the current task has not ended, determine whether the sum of the video memory requested by the currently running execution subtask and the first execution subtask to be run exceeds a second preset value. If so, wait until the task execution phase of the current task ends and run the first execution subtask.
5. The method for fast application switching for GPU computing power scheduling according to claim 1, characterized in that: The video memory allocation stage is used to apply for matching video memory for the execution subtask in the AI program, including: Obtain all execution subtasks in the AI program to be executed, and input the execution subtasks into the pre-trained video memory matching model in sequence; When the video memory matching model receives the execution subtask, it identifies the type information and task division information of the task to which the execution subtask belongs; Based on the type information and the task division information, a corresponding video memory application value is matched for the currently executed subtask; and when the video memory allocation phase of the currently executed subtask is running, a video memory application instruction is generated and sent based on the video memory application value.
6. The method for fast application switching for GPU computing power scheduling according to claim 5, characterized in that: The method of matching the corresponding video memory application value for the currently executed subtask based on the type information and the task division information includes: Based on the type information, a video memory interval corresponding to the type of the task is selected from a preset database of a video memory matching model; The task division number of the task to which the currently executed subtask belongs is identified based on the task division information, and a video memory matching the task division number is selected from the video memory interval as the video memory application value of the currently executed subtask.
7. A fast application switching system for GPU computing power scheduling, characterized in that: A partitioning module is used to divide a task of an AI program running on a GPU into a plurality of execution subtasks, and divide the task preparation phase into a first preparation phase and a second preparation phase based on a preset preparation rule; the first preparation phase includes an environment initialization phase, and the environment initialization phase is used to start multiple task processes and load the PyTorch CUDA runtime and initialize the CUDA context; The second preparation stage includes a video memory allocation stage and a data transmission stage, wherein the video memory allocation stage is used to apply for matching video memory for the execution subtask in the AI program, and the AI program is stored in a preset host memory; The data transmission stage is used to copy the model data of the AI program model and the tasks that need to be run on the GPU from the host memory to the video memory; A first running module is used to synchronously run the environment initialization phase of the next task when the current task starts to run the environment initialization phase based on a preset running rule; A second running module, used for running the second preparation phase of the first execution subtask in the next task when the task execution phase of the current task reaches a preset process; A third running module is used to run the first execution subtask in the next task after the second preparation phase of the first execution subtask is completed; A fourth running module is used to run the second preparation phase of the next execution subtask during the execution of the first execution subtask, and to run the task cleaning phase to clear the tasks that have been completed; The fifth running module is used to run the task cleaning phase after all execution subtasks are completed.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of a method for fast switching of applications applied to GPU computing power scheduling as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, the steps of a method for fast switching of applications applied to GPU computing power scheduling as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Computing power resource allocation method and device, electronic equipment and storage medium
CN116909748A