GPU resource dynamic scheduling method and device for artificial intelligence model
Through task awareness mechanism and real-time monitoring of GPU load, GPU resources are dynamically managed, which solves the problems of low resource utilization, high cost and insufficient flexibility in the existing technology, and realizes efficient GPU resource sharing and flexible resource management.
Patent Information
- Application Number
- CN202510136898.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The existing cloud host rental services adopt exclusive GPU resource allocation solutions, resulting in low resource utilization, high user costs and insufficient flexibility, and the inability to dynamically adjust GPU resource usage according to the actual task needs.
Through the task perception mechanism, identify the resource requirements of artificial intelligence model tasks for GPUs, monitor the GPU load situation in real time, dynamically mount and unload GPU resources, and support a variety of custom recycling strategies.
The time-sharing of GPU resources is realized, which significantly improves GPU utilization, reduces user usage costs, and enhances the flexibility and applicability of the system.
Smart Images

Figure CN120066782A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer information processing, and more specifically, to a method and device for dynamically scheduling GPU resources for an artificial intelligence model. Background Art
[0002] With the rapid development of artificial intelligence technology, tools for generating images and videos (such as ComfyUI, SD-WebUI, Foocus, etc.) have gradually become important tools for individual users and developers to carry out creative design and content generation. These tools are usually based on deep learning models and rely on powerful GPU computing resources to complete complex image and video generation tasks. To meet the needs of users, most current cloud host rental services provide pre-configured images of these tools, and users can directly rent GPU hosts and use these tools to process tasks.
[0003] However, existing cloud host rental services generally adopt an exclusive GPU resource allocation scheme, that is, after a user rents a GPU host, regardless of whether the GPU resources are actually used, the user will monopolize all resources of the host. This resource allocation method has the following problems in practical applications: Low resource utilization: When users use tools for generating images and videos, they usually go through multiple stages, including setting up the software environment, adjusting and designing workflow parameters, running the workflow, and exporting images or videos. Among them, the stages of setting up the software environment and adjusting workflow parameters do not require GPU resources, the running workflow stage must use GPU resources, and the stage of exporting images or videos may require GPU resources. Under the existing exclusive GPU scheme, users will occupy GPU resources throughout the task process, even if they do not need GPU computing power in some stages, resulting in a large amount of GPU resources being idle and reducing resource utilization.
[0004] High user cost: Since users need to pay for GPU resources during the entire rental period, even if they do not need to use the GPU in some stages, users still have to bear high usage costs. This resource allocation method is not economical for users, especially for small and medium-sized users or individual developers, who face greater cost pressure.
[0005] Lack of flexibility: The existing exclusive GPU allocation scheme cannot dynamically adjust the use of GPU resources according to the actual needs of tasks. When users design workflows or install software, GPU resources are occupied ineffectively and cannot be released for other users to use, resulting in a lack of flexibility in resource allocation.
[0006] Therefore, a new method and device for dynamically scheduling GPU resources for an artificial intelligence model are needed.
[0007] The above information disclosed in the background section is only used to enhance the understanding of the background of the present application. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0008] In view of this, the present application provides a method and device for dynamically scheduling GPU resources for an artificial intelligence model, which can achieve time-sharing sharing of GPU resources, significantly improve GPU utilization, reduce user usage costs, and at the same time support multiple custom recycling strategies, enhancing the flexibility and applicability of the system.
[0009] Other features and advantages of the present application will become apparent from the following detailed description, or be learned in part through the practice of the present application.
[0010] According to one aspect of the present application, a method for dynamically scheduling GPU resources for an artificial intelligence model is proposed. The method includes: identifying the resource requirements of the artificial intelligence model task for the GPU through a task awareness mechanism; during the execution of the artificial intelligence model task, monitoring the GPU load situation in real time; determining the resource requirements of the GPU according to the GPU load situation to dynamically mount GPU resources; executing the artificial intelligence model task through the GPU resources to obtain an output result; and dynamically unloading the GPU resources when the GPU recycling strategy is met.
[0011] In an exemplary embodiment of the present application, before identifying the resource requirements of the artificial intelligence model task for the GPU, it further includes: building a software environment for the artificial intelligence model; setting parameters and the workflow of the artificial intelligence model task.
[0012] In an exemplary embodiment of the present application, identifying the resource requirements of the artificial intelligence model task for the GPU through a task awareness mechanism includes: before the execution of the artificial intelligence model task, analyzing the task type and computing requirements of the artificial intelligence model task through the task awareness mechanism; and determining whether there are GPU resource requirements according to the task type and the computing requirements.
[0013] In an exemplary embodiment of the present application, during the execution of the artificial intelligence model task, monitoring the GPU load situation in real time includes: during the execution of the artificial intelligence model task, monitoring the usage status, computing load, and task execution progress of the GPU in real time.
[0014] In an exemplary embodiment of the present application, determining the resource requirements of the GPU according to the GPU load condition to dynamically mount GPU resources includes: judging the current GPU resource requirements according to the GPU load condition; dynamically mounting GPU resources according to the current GPU resource requirements through custom resource definitions and custom controllers in the Kubernetes cluster.
[0015] In an exemplary embodiment of the present application, dynamically mounting GPU resources according to the current GPU resource requirements through custom resource definitions and custom controllers in the Kubernetes cluster includes: applying to mount GPU resources through the Pod interface in the Kubernetes cluster according to the current GPU resource requirements.
[0016] In an exemplary embodiment of the present application, performing the artificial intelligence model task through the GPU resources to obtain an output result includes: obtaining the workflow corresponding to the artificial intelligence model task; during the execution of the workflow, performing calculations based on the GPU resources to obtain an output result.
[0017] In an exemplary embodiment of the present application, dynamically unloading the GPU resources when the GPU recycling policy is satisfied includes: dynamically unloading the GPU resources after the artificial intelligence model task ends; and / or dynamically unloading the GPU resources when the time policy is satisfied; and / or dynamically unloading the GPU resources when the data logging policy is satisfied.
[0018] In an exemplary embodiment of the present application, dynamically unloading the GPU resources includes: dynamically unloading the GPU resources through the Pod interface in the Kubernetes cluster.
[0019] According to one aspect of the present application, there is provided a dynamic scheduling device for GPU resources for an artificial intelligence model, the device including: a sensing module for identifying the resource requirements of the artificial intelligence model task for the GPU through a task sensing mechanism; a monitoring module for real-time monitoring of the GPU load condition during the execution of the artificial intelligence model task; a mounting module for determining the resource requirements of the GPU according to the GPU load condition to dynamically mount GPU resources; a task module for performing the artificial intelligence model task through the GPU resources to obtain an output result; an unloading module for dynamically unloading the GPU resources when the GPU recycling policy is satisfied.
[0020] According to one aspect of the present application, there is provided an electronic device, the electronic device including: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method as described above.
[0021] According to one aspect of the present application, there is provided a computer-readable medium storing a computer program thereon, and when the program is executed by a processor, the above-described method is implemented.
[0022] According to the method and apparatus for dynamically scheduling GPU resources for an artificial intelligence model of the present application, through a task perception mechanism, the resource requirements of the artificial intelligence model task for the GPU are identified; during the execution of the artificial intelligence model task, the GPU load condition is monitored in real time; according to the GPU load condition, the resource requirements of the GPU are determined to dynamically mount GPU resources; the artificial intelligence model task is executed through the GPU resources to obtain an output result; when the GPU recycling policy is satisfied, the GPU resources are dynamically unmounted, which can achieve time-sharing sharing of GPU resources, significantly improve the GPU utilization rate, reduce the user's usage cost, and at the same time support a variety of custom recycling policies, enhancing the flexibility and applicability of the system.
[0023] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] By referring to the drawings and describing its exemplary embodiments in detail, the above and other objects, features and advantages of the present application will become more apparent. The following described drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0025] Figure 1 is a flowchart of a method for dynamically scheduling GPU resources for an artificial intelligence model shown according to an exemplary embodiment.
[0026] Figure 2 is a flowchart of a method for dynamically scheduling GPU resources for an artificial intelligence model shown according to another exemplary embodiment.
[0027] Figure 3 is a schematic diagram of a method for dynamically scheduling GPU resources for an artificial intelligence model shown according to another exemplary embodiment.
[0028] Figure 4 is a block diagram of an apparatus for dynamically scheduling GPU resources for an artificial intelligence model shown according to an exemplary embodiment.
[0029] Figure 5 is a block diagram of an electronic device shown according to an exemplary embodiment.
[0030] Figure 6 is a block diagram of a computer-readable medium shown according to an exemplary embodiment. Detailed implementation manners
[0031] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repeated description will be omitted.
[0032] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will realize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this application.
[0033] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0034] The flowcharts shown in the drawings are merely illustrative and do not necessarily include all the content and operations / steps, nor do they necessarily have to be executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0035] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first component discussed below can be referred to as the second component without departing from the teachings of the concept of this application. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more of them.
[0036] Those skilled in the art can understand that the drawings are only schematic diagrams of the example embodiments, and the modules or processes in the drawings are not necessarily essential for implementing this application, and thus cannot be used to limit the protection scope of this application.
[0037] The technical abbreviations related to this application are explained as follows: GPU (Graphics Processing Unit), the graphics processing unit, is a hardware device specifically designed for processing graphics and parallel computing. In modern artificial intelligence and deep learning tasks, GPUs are widely used to accelerate large-scale matrix operations and neural network training.
[0038] Kubernetes (k8s), an open-source container orchestration platform, is used for automating the deployment, scaling, and management of containerized applications. Kubernetes provides powerful resource scheduling and management capabilities, enabling efficient management of computing resources in distributed systems.
[0039] CRD (Custom Resource Definition), custom resource definition, is an extension mechanism in Kubernetes that allows users to define their own resource types. Through CRD, users can create and manage custom resource objects in Kubernetes to meet specific business requirements.
[0040] Pod, the smallest scheduling unit in Kubernetes, usually contains one or more containers. A Pod is the basic unit for deploying and managing applications in Kubernetes, and each Pod shares the same network and storage resources.
[0041] ComfyUI, a popular tool for generating images and videos, is usually used for inference and generation tasks of artificial intelligence models. ComfyUI provides a user-friendly interface to help users design and execute complex workflows.
[0042] SD-WebUI, Stable Diffusion Web User Interface, is a web interface tool based on the StableDiffusion model for generating high-quality images and videos. Users can easily configure and run generation tasks through the web interface.
[0043] Foocus, a tool for generating images and videos, is usually used for inference and generation tasks of artificial intelligence models. Foocus provides a simple and easy-to-use interface to help users quickly generate image and video content.
[0044] AI (Artificial Intelligence), artificial intelligence, refers to the technology of simulating human intelligence through computers. AI technology is widely applied in fields such as image generation, natural language processing, and speech recognition.
[0045] Load. In a computer system, load usually refers to the current workload or resource usage of a system or hardware device (such as a GPU). Load monitoring is an important basis for resource scheduling and management.
[0046] Figure 1 It is a flowchart of a method for dynamically scheduling GPU resources for an artificial intelligence model shown according to an exemplary embodiment. The method 10 for dynamically scheduling GPU resources for an artificial intelligence model at least includes steps S102 to S110.
[0047] As Figure 1 shown, in S102, through a task awareness mechanism, identify the resource requirements of the artificial intelligence model task for the GPU. For example, before the execution of the artificial intelligence model task, analyze the task type and computing requirements of the artificial intelligence model task through the task awareness mechanism; judge whether there are GPU resource requirements according to the task type and the computing requirements.
[0048] More specifically, in the present application, the task awareness mechanism is a dynamic scheduling technology that can automatically identify the computing resource (such as GPU) requirements of a task according to the characteristics and requirements of the task. By analyzing factors such as task type, computing complexity, and data scale, the task awareness mechanism can judge whether GPU resources are needed.
[0049] Before the execution of the artificial intelligence model task, the system will analyze the task type (such as image generation, video processing, model training, etc.) and computing requirements (such as floating-point operation volume, memory occupancy, etc.). According to the task type and computing requirements, judge whether GPU resources are needed. For example, image generation tasks usually require GPU acceleration, while simple parameter adjustment tasks may not.
[0050] Among them, before identifying the resource requirements of the artificial intelligence model task for the GPU, it further includes: building the software environment of the artificial intelligence model; setting parameters and the workflow of the artificial intelligence model task.
[0051] In a practical application, before the task execution, the user needs to build a software environment suitable for the operation of the artificial intelligence model, including installing deep learning frameworks (such as TensorFlow, PyTorch), dependency libraries (such as CUDA, cuDNN), and necessary toolchains.
[0052] In a practical application, the user needs to configure the parameters of the task (such as model hyperparameters, input data path, etc.) and design the workflow. The workflow refers to the steps and logic of task execution, usually defined in the form of scripts or configuration files.
[0053] In S104, during the execution of the artificial intelligence model task, the GPU load situation is monitored in real time. For example, during the execution of the artificial intelligence model task, the usage status, computing load, and task execution progress of the GPU are monitored in real time.
[0054] GPU load refers to the current workload of the GPU, which is usually measured by indicators such as the GPU utilization rate, memory occupancy rate, and computing load. Monitoring the GPU load in real time can help the system dynamically adjust resource allocation.
[0055] More specifically, during the task execution process, the system will collect the usage status of the GPU in real time (such as GPU utilization rate, memory occupancy rate, temperature, etc.). At the same time, the system will monitor the execution progress of the task (such as the amount of computation completed, the remaining amount of computation, etc.) to determine whether the current GPU resources meet the task requirements.
[0056] In practical applications, the monitoring indicators may include: GPU utilization rate: It represents the utilization rate of the GPU computing core, usually expressed as a percentage.
[0057] GPU memory occupancy rate: It represents the usage situation of the GPU video memory, usually expressed as a percentage.
[0058] Computing load: It represents the current computing task volume of the GPU, usually measured by the floating-point operation volume (FLOPS).
[0059] In S106, the resource requirements of the GPU are determined according to the GPU load situation to dynamically mount GPU resources. For example, the current GPU resource requirements are judged according to the GPU load situation; according to the current GPU resource requirements, the GPU resources are dynamically mounted through custom resource definitions and custom controllers in the Kubernetes cluster. Dynamically mounting GPU resources means that according to the actual needs of the task, GPU resources are temporarily allocated during the task execution process. When the task requires GPU acceleration, the system will automatically mount the GPU resources; when the task does not require the GPU, the system will unmount the GPU resources.
[0060] According to the GPU load situation, the system will judge whether the current task requires additional GPU resources. For example, when the GPU utilization rate is close to 100%, the system may mount more GPU resources.
[0061] More specifically, the GPU resources can be applied for mounting through the Pod interface in the Kubernetes cluster according to the current GPU resource requirements. Through the custom resource definition (CRD) and custom controller in the Kubernetes cluster, the system can dynamically manage the mounting and unmounting of GPU resources.
[0062] In S108, execute the artificial intelligence model task through the GPU resources to obtain an output result. For example, obtain the workflow corresponding to the artificial intelligence model task; during the execution of the workflow, perform calculations based on the GPU resources to obtain the output result.
[0063] After the GPU resources are mounted, the system will execute the artificial intelligence model task based on the GPU resources. The parallel computing ability of the GPU can significantly accelerate the execution of the task.
[0064] More specifically, the system will obtain the workflow corresponding to the task and execute the task according to the steps of the workflow. During the execution of the workflow, the system will use the GPU resources for calculations (such as matrix operations, neural network inferences, etc.) to generate output results (such as images, videos, model parameters, etc.).
[0065] In S110, when the GPU recycling policy is satisfied, dynamically unmount the GPU resources. For example, dynamically unmount the GPU resources through the Pod interface in the Kubernetes cluster. The GPU recycling policy refers to the rules for the system to automatically unmount the GPU resources under specific conditions.
[0066] More specifically, through the Pod interface in the Kubernetes cluster, the system can dynamically unmount the GPU resources. For example, when the task ends, the system will call the Pod interface to release the GPU resources.
[0067] In one embodiment, the pseudo-code for mounting or unmounting the GPU is as follows: Plain Text device_id = request_load_gpu(type) result = reqeust_offload_gpu(device_id) According to the GPU resource dynamic scheduling method for the artificial intelligence model of the present application, through the task awareness mechanism, identify the resource requirements of the artificial intelligence model task for the GPU; during the execution of the artificial intelligence model task, real-time monitor the GPU load situation; determine the resource requirements of the GPU according to the GPU load situation to dynamically mount the GPU resources; execute the artificial intelligence model task through the GPU resources to obtain an output result; when the GPU recycling policy is satisfied, dynamically unmount the GPU resources, which can realize the time-sharing sharing of GPU resources, significantly improve the GPU utilization rate, reduce the user usage cost, and at the same time support multiple custom recycling policies, enhancing the flexibility and applicability of the system.
[0068] It should be clearly understood that this application describes how to form and use specific examples, but the principles of this application are not limited to any details of these examples. Instead, based on the teachings disclosed in this application, these principles can be applied to many other embodiments.
[0069] Figure 2 It is a flowchart of a method for dynamically scheduling GPU resources for an artificial intelligence model shown according to another exemplary embodiment. Figure 2 The shown process 20 is a detailed description of Figure 1 S110 "dynamically unload the GPU resources when the GPU recycling policy is met" in the shown process.
[0070] As Figure 3 shown, for example, when the GPU recycling policy is met, the GPU resources are dynamically unloaded.
[0071] In a practical application, assume that the user is using ComfyUI for an image generation task. During the task execution, the system will dynamically mount GPU resources according to the GPU load. When the task ends, the system will automatically unload the GPU resources so that other users can use these resources.
[0072] For example, after the artificial intelligence model task ends, the GPU resources are dynamically unloaded.
[0073] The task end policy means that when the artificial intelligence model task is completed, the system automatically unloads the GPU resources. This policy is applicable to the situation where the task execution time is short and the GPU resources are no longer needed after the task ends.
[0074] More specifically, the system will monitor the execution status of the task. When the task is completed, the system will call the Pod interface in the Kubernetes cluster to dynamically unload the GPU resources.
[0075] In a practical application, the user uses SD-WebUI for a video generation task. During the task execution, the system mounts the GPU resources to accelerate the video generation. When the video generation task is completed, the system will automatically unload the GPU resources and release them for other users to use.
[0076] For example, when the time policy is met, the GPU resources are dynamically unloaded. The time policy means that when the GPU resources are not used within a preset time, the system automatically unloads the GPU resources. This policy is applicable to the situation where the task execution time is uncertain or the GPU resources may be idle for a long time.
[0077] More specifically, the system will monitor the usage of the GPU resources. If the GPU resources are not used within a preset time (such as 10 minutes), the system will automatically unload the GPU resources.
[0078] In a practical application, the user uses Foocus for image generation tasks. During the task execution, the user pauses the task to adjust parameters. If the GPU resources are not used within 10 minutes, the system will automatically unload the GPU resources to avoid resource waste.
[0079] For example, when the data logging strategy is met, the GPU resources are dynamically unloaded. The data logging strategy means that the user manually triggers the unloading of GPU resources by calling a preset interface during task execution. This strategy is applicable when the user needs to fully control the use of GPU resources.
[0080] More specifically, the system will provide preset interfaces (such as APIs), and the user can call these interfaces during task execution to unload GPU resources. For example, the user can call the interface to unload GPU resources after the task is completed.
[0081] In a practical application, the user uses ComfyUI for image generation tasks. During the task execution, the user finds that the task has been completed, but the system has not automatically unloaded the GPU resources. At this time, the user can call the preset interface to manually unload the GPU resources so that other users can use these resources.
[0082] In addition, there can be various other strategies: Fixed-duration strategy: Unload the GPU if not needed within a certain number of minutes in the future. Under this strategy, it is equivalent to the traditional cloud host rental strategy.
[0083] Custom data logging strategy: Call the interfaces pre-specified by the platform at the places where GPU needs to be mounted and unmounted to fully control the life cycle of the GPU by oneself.
[0084] Figure 3 It is a schematic diagram of a method for dynamically scheduling GPU resources for an artificial intelligence model shown according to another exemplary embodiment. Figure 3 It shows a management process for the entire life cycle of a GPU, such as Figure 3 shown. During the actual task execution, it can be performed, for example, according to the following steps.
[0085] Step 1: Set up the software environment. This stage does not occupy GPU resources and only uses CPU resources.
[0086] Step 2: Adjust and design the workflow and workflow parameters. This stage does not occupy GPU resources and only uses CPU resources; Step 3: Dynamically mount GPU resources through a task awareness mechanism before task execution; Step 4: Run the workflow, which requires the use of GPU resources in this stage; the utilization rate of GPU resources can be monitored in this stage so that the GPU resources can be dynamically unloaded after the task ends or when the occupancy duration exceeds the threshold. Step 5: Obtain the output result, which can be, for example, an exported image or video, and this stage may occupy GPU resources. After unloading the GPU resources or after obtaining the output result, return to Step 2 to continue adjusting and designing the workflow and workflow parameters, forming a loop task process.
[0087] Those skilled in the art can understand that all or part of the steps of implementing the above embodiments are implemented as a computer program executed by a CPU. When the computer program is executed by the CPU, the above functions defined by the above method provided in this application are executed. The program can be stored in a computer-readable storage medium, and the storage medium can be a read-only memory, a disk, an optical disc, etc.
[0088] In addition, it should be noted that the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0089] The following is an embodiment of the device of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the method embodiment of the present application.
[0090] Figure 4 is a block diagram of a device for dynamically scheduling GPU resources for an artificial intelligence model shown according to an exemplary embodiment. As Figure 4 shown, the device 40 for dynamically scheduling GPU resources for an artificial intelligence model includes: a perception module 402, a monitoring module 404, a mounting module 406, a task module 408, and an unloading module 410.
[0091] The perception module 402 is used to identify the resource requirements of the artificial intelligence model task for the GPU through a task perception mechanism; the perception module 402 is also used to analyze the task type and computing requirements of the artificial intelligence model task through the task perception mechanism before the execution of the artificial intelligence model task; and determine whether there are GPU resource requirements according to the task type and the computing requirements.
[0092] The monitoring module 404 is used to monitor the GPU load situation in real time during the execution of the artificial intelligence model task; the monitoring module 404 is also used to monitor the usage status, computing load, and task execution progress of the GPU in real time during the execution of the artificial intelligence model task.
[0093] The mounting module 406 is used to determine the resource requirements of the GPU according to the GPU load condition to dynamically mount GPU resources; the mounting module 406 is also used to judge the current GPU resource requirements according to the GPU load condition; and dynamically mount GPU resources through custom resource definitions and custom controllers in the Kubernetes cluster according to the current GPU resource requirements.
[0094] The task module 408 is used to execute the artificial intelligence model task through the GPU resources to obtain an output result; the task module 408 is also used to obtain the workflow corresponding to the artificial intelligence model task; and perform calculations based on the GPU resources during the execution of the workflow to obtain an output result.
[0095] The unloading module 410 is used to dynamically unload the GPU resources when the GPU recycling policy is satisfied. The unloading module 410 is also used to dynamically unload the GPU resources after the artificial intelligence model task ends; the unloading module 410 is also used to dynamically unload the GPU resources when the time policy is satisfied; the unloading module 410 is also used to dynamically unload the GPU resources when the data logging policy is satisfied.
[0096] The GPU resource dynamic scheduling device for an artificial intelligence model according to the present application can identify the resource requirements of the artificial intelligence model task for the GPU through a task perception mechanism; during the execution of the artificial intelligence model task, the GPU load condition is monitored in real time; determine the resource requirements of the GPU according to the GPU load condition to dynamically mount GPU resources; execute the artificial intelligence model task through the GPU resources to obtain an output result; and dynamically unload the GPU resources when the GPU recycling policy is satisfied, so as to realize time-sharing sharing of GPU resources, significantly improve the GPU utilization rate, reduce the user usage cost, and at the same time support multiple custom recycling policies, enhancing the flexibility and applicability of the system.
[0097] Figure 5 It is a block diagram of an electronic device shown according to an exemplary embodiment.
[0098] Next, refer to Figure 5 to describe the electronic device 500 according to this embodiment of the present application. Figure 5 The shown electronic device 500 is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of the present application.
[0099] As Figure 5As shown, the electronic device 500 is presented in the form of a general-purpose computing device. The components of the electronic device 500 may include, but are not limited to: at least one processing unit 510, at least one storage unit 520, a bus 530 connecting different system components (including the storage unit 520 and the processing unit 510), a display unit 540, etc.
[0100] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 510, so that the processing unit 510 executes the steps according to various exemplary embodiments of the present application described in this specification. For example, the processing unit 510 can execute steps such as Figure 1 , Figure 2 shown in.
[0101] The storage unit 520 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 5201 and / or a cache storage unit 5202, and may further include a read-only storage unit (ROM) 5203.
[0102] The storage unit 520 may further include a program / utility 5204 having a set (at least one) of program modules 5205. Such program modules 5205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0103] The bus 530 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0104] The electronic device 500 can also communicate with one or more external devices 500' (such as a keyboard, a pointing device, a Bluetooth device, etc.), so that the user can communicate with the electronic device 500, and / or the electronic device 500 can communicate with any device that can communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 550. And, the electronic device 500 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 560. The network adapter 560 can communicate with other modules of the electronic device 500 through the bus 530. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0105] From the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, as Figure 6 shown, the technical solution according to the embodiment of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiment of the present application.
[0106] The software product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0107] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0108] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0109] The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the computer-readable medium realizes the following functions: identifying the resource requirements of the artificial intelligence model task for the GPU through a task awareness mechanism; during the execution of the artificial intelligence model task, monitoring the GPU load situation in real time; determining the resource requirements of the GPU according to the GPU load situation to dynamically mount GPU resources; executing the artificial intelligence model task through the GPU resources to obtain an output result; and dynamically unloading the GPU resources when the GPU recycling policy is met.
[0110] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are uniquely different from this embodiment. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.
[0111] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described here can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of this application.
[0112] The above specifically shows and describes the exemplary embodiments of this application. It should be understood that this application is not limited to the detailed structures, setting methods, or implementation methods described here; on the contrary, this application is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.
Claims
1. A GPU resource dynamic scheduling method for an artificial intelligence model, characterized in that: include: Identify the GPU resource requirements of AI model tasks through task awareness mechanism; During the execution of the artificial intelligence model task, real-time monitoring of GPU load; Determining the resource requirements of the GPU according to the GPU load to dynamically mount GPU resources; Execute the artificial intelligence model task by using the GPU resources to obtain an output result; When the GPU recycling policy is met, the GPU resources are dynamically unloaded.
2. The method according to claim 1, characterized in that Before identifying the GPU resource requirements of AI model tasks, it also includes: Build the software environment for artificial intelligence models; Set the parameters and workflow of the AI model task.
3. The method according to claim 1, characterized in that Through the task awareness mechanism, the GPU resource requirements of AI model tasks are identified, including: Before the artificial intelligence model task is executed, analyzing the task type and computing requirements of the artificial intelligence model task through a task perception mechanism; Determine whether there is a GPU resource requirement based on the task type and the computing requirement.
4. The method according to claim 1, characterized in that During the execution of the artificial intelligence model task, the GPU load is monitored in real time, including: During the execution of the artificial intelligence model task, the GPU usage status, computing load and task execution progress are monitored in real time.
5. The method according to claim 1, characterized in that Determining the resource requirement of the GPU according to the GPU load condition to dynamically mount the GPU resources includes: Determine the current GPU resource requirement according to the GPU load condition; Dynamically mount GPU resources through custom resource definitions and custom controllers in the Kubernetes cluster based on current GPU resource requirements.
6. The method according to claim 5, characterized in that Dynamically mount GPU resources based on current GPU resource requirements through custom resource definitions and custom controllers in the Kubernetes cluster, including: Apply for mounting GPU resources through the Pod interface in the Kubernetes cluster based on the current GPU resource requirements.
7. The method according to claim 1, characterized in that Executing the artificial intelligence model task through the GPU resources to obtain output results, including: Obtaining the workflow corresponding to the artificial intelligence model task; During the execution of the workflow, calculations are performed based on GPU resources to obtain output results.
8. The method according to claim 1, characterized in that When the GPU recycling policy is met, the GPU resources are dynamically unloaded, including: After the artificial intelligence model task is completed, dynamically unloading the GPU resources; and / or Dynamically unloading the GPU resources when the time policy is met; and / or When the tracking point strategy is met, the GPU resources are dynamically unloaded.
9. The method according to claim 8, characterized in that Dynamically unloading the GPU resources, including: The GPU resources are dynamically unloaded through the Pod interface in the Kubernetes cluster.
10. A GPU resource dynamic scheduling device for an artificial intelligence model, characterized in that: include: The perception module is used to identify the GPU resource requirements of artificial intelligence model tasks through the task perception mechanism; A monitoring module, used to monitor the GPU load in real time during the execution of the artificial intelligence model task; A mounting module, used to determine the resource requirements of the GPU according to the GPU load to dynamically mount GPU resources; A task module, used to execute the artificial intelligence model task through the GPU resources to obtain an output result; The unloading module is used to dynamically unload the GPU resources when the GPU recycling policy is met.
Citation Information
Patent Citations
GPU card dynamic adjustment method and device, equipment and storage medium
CN113626182A
Cloud computing scheduling system and cloud computing scheduling method
CN118747121A
Dynamic computing power distribution method and system based on reinforcement learning
CN121900959A
User application store, platform, network, discovery, presentation and connecting with or synchronizing or accessing user data of users from / to parent application
WO2022079543A2