Method and device for scheduling GPU resources in AI computing

By obtaining task execution information of multiple GPU devices, determining the target GPU device and formulating task execution strategies, the problem of unreasonable GPU resource scheduling in AI computing is solved, and computing efficiency and accuracy are improved.

CN119806845BActive Publication Date: 2025-05-23SHENZHEN HUMENG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510293261.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-23
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

In AI computing, unreasonable scheduling of GPU resources will affect computing efficiency and accuracy.

Method used

By obtaining task execution information of multiple GPU devices, including the execution of real AI computing tasks and virtual AI computing tasks, the target GPU device is determined, and the task execution strategy is determined based on the information of AI computing tasks and target GPU device, in order to optimize the use of GPU resources.

Benefits of technology

It realizes reasonable scheduling of GPU resources, improves the efficiency and accuracy of AI computing, and ensures that GPU resources have a certain amount of task execution margin.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806845B_ABST
    Figure CN119806845B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and device for scheduling GPU resources in AI computing, and relates to the field of data processing technology; obtaining information representing the execution status of real AI computing tasks and the execution status of virtual AI computing tasks respectively, using the information representing the execution status of real AI computing tasks, combined with the actual situation of AI computing tasks, determining a target GPU device from multiple GPU devices, and combining the actual situation of AI computing tasks and the execution status of virtual AI computing tasks of the target GPU device, determining a task execution strategy so that the target GPU device can execute the AI ​​computing task. By configuring real AI computing tasks and virtual AI computing tasks in GPU devices, the execution status of virtual AI computing tasks can be used as reference information for scheduling GPU resources, and it is also ensured that GPU resources have a certain task execution margin, thereby realizing reasonable scheduling of GPU resources and improving the efficiency and accuracy of AI computing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular, to a method and device for scheduling GPU resources in AI computing. Background Art

[0002] With the development of AI (Artificial Intelligence) technology, AI computing is being used more and more widely. AI computing can improve computing accuracy and efficiency. AI computing requires the support of hardware resources. Currently, the hardware resources for AI computing can be GPUs (Graphics Processing Units). When implementing AI computing through GPUs, it is necessary to schedule GPU resources reasonably. If GPU resource scheduling is unreasonable, it will affect the efficiency and accuracy of AI computing. Summary of the invention

[0003] The purpose of the present invention is to provide a method and device for scheduling GPU resources in AI computing, which improves the efficiency and accuracy of AI computing through GPU resource scheduling.

[0004] In order to achieve the above-mentioned purpose, in a first aspect, the present disclosure provides a scheduling method for GPU resources in AI computing, comprising: in response to receiving an AI computing task, obtaining task execution information corresponding to multiple GPU devices respectively, the task execution information corresponding to each GPU device comprising: first information used to characterize the execution status of a real AI computing task of the GPU device and second information used to characterize the execution status of a virtual AI computing task of the GPU device; determining a target GPU device from the multiple GPU devices according to the first information corresponding to the AI ​​computing task and each GPU device respectively; determining a task execution strategy according to the second information corresponding to the AI ​​computing task and the target GPU device; and executing the AI ​​computing task through the target GPU device according to the task execution strategy.

[0005] Optionally, determining a target GPU device from the multiple GPU devices based on the first information respectively corresponding to the AI ​​computing task and each GPU device includes: determining resource requirement information based on the AI ​​computing task, the resource requirement information including required GPU quantity and required GPU performance; determining first information constraint conditions based on the resource requirement information, wherein if the required GPU quantity is one, the first information constraint conditions are constraint conditions for the first information corresponding to a single GPU device, and if the required GPU quantity is multiple, the first information constraint conditions include constraint conditions for the first information corresponding to a single GPU device and constraint conditions for the first information respectively corresponding to multiple GPU devices; determining a target GPU device that meets the first information constraint conditions from the multiple GPU devices based on the first information respectively corresponding to each GPU device.

[0006] Optionally, the constraint conditions for the first information corresponding to multiple GPU devices include: the similarity between the first information corresponding to the multiple GPU devices is greater than the target similarity, the target similarity is determined according to the required GPU performance and the required number of GPUs, and the target GPU device that satisfies the constraint conditions for the first information is determined from the multiple GPU devices according to the first information corresponding to each GPU device, including: in the case where the required number of GPUs is multiple, a first GPU device that satisfies the constraint conditions for the first information corresponding to a single GPU device is determined from the multiple GPU devices according to the first information corresponding to each GPU device; if the number of the first GPU devices is less than or equal to the required number of GPUs, the first GPU device is determined as the target GPU device; if the number of the first GPU devices is greater than the required number of GPUs, the information similarity between the first information corresponding to each first GPU device and the first information corresponding to other first GPU devices is determined; from the first GPU devices, a second GPU device whose information similarity is greater than the target similarity is determined; and the target GPU device is determined according to the second GPU device.

[0007] Optionally, determining the task execution strategy based on the second information corresponding to the AI ​​computing task and the target GPU device includes: when the number of the target GPU device is one, determining the task similarity between the AI ​​computing task and a virtual AI computing task executed by the target GPU device; determining the task execution priority of the AI ​​computing task in the target GPU device based on the task similarity and the second information; determining the task execution strategy includes the target GPU device executing the AI ​​computing task according to the task execution priority.

[0008] Optionally, determining the task execution strategy according to the second information corresponding to the AI ​​computing task and the target GPU device includes: when there are multiple target GPU devices, determining task coordination information according to the AI ​​computing task, the task coordination information being used to characterize that multiple target GPU devices execute the same computing task in parallel or that multiple target GPU devices execute different computing tasks in series; when the task coordination information characterizes that multiple target GPU devices execute the same computing task in parallel, determining the task similarity between the AI ​​computing task and the virtual AI computing tasks respectively executed by each target GPU device; determining the task execution priority corresponding to the AI ​​computing task in each target GPU device according to the task similarity and the second information corresponding to each target GPU device; determining the task execution strategy includes each target GPU device executing the AI ​​computing task according to the respectively corresponding task execution priority.

[0009] Optionally, the scheduling method also includes: when the task collaboration information represents that multiple target GPU devices execute different computing tasks in series, determining the task similarity between the AI ​​computing task and the virtual AI computing tasks executed by each target GPU device respectively; determining the AI ​​computing subtasks corresponding to each target GPU device respectively based on the task similarity and the AI ​​computing task; determining the task execution priority of each AI computing subtask in the corresponding target GPU device according to the second information corresponding to each target GPU device respectively; and determining the task execution strategy includes each target GPU device executing the corresponding AI computing subtask according to the corresponding task execution priority.

[0010] Optionally, the scheduling method also includes: obtaining device information corresponding to multiple GPU devices respectively; determining the virtual task execution amount and virtual task type allowed by the multiple GPU devices respectively according to the device information corresponding to the multiple GPU devices respectively; and assigning corresponding virtual AI computing tasks to each GPU device respectively according to the virtual task execution amount and the virtual task type, so that each GPU device respectively executes the corresponding virtual AI computing tasks.

[0011] Optionally, the scheduling method also includes: in response to detecting that each GPU device has a virtual AI computing task update request, updating the virtual AI computing task assigned to each GPU device according to the second information corresponding to each GPU device, so that each GPU device executes the corresponding updated virtual AI computing task.

[0012] Optionally, the scheduling method further includes: obtaining third information corresponding to the target GPU device for characterizing the execution status of the AI ​​computing task; determining the AI ​​computing tasks to be temporarily stored and / or the virtual AI computing tasks to be temporarily stored corresponding to the target GPU device according to the third information; and performing temporary storage processing on the AI ​​computing tasks to be temporarily stored and / or the virtual AI computing tasks to be temporarily stored through a pre-configured task register.

[0013] In a second aspect, the present disclosure provides a scheduling device for GPU resources in AI computing, including: an acquisition module, used to obtain task execution information corresponding to multiple GPU devices in response to receiving an AI computing task, the task execution information corresponding to each GPU device including: first information used to characterize the execution status of a real AI computing task of the GPU device and second information used to characterize the execution status of a virtual AI computing task of the GPU device; a determination module, used to determine a target GPU device from the multiple GPU devices based on the first information corresponding to the AI ​​computing task and each GPU device; determine a task execution strategy based on the second information corresponding to the AI ​​computing task and the target GPU device; an execution module, used to execute the AI ​​computing task through the target GPU device according to the task execution strategy.

[0014] Through the above technical solution, information characterizing the execution status of real AI computing tasks and information characterizing the execution status of virtual AI computing tasks are obtained. The information characterizing the execution status of real AI computing tasks is used in combination with the actual situation of AI computing tasks to determine the target GPU device from multiple GPU devices. Further, in combination with the actual situation of AI computing tasks and the execution status of virtual AI computing tasks of the target GPU device, a task execution strategy is determined so that the target GPU device can execute the AI ​​computing task. This technical solution configures real AI computing tasks and virtual AI computing tasks in GPU devices, which not only allows the execution status of virtual AI computing tasks to be used as reference information for scheduling GPU resources, but also ensures that GPU resources have a certain task execution margin, thereby achieving reasonable scheduling of GPU resources and improving the efficiency and accuracy of AI computing.

[0015] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the present disclosure but do not constitute a limitation of the present disclosure. In the accompanying drawings:

[0017] Figure 1 It is a schematic diagram showing an application scenario according to an exemplary embodiment.

[0018] Figure 2 The present invention is a flowchart of a method for scheduling GPU resources in AI computing according to an exemplary embodiment.

[0019] Figure 3 It is a schematic diagram of the first scheduling process according to an embodiment of the present disclosure.

[0020] Figure 4 It is a schematic diagram of a second scheduling process according to an embodiment of the present disclosure.

[0021] Figure 5 It is a block diagram of a scheduling device for GPU resources in AI computing according to an exemplary embodiment.

[0022] Figure 6 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0023] The specific implementation of the present disclosure is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the present disclosure, and is not used to limit the present disclosure.

[0024] With the development of AI technology, AI computing is increasingly being used. AI computing can improve computing accuracy and efficiency. AI computing requires the support of hardware resources. Currently, GPUs can be used as hardware resources for AI computing. When implementing AI computing through GPUs, GPU resources need to be reasonably scheduled. If GPU resource scheduling is unreasonable, the efficiency and accuracy of AI computing will be affected.

[0025] At present, the scheduling scheme for GPU resources in AI computing is generally based on the resource utilization, load rate, etc. of GPU resources. This scheduling scheme requires the scheduling rules to be configured in advance based on these reference information. If the rules are not set reasonably, it will lead to unreasonable resource scheduling, thereby affecting the efficiency and accuracy of AI computing.

[0026] Based on this, the embodiments of the present disclosure provide a technical solution to obtain information characterizing the execution status of a real AI computing task, as well as information characterizing the execution status of a virtual AI computing task. The information characterizing the execution status of the real AI computing task is combined with the actual situation of the AI ​​computing task to determine a target GPU device from multiple GPU devices. Furthermore, the task execution strategy is determined in combination with the actual situation of the AI ​​computing task and the execution status of the virtual AI computing task of the target GPU device so that the target GPU device can execute the AI ​​computing task.

[0027] This technical solution configures real AI computing tasks and virtual AI computing tasks in the GPU device, which not only allows the execution status of the virtual AI computing tasks to be used as reference information for scheduling GPU resources, but also ensures that the GPU resources have a certain task execution margin, thereby achieving reasonable scheduling of GPU resources and improving the efficiency and accuracy of AI computing.

[0028] The technical solution of the disclosed embodiment can be applied to various AI computing scenarios, such as: distributed application scenarios of AI models, distributed training scenarios of AI models. In different scenarios, the AI ​​computing tasks involved may be different. For example, in the distributed application scenarios of AI models, the AI ​​computing tasks are computing tasks involving model application; in the distributed training scenarios of AI models, the AI ​​computing tasks are computing tasks involving model training.

[0029] For the specific implementation of AI computing tasks, please refer to the introduction of subsequent embodiments.

[0030] Figure 1 is a schematic diagram showing an application scenario according to an exemplary embodiment. Figure 1 As shown, in this application scenario, multiple GPU devices and a scheduling platform are included. The performance of multiple GPU devices can be the same or different, and multiple GPU devices can independently perform AI computing tasks or collaborate with other GPU devices to perform AI computing tasks.

[0031] In some embodiments, the scheduling platform can allocate and issue AI computing tasks and schedule resources for multiple GPU devices to ensure the efficiency and accuracy of AI computing.

[0032] In some embodiments, the scheduling platform may be a CPU (Central Processing Unit), an MCU (Microcontroller Unit), etc., which are hardware resources that also have data processing capabilities.

[0033] In some embodiments, the scheduling platform and multiple GPU devices can be integrated into the same device, so that the device has AI computing capabilities.

[0034] Figure 2 is a flowchart of a method for scheduling GPU resources in AI computing according to an exemplary embodiment. The scheduling method can be applied to Figure 1 The scheduling platform shown, such as Figure 2 As shown, the scheduling method includes the following steps:

[0035] Step S21, in response to receiving an AI computing task, obtain task execution information corresponding to multiple GPU devices respectively, where the task execution information corresponding to each GPU device includes: first information used to characterize the execution status of a real AI computing task of the GPU device and second information used to characterize the execution status of a virtual AI computing task of the GPU device.

[0036] Step S22: Determine a target GPU device from multiple GPU devices according to the AI ​​computing task and the first information corresponding to each GPU device.

[0037] Step S23, determining the task execution strategy according to the second information corresponding to the AI ​​computing task and the target GPU device.

[0038] Step S24, executing the AI ​​computing task through the target GPU device according to the task execution strategy.

[0039] In step S21, the AI ​​computing task may be an AI computing task issued by an upper platform of the scheduling platform, for example, an AI computing task issued by a host computer.

[0040] In some embodiments, for multiple GPU devices in the resource pool, in addition to executing AI computing tasks, they can also maintain their own task execution status to generate first information and second information. Thus, the first information and the second information can be synchronized to the scheduling platform.

[0041] In some embodiments, virtual AI computing tasks are pre-configured in each GPU device, and each GPU device is configured to continuously or periodically execute the virtual AI computing tasks. The virtual AI computing tasks may be the same as the real AI computing tasks, but do not involve real task data, that is, the amount of data involved is small, so the size of the virtual AI computing tasks is small, and the GPU resources occupied are also small. These occupied resources can be used as backup resources for each GPU device.

[0042] Therefore, as an optional implementation method, the process of configuring virtual AI computing tasks in each GPU device includes: obtaining device information corresponding to multiple GPU devices respectively; determining the virtual task execution amount and virtual task type allowed by the multiple GPU devices respectively based on the device information corresponding to the multiple GPU devices respectively; and assigning corresponding virtual AI computing tasks to each GPU device respectively based on the virtual task execution amount and the virtual task type, so that each GPU device can execute the corresponding virtual AI computing tasks respectively.

[0043] In some embodiments, the device information may be performance information of the device, which may include: computing speed, computing accuracy, load capacity, etc.

[0044] Correspondingly, according to the performance information of the device, the allowable virtual task execution volume and the corresponding virtual task types that the GPU device can permit can be determined.

[0045] Among them, the virtual task execution volume can be used to allocate the corresponding AI computing virtual task volume to the GPU device, and the virtual task type can be used to allocate the AI computing virtual tasks of the corresponding type to the GPU device.

[0046] Regarding the virtual task types, for example: AI training tasks, AI application tasks, AI data preprocessing tasks, etc.

[0047] In some embodiments, the better the computing speed, computing accuracy, and load capacity, the more virtual task execution volume that can be permitted, and the higher the complexity and difficulty of the virtual task types that can be executed.

[0048] Therefore, different virtual task types and the corresponding device performance requirements for different virtual task types can be pre-configured. According to the pre-configured information and the actual situation of the GPU device, the corresponding virtual task execution volume and virtual task types can be determined.

[0049] Furthermore, based on the virtual task execution volume and virtual task types, the corresponding virtual AI computing tasks can be allocated to each GPU device respectively, so that each GPU device executes the corresponding virtual AI computing tasks respectively.

[0050] In some embodiments, each GPU device can allocate its own resources to execute virtual AI computing tasks or real AI computing tasks.

[0051] In some embodiments, the virtual AI computing tasks are not fixed and can be updated.

[0052] Therefore, as an optional implementation manner, the scheduling method further includes: in response to detecting that each GPU device has a virtual AI computing task update request, updating the virtual AI computing tasks allocated to each GPU device according to the second information corresponding to each GPU device, so that each GPU device executes the corresponding updated virtual AI computing tasks respectively.

[0053] In some embodiments, the virtual AI computing tasks allocated to each GPU device can be updated after waiting for the virtual AI computing tasks to be executed.

[0054] In some embodiments, the second information can be used to determine the update strategy of the virtual AI computing task. For example, if the second information includes a large amount of task execution progress, the virtual AI computing task can be fully updated. For example, in a training scenario, both the training data and the model algorithm can be updated. If the second information includes a small amount of task execution progress, only part of the data involved in the virtual AI computing task can be updated. For example, in a training scenario, the training data, test data, etc. are updated.

[0055] It can be understood that after virtual AI computing tasks are assigned to each GPU device, each GPU device will wait for the assignment of real AI computing tasks while executing the virtual AI computing tasks.

[0056] In step S21, the received AI computing task is an AI computing task to be assigned, and each GPU device may be currently executing an assigned AI computing task. Therefore, for the GPU device, the first information used to characterize the execution status of the real AI computing task and the second information used to characterize the execution status of the virtual AI computing task can be maintained.

[0057] The first information and the second information may include: task execution progress, task execution time, internal resource utilization, internal resource load and other information.

[0058] Further, in step S22, a target GPU device is determined from multiple GPU devices based on the AI ​​computing task and the first information corresponding to each GPU device.

[0059] It can be understood that the target GPU device is a GPU device that ultimately needs to execute the AI ​​computing task, and the target GPU device can be determined based on the AI ​​computing task and the first information.

[0060] In some embodiments, the number of target GPU devices may be one or more.

[0061] As an optional implementation, step S22 includes: determining resource requirement information according to the AI ​​computing task, the resource requirement information including the required number of GPUs and the required GPU performance; determining a first information constraint condition according to the resource requirement information, wherein if the required number of GPUs is one, the first information constraint condition is a constraint condition for the first information corresponding to a single GPU device, and if the required number of GPUs is multiple, the first information constraint condition includes a constraint condition for the first information corresponding to a single GPU device and a constraint condition for the first information corresponding to multiple GPU devices respectively; determining a target GPU device that meets the first information constraint condition from multiple GPU devices according to the first information corresponding to each GPU device.

[0062] In some embodiments, the AI ​​computing task can be analyzed to determine the task information, and then the resource requirement information can be determined based on the task information. For example, the required number of GPUs can be determined by analyzing the task data volume of the AI ​​computing task. The required GPU performance can be determined by analyzing the task complexity of the AI ​​computing task. The specific analysis and determination method can refer to the mature scheduling technology in this field. In the scheduling technology in this field, this information also needs to be applied, so the method of determining this information is relatively mature.

[0063] In some embodiments, the AI ​​computing task may directly carry resource requirement information, that is, the resource requirement information may be known information.

[0064] Furthermore, the first information constraint condition may be determined according to the required number of GPUs and the required GPU performance.

[0065] In some embodiments, when the number of GPUs required is one, corresponding constraints may be configured for the first information. For example, if the first information includes task execution progress, the constraints may include a task execution progress range, and if the first information includes task execution time, the constraints may include a task execution time range.

[0066] In some embodiments, when the number of GPUs required is one, the constraint condition has higher requirements. For example, the task execution progress range may be 90% to 100%, indicating that the task is about to be completed.

[0067] In some embodiments, when the number of GPUs required is multiple, the constraint conditions for the first information of a single GPU device may be included, which may be the same as the above embodiment, and the constraint conditions for the first information corresponding to multiple GPU devices may also be included.

[0068] As an example, the constraint conditions for the first information corresponding to multiple GPU devices include: the similarity between the first information corresponding to the multiple GPU devices is greater than the target similarity, and the target similarity is determined according to the required GPU performance and the required GPU quantity.

[0069] In this example, the similarity between the first information of each GPU device performing the AI ​​computing task needs to be greater than the target similarity. Through this implementation, the GPU device performing the AI ​​computing task can be a GPU device with similar execution conditions of the actual AI computing task, so as to avoid the situation where there is a particularly strong or particularly poor GPU performing the AI ​​computing task, so as to ensure the stability of task execution.

[0070] In some embodiments, the target similarity may be determined based on the required GPU performance and the required GPU quantity.

[0071] As an example, target similarity = (required number of GPUs / total number of GPUs)%-(required GPU performance-average GPU performance)×performance weight. The number of GPUs is the total number of GPU devices, and the average GPU performance is the average performance of each GPU device. The performance can be represented by a performance evaluation value obtained by weighted summing up various performance information. And the performance weight is used to convert the difference between the required GPU performance and the average GPU performance into a similarity impact value, and the performance weight can be pre-configured. For example, the greater the difference between the required GPU performance and the average GPU performance, the smaller the performance weight, the smaller the negative impact on the target similarity, and the greater the performance weight, the greater the negative impact on the target similarity. Therefore, if the difference between the required GPU performance and the average GPU performance is large, the higher the target similarity finally determined.

[0072] Furthermore, based on the determination of the first information constraint condition, the first information corresponding to each GPU device can be used to determine a target GPU device that meets the constraint condition.

[0073] In some embodiments, when the number of GPUs required is one, a GPU device that satisfies the constraint condition of the first information corresponding to a single GPU device can be determined from multiple GPU devices according to the first information corresponding to each GPU device. If there are multiple GPU devices determined, the GPU device with the best performance is determined as the target GPU device. If there is only one GPU device determined, the one GPU device is determined as the target GPU device.

[0074] In some embodiments, a target GPU device that satisfies the constraint conditions of the first information is determined from multiple GPU devices based on the first information corresponding to each GPU device, including: when a plurality of GPUs are required, a first GPU device that satisfies the constraint conditions of the first information corresponding to a single GPU device is determined from multiple GPU devices based on the first information corresponding to each GPU device; if the number of first GPU devices is less than the required number of GPUs, the first GPU device is determined as the target GPU device; if the number of first GPU devices is greater than or equal to the required number of GPUs, the information similarity between the first information corresponding to each first GPU device and the first information corresponding to other first GPU devices is determined; from the first GPU devices, a second GPU device whose information similarity is greater than the target similarity is determined; and the target GPU device is determined based on the second GPU device.

[0075] In this implementation, when the number of GPUs is multiple, a first GPU device that satisfies the constraint condition of the first information for a single GPU device is first determined. For example, the task execution progress is within the task execution progress range, and the task execution time is within the task execution time range.

[0076] Further, if the number of first GPU devices is less than or equal to the required number of GPUs, all first GPU devices may be determined as target GPU devices.

[0077] If the number of first GPU devices is greater than the required number of GPUs, further screening can be performed based on similarity. Specifically, the similarity of the first information of each first GPU device is determined to obtain a second GPU device with a similarity greater than the target similarity, and then the target GPU device is determined based on the second GPU device.

[0078] In some embodiments, the similarity of the first information can be determined based on specific information, for example, the similarity of the first information can be determined based on the difference between the task execution progress and the difference between the task execution time.

[0079] In some embodiments, for any first GPU device, the first information similarity between the device and each of the other first GPU devices may be determined respectively, and then the obtained multiple first information similarities are averaged to obtain the final information similarity.

[0080] Furthermore, if the number of second GPU devices is less than or equal to the required number of GPUs, all second GPU devices are directly determined as target GPU devices. If the number of second GPU devices is greater than the required number of GPUs, the second GPU devices with higher information similarity are determined as target GPU devices according to the information similarity ranking.

[0081] This implementation method can avoid excessive occupation of GPU resources and achieve reasonable occupation of GPU resources.

[0082] In step S23, a task execution strategy is further determined based on the second information corresponding to the AI ​​computing task and the target GPU device.

[0083] It can be understood that after the target GPU device is determined, only the GPU device used to execute the AI ​​computing task is determined. How to execute it specifically (i.e., the task execution strategy) needs to be further determined. Therefore, the second information can be used to determine the task execution strategy.

[0084] In some embodiments, the second information characterizing the execution status of the virtual AI computing task may include: task execution progress, task execution time, etc., which is similar to the first information.

[0085] In some embodiments, determining a task execution strategy based on the second information corresponding to the AI ​​computing task and the target GPU device may include: when the number of target GPU devices is one, determining the task similarity between the AI ​​computing task and the virtual AI computing task executed by the target GPU device; determining the task execution priority of the AI ​​computing task in the target GPU device based on the task similarity and the second information; and determining that the task execution strategy includes the target GPU device executing the AI ​​computing task according to the task execution priority.

[0086] In this implementation, if the number of target GPU devices is one, the task execution priority of the AI ​​computing task can be determined based on the task similarity between the AI ​​computing task and the virtual AI computing task executed by the target GPU device and the second information, so that the target GPU device executes the AI ​​computing task according to the task execution priority.

[0087] The task similarity can be determined by the difference between the task type and the task data volume. For example, the greater the difference in task type and the task data volume, the lower the similarity.

[0088] In some embodiments, when the task similarity is high (for example, higher than 90%), the execution priority is high, and the AI ​​computing task can replace the virtual AI computing task, that is, the virtual AI computing task is no longer executed.

[0089] In some embodiments, when the task similarity is low (for example, less than 90%), the execution priority is low, and the AI ​​computing task is executed after the currently executing task is completed, without replacing the virtual AI computing task, so as to maintain the synchronous execution of the virtual AI computing task.

[0090] In some embodiments, a task execution strategy is determined based on the AI ​​computing task and the second information corresponding to the target GPU device, including: when there are multiple target GPU devices, determining task coordination information based on the AI ​​computing task, the task coordination information being used to characterize that multiple target GPU devices execute the same computing task in parallel or that multiple target GPU devices execute different computing tasks in series; when the task coordination information characterizes that multiple target GPU devices execute the same computing task in parallel, determining the task similarity between the AI ​​computing task and the virtual AI computing tasks respectively executed by each target GPU device; determining the task execution priorities respectively corresponding to the AI ​​computing task in each target GPU device based on the task similarity and the second information respectively corresponding to each target GPU device; determining the task execution strategy includes each target GPU device executing the AI ​​computing task according to the respectively corresponding task execution priority.

[0091] In this implementation, when there are multiple target GPU devices, task coordination information can be determined first based on the AI ​​computing task.

[0092] For example, if the AI ​​computing task data volume is large, it can be determined that multiple target GPU devices execute different computing tasks in series. If the AI ​​computing task data volume is small, it can be determined that multiple target GPU devices execute the same computing task in parallel.

[0093] For example, if the AI ​​computing task is a distributed training task, multiple target GPU devices can be determined to execute different computing tasks in series. If the AI ​​computing task is an application task, multiple target GPU devices can execute the same computing task in parallel.

[0094] Furthermore, in the case of executing tasks in parallel, the task execution strategy may be determined for each target GPU device according to the aforementioned method of determining the task execution strategy for a single target GPU device.

[0095] In some embodiments, the scheduling method also includes: when the task collaboration information represents that multiple target GPU devices execute different computing tasks in series, determining the task similarity between the AI ​​computing task and the virtual AI computing tasks executed by each target GPU device respectively; determining the AI ​​computing subtasks corresponding to each target GPU device respectively based on the task similarity and the AI ​​computing task; determining the task execution priority of each AI computing subtask in the corresponding target GPU device according to the second information corresponding to each target GPU device respectively; and determining the task execution strategy includes each target GPU device executing the corresponding AI computing subtask according to the corresponding task execution priority.

[0096] In this implementation, if it is determined that multiple GPU devices execute tasks serially, the AI ​​computing task can be divided into AI computing subtasks corresponding to each target GPU device according to task similarity.

[0097] In some embodiments, the higher the task similarity, the larger the amount of data of the corresponding AI computing subtask and the more complex the task type.

[0098] Furthermore, after the AI ​​computing subtasks are determined, the corresponding task execution priorities are determined respectively according to the task execution priority determination method introduced in the aforementioned embodiment.

[0099] Through this implementation, AI computing tasks can be reasonably allocated to each target GPU device.

[0100] Figure 3 is a schematic diagram of a first scheduling process according to an embodiment of the present disclosure, such as Figure 3As shown in the figure, when GPU resource scheduling is performed, the GPU device used to perform AI computing tasks is first selected from multiple GPU devices based on the execution status of the real AI computing tasks. Then, the task execution strategies of these GPU devices are determined based on the execution status of the virtual AI computing tasks.

[0101] Figure 4 is a schematic diagram of a second scheduling process according to an embodiment of the present disclosure, such as Figure 4 As shown in the figure, when determining the task execution strategy, for the case where multiple GPU devices execute AI computing tasks, serial task execution and parallel task execution can be used. For the parallel task execution method, it is necessary to determine the task priority for each GPU device. For the serial task execution method, it is necessary to first divide the subtasks and then determine the task priority for each GPU device.

[0102] Furthermore, for the target GPU device, it can execute the corresponding tasks according to the AI ​​computing tasks assigned to it and the corresponding task execution priority, thereby realizing GPU resource scheduling.

[0103] In some embodiments, the scheduling method may also include: obtaining third information corresponding to the target GPU device for characterizing the execution status of the AI ​​computing task; determining the AI ​​computing tasks to be temporarily stored and / or the virtual AI computing tasks to be temporarily stored corresponding to the target GPU device based on the third information; and performing temporary storage processing on the AI ​​computing tasks to be temporarily stored and / or the virtual AI computing tasks to be temporarily stored through a pre-configured task register.

[0104] In this implementation, the AI ​​computing tasks to be temporarily stored and / or the virtual AI computing tasks to be temporarily stored may be determined based on the third information. The AI ​​computing tasks to be temporarily stored and / or the virtual AI computing tasks to be temporarily stored may be tasks with execution errors or tasks with lower task priorities.

[0105] In some embodiments, if the task execution progress does not change for a long time, it can be determined as a task with an execution error. If the task execution progress is 0 for a long time, it can be determined as a task with a low task priority.

[0106] Furthermore, the AI ​​computing tasks to be temporarily stored and / or the virtual AI computing tasks to be temporarily stored can be temporarily stored in a task register, which is a register shared by multiple GPU devices and can play a role in task temporary storage.

[0107] Furthermore, the temporary storage process may have a corresponding temporary storage duration. After the temporary storage duration is reached, the temporarily stored task needs to be executed or task error processing needs to be fed back.

[0108] Through this implementation, when an error occurs in task execution or the task progress is slow, task information can be temporarily stored in time to avoid task loss.

[0109] Correspondingly, the scheduling platform can periodically query the status of tasks temporarily stored in the task register and process these temporarily stored tasks in a timely manner.

[0110] In some embodiments, the GPU device may temporarily store tasks with the permission of the scheduling platform. For example, when there is a temporary storage requirement, a temporary storage request may be initiated to the scheduling platform. Tasks may be temporarily stored only after the temporary storage request is approved.

[0111] In some embodiments, the scheduling platform can determine whether to allow temporary storage based on the performance of the GPU device and the status of the virtual AI computing task. For example, if the performance is significantly better than other GPU devices, temporary storage is not allowed. If the virtual AI computing task is still being executed, temporary storage is not allowed, etc.

[0112] In some embodiments, in the case of an error task, the error cause (such as resource exhaustion, task data, etc.) needs to be fed back. If the error cause is confirmed to be a storable error cause, the GPU device can be allowed to storable the task. Storable error causes include, for example, resource exhaustion, task data error, etc.

[0113] Figure 5 is a block diagram of a scheduling device 500 for GPU resources in AI computing according to an exemplary embodiment, such as Figure 5 As shown, the device comprises:

[0114] The acquisition module 501 is used to obtain task execution information corresponding to multiple GPU devices in response to receiving an AI computing task, and the task execution information corresponding to each GPU device includes: first information used to characterize the execution status of a real AI computing task of the GPU device and second information used to characterize the execution status of a virtual AI computing task of the GPU device.

[0115] The determination module 502 is used to determine the target GPU device from the multiple GPU devices according to the first information corresponding to the AI ​​computing task and each GPU device respectively; and determine the task execution strategy according to the second information corresponding to the AI ​​computing task and the target GPU device.

[0116] The execution module 503 is used to execute the AI ​​computing task through the target GPU device according to the task execution strategy.

[0117] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0118] Figure 6 FIG. 6 is a block diagram of an electronic device 600 according to an exemplary embodiment. Figure 6 As shown, the electronic device 600 may include: a processor 601 , a memory 602 . The electronic device 600 may also include one or more of a multimedia component 603 , an input / output (I / O) interface 604 , and a communication component 605 .

[0119] The processor 601 is used to control the overall operation of the electronic device 600 to complete all or part of the steps in the scheduling method of GPU resources in the above-mentioned AI calculation. The memory 602 is used to store various types of data to support the operation of the electronic device 600, and these data may include, for example, instructions for any application or method used to operate on the electronic device 600, and application-related data, such as contact data, sent and received messages, pictures, audio, video, etc. The memory 602 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, referred to as EPROM), programmable read-only memory (Programmable Read-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic memory, flash memory, disk or optical disk. The multimedia component 603 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, which is used to receive external audio signals. The received audio signal may be further stored in the memory 602 or sent through the communication component 605. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 604 provides an interface between the processor 601 and other interface modules, and the above-mentioned other interface modules may be keyboards, mice, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 605 is used for wired or wireless communication between the electronic device 600 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more of them, so the corresponding communication component 605 may include: Wi-Fi module, Bluetooth module, NFC module.

[0120] In an exemplary embodiment, the electronic device 600 can be implemented by one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned method for scheduling GPU resources in AI computing.

[0121] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, and when the program instructions are executed by the processor, the steps of the above-mentioned method for scheduling GPU resources in AI computing are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 602 including program instructions, and the above-mentioned program instructions can be executed by the processor 601 of the electronic device 600 to complete the above-mentioned method for scheduling GPU resources in AI computing.

[0122] In another exemplary embodiment, a computer program product is also provided, which includes a computer program that can be executed by a processor, and when the computer program is executed by the processor, the steps of the above-mentioned method for scheduling GPU resources in AI computing are implemented.

[0123] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings; however, the present disclosure is not limited to the specific details in the above embodiments. Within the technical concept of the present disclosure, a variety of simple modifications can be made to the technical solution of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0124] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.

[0125] In addition, various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.

Claims

1. A method for scheduling GPU resources in AI computing, characterized in that: include: In response to receiving the AI ​​computing task, acquiring task execution information corresponding to each of the plurality of GPU devices, wherein the task execution information corresponding to each of the GPU devices includes: first information for characterizing the execution status of a real AI computing task of the GPU device and second information for characterizing the execution status of a virtual AI computing task of the GPU device; Determine a target GPU device from the multiple GPU devices according to the first information corresponding to the AI ​​computing task and each GPU device respectively; Determining a task execution strategy according to the second information corresponding to the AI ​​computing task and the target GPU device; Executing the AI ​​computing task according to the task execution strategy through the target GPU device; The step of determining a target GPU device from the plurality of GPU devices according to the first information corresponding to the AI ​​computing task and each GPU device respectively includes: Determine resource requirement information according to the AI ​​computing task, where the resource requirement information includes the required number of GPUs and the required GPU performance; Determine a first information constraint condition according to the resource requirement information, wherein if the required number of GPUs is one, the first information constraint condition is a constraint condition of the first information corresponding to a single GPU device, and if the required number of GPUs is multiple, the first information constraint condition includes a constraint condition of the first information corresponding to a single GPU device and a constraint condition of the first information corresponding to multiple GPU devices respectively; According to the first information respectively corresponding to each GPU device, determine a target GPU device that meets the constraint condition of the first information from the multiple GPU devices; The constraint conditions for the first information corresponding to the multiple GPU devices include: the similarity between the first information corresponding to the multiple GPU devices is greater than the target similarity, the target similarity is determined according to the required GPU performance and the required GPU quantity, and the target GPU device that satisfies the constraint conditions of the first information is determined from the multiple GPU devices according to the first information corresponding to each GPU device, including: In the case where the number of GPUs required is multiple, determining, according to the first information corresponding to each GPU device, a first GPU device that satisfies the constraint condition of the first information corresponding to a single GPU device from the multiple GPU devices; If the number of the first GPU devices is less than or equal to the required number of GPUs, determine the first GPU devices as the target GPU devices; If the number of the first GPU devices is greater than the required number of GPUs, determine the information similarity between the first information corresponding to each first GPU device and the first information corresponding to other first GPU devices; from the first GPU devices, determine the second GPU device whose information similarity is greater than the target similarity; and determine the target GPU device based on the second GPU device.

2. The scheduling method according to claim 1, characterized in that: The determining of the task execution strategy according to the second information corresponding to the AI ​​computing task and the target GPU device includes: When the number of the target GPU device is one, determining a task similarity between the AI ​​computing task and a virtual AI computing task executed by the target GPU device; Determining a task execution priority of the AI ​​computing task in the target GPU device according to the task similarity and the second information; Determining the task execution strategy includes the target GPU device executing the AI ​​computing task according to the task execution priority.

3. The scheduling method according to claim 1, characterized in that: The determining of the task execution strategy according to the second information corresponding to the AI ​​computing task and the target GPU device includes: In the case where there are multiple target GPU devices, determining task coordination information according to the AI ​​computing task, wherein the task coordination information is used to characterize that multiple target GPU devices execute the same computing task in parallel or that multiple target GPU devices execute different computing tasks in series; When the task coordination information represents that multiple target GPU devices execute the same computing task in parallel, determining the task similarity between the AI ​​computing task and the virtual AI computing tasks respectively executed by each target GPU device; Determine, according to the task similarity and the second information corresponding to each target GPU device, the task execution priority corresponding to each target GPU device of the AI ​​computing task; Determining the task execution strategy includes each target GPU device executing the AI ​​computing task according to the corresponding task execution priority.

4. The scheduling method according to claim 3, characterized in that: The scheduling method further includes: In a case where the task coordination information represents that a plurality of target GPU devices serially execute different computing tasks, determining a task similarity between the AI ​​computing task and a virtual AI computing task respectively executed by each target GPU device; Determine, according to the task similarity and the AI ​​computing task, the AI ​​computing subtasks corresponding to the target GPU devices; Determine, according to the second information corresponding to each target GPU device, the task execution priority of each AI computing subtask in the corresponding target GPU device; Determining the task execution strategy includes each target GPU device executing the corresponding AI computing subtask according to the corresponding task execution priority.

5. The scheduling method according to claim 1, characterized in that: The scheduling method further includes: Get the device information corresponding to multiple GPU devices; Determine, according to the device information respectively corresponding to the multiple GPU devices, the virtual task execution amounts and virtual task types respectively allowed by the multiple GPU devices; According to the virtual task execution amount and the virtual task type, corresponding virtual AI computing tasks are respectively allocated to each GPU device, so that each GPU device respectively executes the corresponding virtual AI computing tasks.

6. The scheduling method according to claim 5, characterized in that: The scheduling method further includes: In response to detecting that each GPU device has a virtual AI computing task update request, the virtual AI computing task assigned to each GPU device is updated according to the second information corresponding to each GPU device, so that each GPU device executes the corresponding updated virtual AI computing task.

7. The scheduling method according to claim 1, characterized in that: The scheduling method further includes: Obtaining third information corresponding to the target GPU device and used to characterize the execution status of the AI ​​computing task; Determine, according to the third information, the AI ​​computing tasks to be temporarily stored and / or the virtual AI computing tasks to be temporarily stored corresponding to the target GPU device; The AI ​​computing task to be temporarily stored and / or the virtual AI computing task to be temporarily stored are temporarily stored through a pre-configured task temporary register.

8. A scheduling device for GPU resources in AI computing, characterized in that: include: an acquisition module, configured to acquire task execution information corresponding to a plurality of GPU devices respectively in response to receiving an AI computing task, wherein the task execution information corresponding to each GPU device includes: first information for characterizing a real AI computing task execution status of the GPU device and second information for characterizing a virtual AI computing task execution status of the GPU device; A determination module, configured to determine a target GPU device from the plurality of GPU devices according to first information corresponding to the AI ​​computing task and each GPU device respectively; and determine a task execution strategy according to second information corresponding to the AI ​​computing task and the target GPU device; An execution module, configured to execute the AI ​​computing task through the target GPU device according to the task execution strategy; The step of determining a target GPU device from the plurality of GPU devices according to the first information corresponding to the AI ​​computing task and each GPU device respectively includes: Determine resource requirement information according to the AI ​​computing task, where the resource requirement information includes the required number of GPUs and the required GPU performance; Determine a first information constraint condition according to the resource requirement information, wherein if the required number of GPUs is one, the first information constraint condition is a constraint condition of the first information corresponding to a single GPU device, and if the required number of GPUs is multiple, the first information constraint condition includes a constraint condition of the first information corresponding to a single GPU device and a constraint condition of the first information corresponding to multiple GPU devices respectively; According to the first information respectively corresponding to each GPU device, determine a target GPU device that meets the constraint condition of the first information from the multiple GPU devices; The constraint conditions for the first information corresponding to the multiple GPU devices include: the similarity between the first information corresponding to the multiple GPU devices is greater than the target similarity, the target similarity is determined according to the required GPU performance and the required GPU quantity, and the target GPU device that satisfies the constraint conditions of the first information is determined from the multiple GPU devices according to the first information corresponding to each GPU device, including: In the case where the number of GPUs required is multiple, determining, according to the first information corresponding to each GPU device, a first GPU device that satisfies the constraint condition of the first information corresponding to a single GPU device from the multiple GPU devices; If the number of the first GPU devices is less than or equal to the required number of GPUs, determine the first GPU devices as the target GPU devices; If the number of the first GPU devices is greater than the required number of GPUs, determine the information similarity between the first information corresponding to each first GPU device and the first information corresponding to other first GPU devices; from the first GPU devices, determine the second GPU device whose information similarity is greater than the target similarity; and determine the target GPU device based on the second GPU device.

Citation Information

Patent Citations

  • Task scheduling method and device in heterogeneous cluster and electronic equipment

    CN110489223A

  • Resource scheduling method and device, electronic equipment and storage medium

    CN112346859A