Resource allocation method and device, computer equipment, storage medium and program product

By dynamically controlling the allocation of shared GPU resources by offline tasks, the problem of unbalanced demand for GPU resources by online inference tasks and offline training tasks is solved, improving resource utilization and avoiding waste.

CN120029774APending Publication Date: 2025-05-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510121434.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Online inference tasks and offline training tasks have uneven demands on GPU resources, resulting in uneven allocation of GPU resources and waste of resources.

Method used

Provide a resource allocation method, by obtaining resource allocation requests for shared GPU resources by offline tasks, determining offline usage information and expected usage information of offline tasks, and correlating it with online usage information of online tasks, performing resource planning, and dynamically regulating GPU resource allocation.

Benefits of technology

It effectively avoids competition between online and offline tasks for shared GPU resources, improves the utilization rate of GPU resources, and avoids resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029774A_ABST
    Figure CN120029774A_ABST
Patent Text Reader

Abstract

The invention relates to a resource allocation method and device, computer equipment, a storage medium and a program product. The method comprises the following steps: acquiring a resource allocation request of an off-line task for shared GPU resources, wherein the shared GPU resources are shared by the off-line task and an on-line task; off-line use information of the off-line task on the shared GPU resources is determined; obtaining expected use information of the shared GPU resource by the offline task, wherein the expected use information is related to online use information of the shared GPU resource by the online task; according to the offline use information and the expected use information, resource planning is carried out on the shared GPU resources, and to-be-allocated GPU resource information for the offline task is obtained; and allocating GPU resources to the offline task from the shared GPU resources according to the information of the GPU resources to be allocated. By adopting the method, the GPU resources distributed for the offline tasks can be dynamically regulated and controlled, and the utilization rate of the GPU resources is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a resource allocation method, apparatus, computer equipment, storage medium and computer program product. Background Art

[0002] With the promotion of large models, a large number of online reasoning tasks and offline training tasks based on large models have emerged, and the demand for GPU resources has become increasingly large, making it difficult to meet the computing needs of online reasoning tasks and offline training tasks. In the traditional mode, online reasoning tasks and offline training tasks are generally run on different GPU cards in different nodes to prevent mutual interference.

[0003] However, online reasoning tasks have high demand for GPU resources in some periods and low demand for GPU resources in other periods. If the same GPU resources are allocated to online reasoning tasks, there will be insufficient supply of GPU resources in the high demand period, while a large amount of GPU resources will remain in the low demand period, resulting in unbalanced allocation of GPU resources and waste of GPU resources. Summary of the invention

[0004] Based on this, it is necessary to provide a resource allocation method, device, computer equipment, computer-readable storage medium and computer program product that can dynamically control GPU resources in order to solve the above technical problems.

[0005] In one aspect, the present application provides a resource allocation method. The method comprises:

[0006] Obtaining a resource allocation request of an offline task for a shared GPU resource, where the shared GPU resource is shared by the offline task and the online task;

[0007] Determining offline usage information of the shared GPU resource by the offline task;

[0008] Acquire expected usage information of the shared GPU resource by the offline task, where the expected usage information is related to online usage information of the shared GPU resource by the online task;

[0009] Perform resource planning on the shared GPU resources according to the offline usage information and the expected usage information to obtain GPU resource information to be allocated for the offline task;

[0010] According to the information of the GPU resources to be allocated, GPU resources are allocated to the offline task from the shared GPU resources.

[0011] On the other hand, the present application also provides a resource allocation device. The device includes:

[0012] A request acquisition module, used to acquire a resource allocation request of an offline task for a shared GPU resource, where the shared GPU resource is shared by the offline task and the online task;

[0013] A determination module, used to determine offline usage information of the shared GPU resource by the offline task;

[0014] an expectation acquisition module, configured to acquire expected usage information of the shared GPU resource by the offline task, wherein the expected usage information is related to online usage information of the shared GPU resource by the online task;

[0015] A planning module, configured to perform resource planning on the shared GPU resources according to the offline usage information and the expected usage information, and obtain GPU resource information to be allocated for the offline task;

[0016] The allocation module is used to allocate GPU resources to the offline task from the shared GPU resources according to the information of the GPU resources to be allocated.

[0017] On the other hand, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0018] Obtaining a resource allocation request of an offline task for a shared GPU resource, where the shared GPU resource is shared by the offline task and the online task;

[0019] Determining offline usage information of the shared GPU resource by the offline task;

[0020] Acquire expected usage information of the shared GPU resource by the offline task, where the expected usage information is related to online usage information of the shared GPU resource by the online task;

[0021] Perform resource planning on the shared GPU resources according to the offline usage information and the expected usage information to obtain GPU resource information to be allocated for the offline task;

[0022] According to the information of the GPU resources to be allocated, GPU resources are allocated to the offline task from the shared GPU resources.

[0023] On the other hand, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0024] Obtaining a resource allocation request of an offline task for a shared GPU resource, where the shared GPU resource is shared by the offline task and the online task;

[0025] Determining offline usage information of the shared GPU resource by the offline task;

[0026] Acquire expected usage information of the shared GPU resource by the offline task, where the expected usage information is related to online usage information of the shared GPU resource by the online task;

[0027] Perform resource planning on the shared GPU resources according to the offline usage information and the expected usage information to obtain GPU resource information to be allocated for the offline task;

[0028] According to the information of the GPU resources to be allocated, GPU resources are allocated to the offline task from the shared GPU resources.

[0029] On the other hand, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0030] Obtaining a resource allocation request of an offline task for a shared GPU resource, where the shared GPU resource is shared by the offline task and the online task;

[0031] Determining offline usage information of the shared GPU resource by the offline task;

[0032] Acquire expected usage information of the shared GPU resource by the offline task, where the expected usage information is related to online usage information of the shared GPU resource by the online task;

[0033] Perform resource planning on the shared GPU resources according to the offline usage information and the expected usage information to obtain GPU resource information to be allocated for the offline task;

[0034] According to the information of the GPU resources to be allocated, GPU resources are allocated to the offline task from the shared GPU resources.

[0035] The resource allocation method, device, computer equipment, computer-readable storage medium and computer program product described above share GPU resources between offline tasks and online tasks. By obtaining the resource allocation request of the offline task for the shared GPU resources, the offline usage information of the offline task for the shared GPU resources is determined to obtain the historical usage information of the offline task for the shared GPU resources. The expected usage information of the offline task for the shared GPU resources is obtained to determine the GPU resource usage information planned for the offline task, and the planned usage information is related to the online usage information of the shared GPU resources by the online task, so that the use of the GUP resources by the online task is taken into account when planning the GPU resource usage information of the offline task, and the competition for the shared GUP resources when the online task and the offline task are executed can be avoided. According to the offline usage information and the expected usage information, the shared GPU resources are planned to estimate how many GUP resources need to be allocated to the offline task so that the usage information of the offline task for the shared GPU resources is consistent with the expected usage information. According to the GPU resource information to be allocated, the GPU resources are allocated to the offline task from the shared GPU resources, so that the use of the GPU resources by the offline task is controlled within the expected range, and the dynamic regulation of the GPU resources is realized. Moreover, when online tasks and offline tasks share GPU resources, the utilization of GPU resources can be improved by dynamically adjusting the allocation of GPU resources to offline tasks, which can effectively avoid the waste of GPU resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 An application environment diagram of a resource allocation method in an embodiment;

[0037] Figure 2 A schematic diagram of a flow chart of a resource allocation method in an embodiment;

[0038] Figure 3 A schematic diagram of related information of the first stage and the second stage in one embodiment;

[0039] Figure 4 A schematic diagram of a process of allocating resources for an offline task in a first stage and a second stage in one embodiment;

[0040] Figure 5 is a timing diagram of a resource allocation method in an embodiment;

[0041] Figure 6 An interactive schematic diagram of a resource allocation method in one embodiment;

[0042] Figure 7 It is an interactive schematic diagram of a resource allocation method in another embodiment;

[0043] Figure 8It is an interactive schematic diagram of a resource allocation method in another embodiment;

[0044] Fig. 9 It is an interactive schematic diagram of a resource allocation method in another embodiment;

[0045] Fig.10 A schematic diagram of the architecture of a resource scheduler in one embodiment;

[0046] Fig.11 is a timing diagram of a resource allocation method in another embodiment;

[0047] Fig.12 It is a diagram of the architecture usage of a resource allocation method in another embodiment;

[0048] Fig.13 is a structural block diagram of a resource allocation device in an embodiment;

[0049] Fig.14 is an internal structure diagram of a computer device in one embodiment;

[0050] Fig.15 FIG. 4 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0052] The resource allocation method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other servers. Both the terminal 102 and the server 104 can independently execute the resource allocation method provided in the embodiment of the present application. The terminal 102 and the server 104 can also be used in conjunction to execute the resource allocation method provided in the embodiment of the present application. When the terminal 102 and the server 104 are used in conjunction to execute the resource allocation method provided in the embodiment of the present application, the terminal 102 obtains the resource allocation request of the offline task for the shared GPU resource, and the shared GPU resource is shared by the offline task and the online task. The terminal 102 sends the resource allocation request to the server 104. The server 104 determines the offline usage information of the shared GPU resource by the offline task, obtains the expected usage information of the shared GPU resource by the offline task, and the expected usage information is related to the online usage information of the shared GPU resource by the online task. The server 104 performs resource planning for the shared GPU resource based on the offline usage information and the expected usage information, and obtains the GPU resource information to be allocated for the offline task. The server 104 allocates GPU resources to offline tasks from the shared GPU resources according to the information of the GPU resources to be allocated. The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers.

[0053] For ease of understanding, some terms involved in the following embodiments are explained illustratively.

[0054] GPU (Graphics Processing Unit): Graphics processor.

[0055] GPU colocation: used to deploy online reasoning tasks and offline training tasks in a co-located manner, so that online reasoning tasks and offline training tasks share GPU resources to improve GPU resource utilization. From the node dimension, colocation is to deploy multiple containers on the same node. The applications in these containers include both online applications and offline applications.

[0056] Kubernetes: K8S for short, is a system for running and coordinating containerized applications on computer devices, providing mechanisms for application deployment, planning, updating and maintenance. All resource allocation methods provided in this solution can be implemented on the K8S platform.

[0057] Pod: It is the basic unit of the Kubernetes system. It is the smallest component created or deployed by the user. It is a combination of one or more containers and a resource object for running containerized applications on the Kubernetes system.

[0058] CUDA (Compute Unified Device Architecture) is a parallel computing platform and programming model framework that provides developers with full support for efficient parallel computing using GPU resources.

[0059] In one embodiment, Figure 2 As shown, a resource allocation method is provided, which is applied to a computer device (computer device such as Figure 1 The terminal or server shown in FIG. 1 is used as an example to illustrate the method, which includes the following steps:

[0060] Step S202, obtaining a resource allocation request of an offline task for shared GPU resources, where the shared GPU resources are shared by the offline task and the online task.

[0061] Among them, online tasks and offline tasks refer to two different types of tasks performed in computer devices. Online tasks refer to tasks that need to process and respond to real-time data requests, and usually require a fast response. Online tasks include real-time recommendation systems, online advertising, chatbots, etc. Online tasks need to quickly process user requests and return results.

[0062] Offline tasks refer to tasks that process a large amount of data in batches without real-time data input. They usually have no time limit and can run in the background. Offline tasks include data mining, machine learning model training, log analysis, data cleaning, and other tasks.

[0063] In terms of demand for GPU resources, online tasks focus on processing real-time data streams, so sufficient GPU resources need to be provided in a timely manner to meet real-time processing and real-time response. Offline tasks, on the other hand, focus more on batch processing and analysis of data, and usually do not require real-time processing or real-time response. They can run in the background or be processed when time permits or GPU resources permit. Therefore, offline tasks have lower demand for GPU resources than online tasks.

[0064] GPU resources refer to the computing resources of the graphics processor. GPU resources may include GPU kernels. The GPU kernel refers to the processing elements (PEs) in the GPU, which are mainly responsible for performing various computing tasks.

[0065] ‌The number of GPU kernels‌ refers to the number of processing units in the GPU, which are responsible for executing the computing tasks in the GPU kernel. The number of GPU cores directly affects its parallel computing capabilities and processing speed. The more GPU cores, the faster the processing speed and the faster the response speed. Various computing tasks, such as the computing tasks in the offline tasks mentioned above, or the computing tasks in the online tasks, etc. ‌

[0066] Shared GPU resources refer to GPU resources shared by offline tasks and online tasks.

[0067] Specifically, offline tasks and online tasks are deployed on the same device, and the computer device is configured with shared GPU resources for the offline tasks and the online tasks.

[0068] In this embodiment, the offline task and the online task are deployed together on the computer device.

[0069] In one embodiment, the target application is deployed on a computer device, and the target application is used to process online tasks and offline tasks. When the target application needs to process offline tasks and needs to schedule GPU resources, the target application will initiate a resource allocation request for shared GPU resources from the offline task. When the target application needs to process online tasks and needs to schedule GPU resources, the target application will initiate a resource allocation request for shared GPU resources from the online task.

[0070] In one embodiment, the target application includes an offline application and an online application. The offline application is used to process offline tasks, and the online application is used to process online tasks. When the offline application needs to process offline tasks, the offline application initiates a resource allocation request.

[0071] In this embodiment, the online application may be a client installed in the terminal, or may refer to an application that does not require installation, that is, an application that can be used without downloading and installing. An application that does not require installation may also be called a mini-program or a sub-application, which is usually run as a sub-program in a client, and the client is called a parent application. An online application may also refer to a web application opened through a browser, etc., but is not limited thereto.

[0072] The offline application may be a client installed in the terminal.

[0073] Step S204: determining offline usage information of shared GPU resources by offline tasks.

[0074] The offline usage information refers to the historical usage information of the shared GPU resources by the offline tasks, and includes at least one of the offline resource usage amount and the offline resource utilization rate.

[0075] Offline resource usage refers to the historical usage of shared GPU resources by offline tasks. Offline resource utilization refers to the historical utilization of shared GPU resources by offline tasks.

[0076] Specifically, after obtaining a resource allocation request for an offline task, the computer device determines offline usage information of the offline task on the shared GPU resources.

[0077] In this embodiment, the computer device may determine offline usage information of the shared GPU resources by the offline task within a historical time period.

[0078] In one embodiment, the computer device may determine the GPU card to which the shared GPU resources belong. The array corresponding to the GPU card records the offline usage information of the shared GPU resources of the GPU card by the offline task. The computer device may obtain the offline usage information of the offline task from the array.

[0079] In one embodiment, after obtaining the resource allocation task, the computer device may detect the current stage of the online task. The current stage of the online task may be the first stage or the second stage. The demand for shared GPU resources by the online task in the first stage is greater than the demand for shared GPU resources in the second stage. That is, the usage of shared GPU resources by the online task in the first stage is greater than the usage of shared GPU resources in the second stage. For example, the first stage is the peak period of the online task, and the second stage is the trough period of the online task.

[0080] The offline usage information includes at least one of the following: first offline usage information of the offline task in the first stage, or second offline usage information of the offline task in the second stage.

[0081] When the current stage of the online task is the first stage, the computer device can obtain the first offline usage information of the first stage of the offline task in the historical time period. When the current stage of the online task is the second stage, the computer device can obtain the second offline usage information of the second stage of the offline task in the historical time period.

[0082] Step S206, obtaining expected usage information of the offline task on the shared GPU resources, where the expected usage information is related to the online usage information of the online task on the shared GPU resources.

[0083] The expected usage information refers to the usage information of the shared GPU resources planned for the offline task, that is, the expected usage information is the GPU resource usage information that is expected to be achieved in an ideal situation.

[0084] The expected usage information includes at least one of an expected resource usage or an expected resource utilization rate. The expected resource usage rate refers to the usage rate of the shared GPU resources planned for the offline task. The expected resource utilization rate refers to the utilization rate of the shared GPU resources planned for the offline task.

[0085] In one embodiment, the expected usage information can be set by the business party according to the online usage information of the online task, or can be obtained by the computer device through dynamic planning according to the online usage information of the online task. Therefore, the expected resource usage and the expected resource utilization can be set by the business party itself, or can be obtained by the computer device through dynamic planning.

[0086] For example, the expected resource utilization of offline tasks is set by the business side: the offline resource utilization of shared GPU resources by offline tasks is 20%, and the business side hopes that the utilization of shared GPU resources by offline tasks can reach 40%, so the expected resource utilization is set to 40%.

[0087] Specifically, the computer device may obtain expected usage information of the offline task on the shared GPU resources, where the expected usage information is related to online usage information of the online task. The online usage information refers to historical usage information of the online task on the shared GPU resources.

[0088] In this embodiment, the online usage information may include at least one of online resource usage or online resource utilization. Online resource usage refers to the historical usage of shared GPU resources by online tasks. Online resource utilization refers to the historical utilization of shared GPU resources by online tasks.

[0089] In one embodiment, the computer device obtains online usage information of the shared GPU resources by the online task, and determines expected usage information of the shared GPU resources by the offline task based on the offline usage information and the online usage information.

[0090] In one embodiment, the computer device may send a resource feedback request to the online application, and the online application responds to the resource feedback request and returns the online usage information of the online task for the shared GPU. The computer device performs planning based on the online usage information and the offline usage information to obtain the expected usage information of the offline task for the shared GPU resources.

[0091] Step S208: performing resource planning for shared GPU resources according to the offline usage information and the expected usage information, and obtaining GPU resource information to be allocated for offline tasks.

[0092] The GPU resource information to be allocated refers to the information about the GPU resources that are expected to be allocated to the offline task in order to make the usage information of the offline task on the shared GPU resources consistent with the expected usage information. That is, it is estimated how many GPU resources need to be allocated to the offline task in order to make the usage information of the offline task on the shared GPU resources consistent with the expected usage information.

[0093] In one embodiment, the information of GPU resources to be allocated includes a target number of GPU resources to be allocated to the offline task.

[0094] Specifically, the computer device performs resource planning on the shared GPU resources according to the offline usage information and the expected usage information, and obtains GPU resource information to be allocated for the offline task.

[0095] In this embodiment, the computer device can obtain the expected usage information of the online task, perform resource planning for the shared GPU resources based on the offline usage information and expected usage information of the offline task and the expected usage information of the online task, and obtain the GPU resource information to be allocated.

[0096] In one embodiment, the computer device can determine the idle GPU resources of the shared GPU resources according to the offline usage information and the online usage information. The computer device performs resource planning on the idle GPU resources according to the expected usage information of the offline task to obtain the GPU resource information to be allocated.

[0097] For example, if the offline resource utilization rate of shared GPU resources by offline tasks is 20% and the expected resource utilization rate is 40%, the computer device needs to calculate how much more GPU resources need to be allocated to offline tasks so that the offline task utilization rate of shared GPU resources reaches 40%.

[0098] Step S210 , allocating GPU resources to offline tasks from shared GPU resources according to the information of GPU resources to be allocated.

[0099] Specifically, the computer device allocates GPU resources to the offline task from the shared GPU resources according to the information of the GPU resources to be allocated.

[0100] In the traditional scheme, the GPU resources used by online tasks and offline tasks do not interfere with each other, but the requirements of online tasks for GPU resources are different in different time periods, which easily leads to waste of GPU resources. Therefore, in the resource allocation method of this embodiment, GPU resources shared by offline tasks and online tasks are set. By obtaining the resource allocation request of the offline task for the shared GPU resources, the offline usage information of the offline task for the shared GPU resources is determined to obtain the historical usage information of the offline task for the shared GPU resources. The expected usage information of the offline task for the shared GPU resources is obtained to determine the GPU resource usage information planned for the offline task, and the planned usage information is related to the online usage information of the shared GPU resources by the online task, so that the use of the GUP resources by the online task is taken into account when planning the GPU resource usage information of the offline task, and the competition for the shared GUP resources when the online task and the offline task are executed can be avoided. According to the offline usage information and the expected usage information, the shared GPU resources are resource planned to estimate how many GUP resources need to be allocated to the offline task, so that the usage information of the offline task for the shared GPU resources is consistent with the expected usage information. According to the information of GPU resources to be allocated, GPU resources are allocated to offline tasks from shared GPU resources, so that the use of GPU resources by offline tasks is controlled within the expected range, and dynamic regulation of GPU resources is achieved. In addition, when online tasks and offline tasks share GPU resources, the allocation of GPU resources to offline tasks can be dynamically regulated to improve the utilization of GPU resources and effectively avoid the waste of GPU resources.

[0101] In one embodiment, the offline usage information includes offline resource usage and offline resource utilization, the expected usage information includes expected resource utilization, and the to-be-allocated GPU resource information includes a target number of to-be-allocated GPU resources;

[0102] According to the offline usage information and the expected usage information, resource planning is performed on the shared GPU resources to obtain the GPU resource information to be allocated for the offline task, including: according to the offline resource usage, the offline resource utilization rate and the expected resource utilization rate, resource planning is performed on the shared GPU resources to obtain the target number of GPU resources to be allocated when the resource utilization rate of the offline task reaches the expected resource utilization rate;

[0103] According to the information of the GPU resources to be allocated, the GPU resources are allocated to the offline task from the shared GPU resources, including: allocating a target number of GPU resources to the offline task from the shared GPU resources.

[0104] Specifically, the computer device obtains the offline resource usage and offline resource utilization of the shared GPU resources by the offline task, and obtains the expected resource utilization of the shared GPU resources by the offline task. The computer device calculates how many GPU resources need to be allocated to the offline task based on the offline resource usage, offline resource utilization and expected resource utilization, so that the resource utilization of the shared GPU resources by the offline task reaches the expected resource utilization. The calculated number of GPU resources that need to be allocated to the offline task is the target number.

[0105] The computer device allocates a target number of GPU resources to the offline task from the shared GPU resources.

[0106] In one embodiment, the computer device may determine the total number of shared GPU resources, calculate the difference between the expected resource utilization and the offline resource utilization, calculate the product of the difference and the total number, and the difference between the product and the offline resource usage is the target number.

[0107] For example, the target quantity may be calculated using the following formula: ratio of expected resource utilization - offline resource utilization = (offline resource usage + target quantity) / total quantity.

[0108] In other embodiments, the difference between the expected resource utilization and the offline resource utilization may be calculated, and the product of the difference and the total quantity may be calculated, and the product is the target quantity.

[0109] In one embodiment, the offline usage information includes offline resource usage. According to the offline resource usage and the expected resource utilization, resource planning is performed on the shared GPU resources to obtain the target number of GPU resources to be allocated when the resource utilization of the offline task reaches the expected resource utilization.

[0110] Specifically, the computer device can calculate the product of the total number of shared GPU resources and the expected resource utilization, and the difference between the product and the offline resource usage is the target number. The product represents the total usage of shared GPU resources by offline tasks when the resource utilization of offline tasks reaches the expected resource utilization.

[0111] For example, the target quantity can be calculated using the following formula: expected resource utilization = (offline resource usage + target quantity) / total quantity ratio.

[0112] In other embodiments, resource planning is performed on shared GPU resources based on offline resource usage and expected resource utilization to obtain a target number of GPU resources to be allocated when the resource utilization of offline tasks reaches the expected resource utilization, including: resource planning is performed on shared GPU resources based on periodic offline resource usage of offline tasks in a historical period and periodic expected resource utilization of offline tasks in a planned period to obtain a periodic number of GPU resources to be allocated when the resource utilization of offline tasks in the planned period reaches the periodic expected resource utilization.

[0113] In one embodiment, the shared GPU resource includes a shared GPU core, the offline resource usage includes the number of offline cores allocated to the offline task, the offline resource utilization includes the offline core utilization of the offline task, and the expected resource utilization includes the expected core utilization for the offline task. The computer device calculates how many GPU cores need to be allocated to the offline task based on the number of offline cores, the offline core utilization, and the offline core utilization so that the utilization of the shared GPU core by the offline task reaches the offline core utilization. The calculated number of GPU cores is the target number, and the computer device allocates the target number of GPU cores to the offline task.

[0114] In this embodiment, resource planning is performed on shared GPU resources based on offline resource usage, offline resource utilization and expected resource utilization, so that the target number of GPU resources required for offline tasks to achieve the expected resource utilization can be accurately calculated based on the historical usage, historical utilization and expected utilization of shared GPU resources by offline resources, thereby controlling the GPU resources allocated to offline tasks within expectations and realizing dynamic regulation of GPU resources for offline tasks.

[0115] In one embodiment, the offline resource usage includes the periodic offline resource usage of the offline task in the historical period, the offline resource utilization includes the periodic offline resource utilization of the offline task in the historical period, the expected resource utilization includes the periodic expected resource utilization of the offline task in the planned period, and the target number of GPU resources to be allocated includes the periodic number of GPU resources to be allocated when the resource utilization of the offline task in the planned period reaches the expected resource utilization;

[0116] Allocating a target number of GPU resources to the offline task from the shared GPU resources includes: allocating a period number of GPU resources to the offline task from the shared GPU resources within the planning period.

[0117] The historical period refers to the time period for allocating GPU resources to offline tasks in the past, which can be in units of days, weeks, months, hours, minutes, seconds, milliseconds, etc. For example, the past day is regarded as a historical period. The historical period can be divided into multiple time periods, such as 8:00-9:00, 9:00-12:00, etc. in the past day. Different time periods can represent different stages, for example, 8:00-9:00 is the peak period, 12:00-14:00 is the trough period, etc.

[0118] The planning cycle is the planning allocation cycle, which refers to the time period for allocating GPU resources to offline tasks.

[0119] Specifically, the computer device obtains the periodic offline resource usage and periodic offline resource utilization of the shared GPU resources by the offline task in the historical period, and obtains the periodic expected resource utilization of the shared GPU resources by the offline task in the planned period. The computer device calculates the number of GPU resources that need to be allocated when the resource utilization of the offline task in the planned period reaches the expected resource utilization, that is, the number of cycles, based on the periodic offline resource usage, the periodic offline resource utilization, and the periodic expected resource utilization.

[0120] The computer device allocates GPU resources of the number of cycles to the offline task from the shared GPU resources within the planning cycle.

[0121] In one embodiment, the periodic offline resource usage includes the stage offline resource usage, the periodic offline resource utilization includes the stage offline resource utilization, and the periodic expected resource utilization includes the stage expected resource utilization.

[0122] The computer device can obtain the offline resource usage and offline resource utilization of the offline task at the same stage of multiple historical cycles, and obtain the expected resource utilization of the offline task at the same stage of the planned cycle. The computer device calculates the number of GPU resources that need to be allocated when the resource utilization of the shared GPU resources of the offline task reaches the expected resource utilization in the same stage of the planned cycle, i.e., the number of stages, based on the offline resource usage, offline resource utilization, and expected resource utilization.

[0123] In this embodiment, the stage offline resource usage may include at least one of the first stage offline resource usage or the second stage offline resource usage. The stage offline resource utilization rate includes at least one of the first stage offline resource utilization rate or the second stage offline resource utilization rate, and the stage expected resource utilization rate includes at least one of the first stage expected resource utilization rate or the second stage expected resource utilization rate. For example, the first stage is a peak period, and the second stage is a valley period.

[0124] In this embodiment, the calculation method of the cycle number and the stage number can refer to the calculation method of the target number in the above embodiments.

[0125] In this embodiment, based on the usage and utilization rate of shared GPU resources by offline resources in historical cycles and the utilization rate planned to be achieved in the next cycle, the estimated number of GPU resources to be allocated when the offline tasks reach the planned resource utilization rate in the next cycle is accurately calculated, thereby periodically controlling the GPU resources allocated to offline tasks and realizing dynamic regulation of the GPU resources of offline tasks.

[0126] In one embodiment, the shared GPU resources include shared GPU cores; periodic offline resource usage, including the number of offline cores allocated to offline tasks in a historical period; periodic offline resource utilization, including offline core utilization for offline tasks in a historical period;

[0127] The cycle expected resource utilization includes the expected core utilization for offline tasks in the planned cycle; the cycle number of GPU resources to be allocated includes the cycle number of GPU cores to be allocated when the GPU core utilization of offline tasks in the planned cycle reaches the expected core utilization.

[0128] Specifically, the computer device obtains the number of offline cores allocated to the offline task in the historical period and the offline core utilization of the offline task on the shared GPU core. The computer device obtains the offline core utilization for the offline task in the historical period, and calculates how many GPU cores need to be allocated to the offline task in the planned period according to the number of offline cores in the historical period, the offline core utilization and the expected core utilization in the planned period, so that the GPU core utilization of the offline task reaches the expected core utilization, which is the number of cycles.

[0129] In this embodiment, the number of GPU cores allocated in the next period is predicted by periodically calculating the number and utilization of GPU cores allocated, as well as the expected core utilization. This allows for periodic regulation of the allocation of GPU cores to offline tasks, making the allocation of GPU cores to offline tasks controllable.

[0130] In one embodiment, within a planning period, a period number of GPU resources are allocated to offline tasks from shared GPU resources, including:

[0131] According to the planned cycle and the number of GPU cores to be allocated in the cycle, the core allocation rate in the planned cycle is determined; in the planned cycle, GPU cores are allocated to offline tasks multiple times according to the core allocation rate until the number of allocated GPU cores reaches the number of cycles.

[0132] Specifically, the computer device can determine the duration of the planning cycle, and calculate the core allocation rate per unit time according to the duration and the number of GPU cores to be allocated during the cycle. When the planning cycle is reached, the computer device allocates GPU cores to the offline task multiple times within the planning cycle according to the core allocation rate until the number of allocated GPU cores reaches the number of cycles.

[0133] In this embodiment, the core allocation rate within the planned cycle is determined according to the planned cycle and the number of GPU cores to be allocated in the cycle, so that within the planned cycle, GPU cores are allocated to offline tasks multiple times according to the core allocation rate until the number of allocated GPU cores reaches the number of cycles, thereby controlling the rate at which GPU cores are issued within the cycle and making better use of the GPU cores.

[0134] In one embodiment, obtaining expected usage information of shared GPU resources by offline tasks includes:

[0135] Obtain online usage information of shared GPU resources by online tasks; determine idle GPU resources of shared GPU resources based on offline usage information and online usage information; plan expected usage information of shared GPU resources by offline tasks based on offline usage information, online usage information and idle GPU resources.

[0136] Specifically, the computer device obtains online usage information of shared GPU resources by online tasks, determines idle GPU resources of shared GPU resources according to offline usage information and online usage information, and plans expected usage information of shared GPU resources by offline tasks according to offline usage information, online usage information and idle GPU resources.

[0137] In one of the embodiments, the method further includes: configuring expected usage information for shared GPU resources for online tasks based on offline usage information, online usage information and idle GPU resources.

[0138] In one embodiment, the computer device plans the expected usage information of the offline task for the shared GPU resources based on the offline usage information, the online usage information, the idle GPU resources of the offline task, and the expected usage information of the online task.

[0139] In this embodiment, historical usage information of shared GPU resources by online tasks is obtained, and idle GPU resources of shared GPU resources are calculated based on historical usage information of offline tasks and historical usage information of online tasks. Based on offline usage information, online usage information and idle GPU resources, usage information of shared GPU resources can be accurately planned for offline tasks, so as to control the usage of GPU resources by offline tasks within an expected range and avoid insufficient supply of GPU resources.

[0140] like Figure 3 As shown, in one embodiment, the online usage information includes at least one of first online usage information of the online task in the first stage or second online usage information of the online task in the second stage, and the usage of the shared GPU resources by the online task in the first stage is greater than the usage of the shared GPU resources in the second stage;

[0141] The offline usage information includes at least one of the first offline usage information of the offline task in the first stage or the second offline usage information in the second stage, the idle GPU resources include at least one of the first idle GPU resources in the first stage or the second idle GPU resources in the second stage, the expected usage information includes at least one of the first expected usage information in the first stage or the second expected usage information in the second stage. The GPU resource information to be allocated includes at least one of the first GPU resource information to be allocated in the first stage or the second GPU resource information to be allocated in the second stage.

[0142] The first online usage information refers to the usage information of the shared GPU resources by the online tasks in the first phase, and the second online usage information refers to the usage information of the shared GPU resources by the online tasks in the second phase.

[0143] The first offline usage information refers to the usage information of the shared GPU resources by the offline task in the first phase. The second offline usage information refers to the usage information of the shared GPU resources by the offline task in the second phase.

[0144] The first expected usage information refers to the usage information of the shared GPU resources achieved in the first phase of offline task planning. The second expected usage information refers to the usage information of the shared GPU resources achieved in the second phase of offline task planning.

[0145] like Figure 4 As shown, a resource allocation request of an offline task for shared GPU resources is obtained to determine the current stage of the online task. When in the first stage, first offline usage information and first expected usage information of the offline task for shared GPU resources in the first stage are obtained. According to the first offline usage information and the first expected usage information, resource planning is performed on the shared GPU resources to obtain first GPU resource information to be allocated for the offline task. According to the first GPU resource information to be allocated, GPU resources are allocated to the offline task from the shared GPU resources.

[0146] When in the second stage, second offline usage information and second expected usage information of the shared GPU resources by the offline task in the second stage are obtained. According to the second offline usage information and the second expected usage information, resource planning is performed on the shared GPU resources to obtain second GPU resource information to be allocated for the offline task. According to the second GPU resource information to be allocated, GPU resources are allocated to the offline task from the shared GPU resources.

[0147] In this embodiment, the first stage is the peak period of the demand for shared GPU resources by online tasks, and the second stage is the trough period of the demand for shared GPU resources by online tasks. According to the usage of GPU resources by online tasks during the peak period and the trough period, the usage of GPU resources by offline tasks during the peak period and the trough period is planned, so as to avoid the competition for GPU resources between online tasks and offline tasks during the peak period and the trough period. In addition, by planning the GPU resources allocated to offline tasks, it is possible to dynamically adjust the idle GPU resources of online tasks to offline tasks on the basis of meeting the demand for GPU resources by online tasks, thereby improving the utilization rate of GPU resources.

[0148] In one embodiment, a resource allocation method is provided, which is applied to a computer device, comprising:

[0149] Get the kernel allocation request of the offline task for the shared GPU kernel. The shared GPU kernel is shared by the offline task and the online task.

[0150] The current stage of the online task is determined. The current stage may be the first stage or the second stage. The demand of the online task for the shared GPU core in the first stage is greater than the demand for the shared GPU core in the second stage.

[0151] When the online task is in the first stage, the first offline kernel quantity and the first offline kernel utilization rate of the offline task for the shared GPU kernel in the first stage are obtained, and the first online kernel quantity and the first online kernel utilization rate of the online task for the shared GPU kernel in the first stage are obtained.

[0152] A first idle GPU core number of the shared GPU core is determined according to the first offline core number and the first online core number.

[0153] A first expected core utilization of the shared GPU cores for the offline task in the first phase is planned according to the first number of offline cores, the first offline core utilization, the first number of online cores, the first online core utilization and the first number of idle GPU cores.

[0154] According to the first number of offline cores, the first offline core utilization and the first expected core utilization, a first stage number of GPU cores to be allocated when the GPU core utilization of the offline task reaches the first expected core utilization is calculated.

[0155] According to the duration of the first stage and the number of GPU cores to be allocated in the first stage, the first core allocation rate in the first stage is determined. During the first stage, GPU cores are allocated to offline tasks multiple times according to the first core allocation rate until the number of allocated GPU cores reaches the number in the first stage.

[0156] Similarly, when the online task is in the second stage, the second offline core quantity and the second offline core utilization of the shared GPU core by the offline task in the second stage are obtained, and the second online core quantity and the second online core utilization of the shared GPU core by the online task in the second stage are obtained.

[0157] According to the second number of offline cores and the second number of online cores, a second number of idle GPU cores of the shared GPU core is determined.

[0158] A second expected core utilization of the shared GPU core for the offline task in the second phase is planned according to the second number of offline cores, the second offline core utilization, the second number of online cores, the second online core utilization and the second number of idle GPU cores.

[0159] The second stage number of GPU cores to be allocated when the GPU core utilization of the offline task reaches the second expected core utilization is calculated according to the second offline core number, the second offline core utilization and the second expected core utilization.

[0160] According to the duration of the second stage and the number of GPU cores to be allocated in the second stage, the second core allocation rate in the second stage is determined. During the second stage, GPU cores are allocated to offline tasks multiple times according to the second core allocation rate until the number of allocated GPU cores reaches the number in the second stage.

[0161] In this embodiment, the first stage is the peak period of the online task's demand for the shared GPU kernel, and the second stage is the trough period of the online task's demand for the shared GPU kernel. According to the historical usage and utilization of the GPU resources by the online task at the peak and trough periods, and the historical usage and utilization of the shared GPU kernel by the offline task, the utilization of the GPU kernel resources by the offline task at the peak and trough periods is planned, so as to accurately calculate the number of GPU kernels expected to be allocated when the offline task reaches the planned GPU kernel utilization, so as to control the GPU kernels allocated to the offline task at different stages, and realize the dynamic regulation of the GPU kernel resources of the offline task. In this way, not only can the competition for GPU kernel resources between the online task and the offline task at the peak and trough periods be avoided. In addition, the number of GPU kernels allocated to the offline task can be planned, and the idle GPU kernels of the online task can be dynamically regulated to the offline task on the basis of meeting the demand of the online task for the GPU kernel, so as to improve the utilization of the GPU kernel resources.

[0162] In one embodiment, obtaining a resource allocation request of an offline task for a shared GPU resource includes:

[0163] The resource allocation request sent to the resource scheduler is intercepted by the interception tool. The resource allocation request is used to request to allocate GPU resources for offline tasks from shared GPU resources.

[0164] According to the information of GPU resources to be allocated, GPU resources are allocated to offline tasks from shared GPU resources, including:

[0165] The resource scheduler feeds back the information of GPU resources to be allocated to the interception tool; the interception tool allocates GPU resources to offline tasks from the shared GPU resources according to the information of GPU resources to be allocated.

[0166] Specifically, a resource scheduler and an interception tool are deployed on the computer device. The resource scheduler is responsible for managing the usage information of the shared GPU resources on the computer device, such as the usage information of the shared GPU resources by offline tasks and the usage information of the GPU resources by online tasks.

[0167] The interception tool is used to intercept resource allocation requests from offline tasks, allocate GPU resources to offline tasks according to GPU resource allocation information provided by the resource scheduler, and feed back resource allocation results to the resource scheduler.

[0168] In this embodiment, Figure 5 As shown, offline applications and online applications are deployed on the computer device, the offline applications are used to execute offline tasks, and the online applications are used to execute online tasks. The shared GPU resources are shared by the offline applications and the online applications.

[0169] The offline application sends a resource allocation request for the offline task to the resource scheduler. The interception tool detects the request sent by the offline application in real time or periodically. When the resource allocation request sent by the offline application is detected, the resource allocation request is intercepted, and a resource planning request is sent to the resource scheduler based on the resource allocation request. In response to the resource planning request, the resource scheduler determines the offline usage information and expected usage information of the offline task for the shared GPU resources, and determines the GPU resource information to be allocated for the offline task based on the offline usage information and expected usage information. The resource scheduler feeds back the GPU resource information to be allocated to the interception tool. The interception tool allocates GPU resources to the offline application from the shared GPU resources according to the GPU resource information to be allocated. The offline application uses the allocated GPU resources to execute the offline task.

[0170] Furthermore, if Figure 6 As shown, the resource scheduler obtains online usage information of online tasks on shared GPU resources from online applications in response to a resource planning request. The resource scheduler plans the expected usage information of offline tasks for shared GPU resources based on the offline usage information and the online usage information.

[0171] In one embodiment, Figure 7 As shown, the resource scheduler may include an offline resource scheduler and an online resource scheduler. The offline resource scheduler is used for offline application usage information for shared GPU resources, and the online resource scheduler is used to manage online application usage information for shared GPU resources. The offline resource scheduler and the online resource scheduler can communicate with each other, and the online resource scheduler obtains the online usage information of the online task from the online application and feeds it back to the offline resource scheduler. The offline resource scheduler obtains the offline usage information of the offline task from the offline application, and plans the GPU resource information to be allocated in combination with the online usage information and the expected usage information. The offline resource scheduler feeds back the GPU resource information to be allocated to the interception tool, and the interception tool allocates GPU resources to the offline application according to the GPU resource information to be allocated.

[0172] In this embodiment, the resource allocation request sent to the resource scheduler is intercepted by an interception tool, so as to realize the request timing of the resource allocation request for the offline task through the interception tool. The GPU resource information that can be allocated to the offline task this time is planned by the resource scheduler to realize the dynamic regulation of the GPU resources of the offline task. The information of the GPU resources to be allocated is fed back to the interception tool, and the GPU resources are allocated to the offline task from the shared GPU resources according to the information of the GPU resources to be allocated through the interception tool, so that the timing of the GPU resources being issued can be controlled through the interception tool.

[0173] In one embodiment, the interception tool includes a statistical unit and a kernel processing unit; determining offline usage information of shared GPU resources by offline tasks includes:

[0174] Obtain offline usage information of shared GPU resources by offline tasks through the statistical unit, and report the offline usage information to the resource scheduler;

[0175] Through the interception tool, according to the information of GPU resources to be allocated, GPU resources are allocated to offline tasks from shared GPU resources, including:

[0176] The kernel processing unit receives the GPU resource information to be allocated fed back by the resource scheduler; and the kernel processing unit allocates the GPU resources to the offline task from the shared GPU resources according to the GPU resource information to be allocated.

[0177] In this embodiment, the interception tool includes a statistical unit and a kernel processing unit. Through the statistical unit, offline usage information of shared GPU resources by offline tasks is obtained, and the offline usage information is reported to the resource scheduler to provide the resource scheduler with the latest GPU resource allocation information in a timely manner. Through the kernel processing unit, the to-be-allocated GPU resource information fed back by the resource scheduler is received, and through the kernel processing unit, according to the to-be-allocated GPU resource information, GPU resources are allocated to offline tasks from shared GPU resources, so that dynamic regulation of GPU resources for offline tasks in GPU resource sharing scenarios can be achieved through interaction between the kernel processing unit and the resource scheduler.

[0178] In one embodiment, the method further comprises:

[0179] After allocating GPU resources to offline tasks, the statistical unit feeds back the resource allocation result to the resource scheduler; and the resource scheduler updates the offline usage information of the offline tasks on the shared GPU resources based on the resource allocation result.

[0180] like Figure 8 As shown, the core processing unit intercepts the request sent by the offline application in real time or periodically. When the resource allocation request sent by the offline application is detected, the resource allocation request is intercepted, and a resource planning request is sent to the resource scheduler based on the resource allocation request. In response to the resource planning request, the resource scheduler determines the offline usage information and expected usage information of the offline task for the shared GPU resources, and determines the GPU resource information to be allocated for the offline task based on the offline usage information and the expected usage information. The resource scheduler feeds back the GPU resource information to be allocated to the core processing unit. The core processing unit allocates GPU resources to the offline application from the shared GPU resources according to the GPU resource information to be allocated. The offline application uses the allocated GPU resources to execute the offline task.

[0181] After allocating GPU resources to offline tasks, the core processing unit feeds back the resource allocation result to the statistical unit. The statistical unit feeds back the resource allocation result to the resource scheduler. The resource scheduler updates the offline usage information of the offline tasks for the shared GPU resources based on the resource allocation result.

[0182] In one embodiment, Fig. 9 As shown, the resource scheduler includes a communication unit and a planning unit. The communication unit communicates with the kernel processing unit and the statistical unit respectively, and the communication unit is used to receive the resource planning request sent by the kernel processing unit, feedback the GPU resource information to be allocated to the kernel processing unit, and receive the resource allocation result feedback from the statistical unit. The planning unit is used to determine the GPU resource information to be allocated for the offline task according to the offline usage information and the expected usage information. The planning unit is also used to update the offline usage information of the offline task for the shared GPU resources based on the resource allocation result.

[0183] like Fig.10 As shown, each GUP card corresponds to an array, which records the usage information of the GUP resources in the GPU card by offline tasks. The planning unit can obtain the offline usage information from the array and update the offline usage information in the array.

[0184] In this embodiment, after allocating GPU resources to offline tasks, the statistical unit feeds back the resource allocation results to the resource scheduler. The resource scheduler timely updates the offline usage information of the offline tasks for the shared GPU resources based on the resource allocation results, thereby ensuring the accuracy of the data.

[0185] In one embodiment, the method is applied to nodes in a distributed cluster, and offline applications and online applications are deployed on the nodes. Offline tasks include offline training tasks of the model, and online tasks include online reasoning tasks of the model. The offline application is used to execute the offline training tasks of the model, and the online application is used to execute the online reasoning tasks of the model.

[0186] Offline training tasks of models include offline training tasks of various models, such as image recognition models, feature extraction models, information recommendation models, etc., but not limited to these.

[0187] like Fig.10As shown, the resource allocation method is applied to any node in a distributed cluster. The node is a computer device. Offline applications and online applications are used on each node. Offline applications are used to execute offline training tasks of the model, and online applications are used to execute online reasoning tasks of the model, so that offline training tasks and online reasoning tasks of the same model can be mixed and deployed on one node, so that the offline training tasks and online reasoning tasks of the model share GPU resources. In addition, in order to ensure the response quality of the online reasoning tasks of the model, the GPU resources of the offline training tasks of the model can be dynamically regulated. On the premise of meeting the demand for GPU resources of the online reasoning tasks of the model, GPU resources are allocated to the offline training tasks of the model to make full use of GPU resources and avoid waste of GPU resources.

[0188] It is understandable that each node can implement each step of the resource allocation method in the above embodiments. Figure 6 The interaction process between interception tools, resource schedulers, offline applications and online applications.

[0189] In one embodiment, Fig.11 As shown, a timing diagram of a resource allocation method is provided, which is applied to a computer device. The resource scheduler includes a communication unit and a planning unit, and the interception tool includes a kernel processing unit and a statistical unit. The processing process of the resource allocation method is as follows:

[0190] 1) The offline application sends a resource allocation request to the communication unit of the resource scheduler. The resource allocation request is used to request the allocation of GPU resources for offline tasks from the shared GPU resources. The shared GPU resources are shared by offline applications and online applications. Offline applications are used to execute offline tasks, and online applications are used to execute online tasks.

[0191] 2) The core processing unit of the interception tool intercepts the resource allocation request and sends a resource planning request to the statistical unit.

[0192] 3) The statistical unit of the resource scheduler obtains offline usage information of shared GPU resources by offline tasks from offline applications.

[0193] 4) The statistical unit of the resource scheduler obtains the online usage information of online tasks on shared GPU resources from online applications.

[0194] 5) The statistical unit sends the offline usage information and the online usage information to the communication unit.

[0195] 6) The communication unit feeds back the offline usage information and online usage information to the planning unit.

[0196] 7) The planning unit determines the idle GPU resources of the shared GPU resources according to the offline usage information and the online usage information; and plans the expected usage information of the offline tasks for the shared GPU resources according to the offline usage information, the online usage information and the idle GPU resources.

[0197] 8) The planning unit performs resource planning for the shared GPU resources according to the offline usage information and the expected usage information, and obtains the GPU resource information to be allocated for the offline task.

[0198] 9) The planning unit transmits the GPU resource information to be allocated to the communication unit.

[0199] 10) The communication unit feeds back the information of GPU resources to be allocated to the core processing unit.

[0200] 11) The core processing unit allocates GPU resources to the offline application from the shared GPU resources according to the information of the GPU resources to be allocated.

[0201] In one embodiment, a resource allocation method is provided, which is applied to a computer device, wherein the computer device can run offline applications and online applications, the offline applications are used to execute offline tasks, the online applications are used to execute online tasks, and the shared GPU resources are shared by the offline applications and the online applications.

[0202] A resource scheduler and an interception tool are deployed on the computer device. The resource scheduler includes a communication unit and a planning unit, and the interception tool includes a kernel processing unit and a statistical unit. The processing process of the resource allocation method is as follows:

[0203] The offline application sends a resource allocation request to the communication unit of the resource scheduler, where the resource allocation request is used to request allocation of a GPU core to the offline task from the shared GPU core.

[0204] The core processing unit of the interception tool intercepts the resource allocation request and sends a resource planning request to the statistics unit.

[0205] The statistical unit of the resource scheduler determines the current stage of the online task in response to the resource planning request.

[0206] When the current stage of the online task is the first stage, the statistical unit obtains the first offline core quantity and the first offline core utilization rate of the shared GPU core of the offline task in the first stage from the offline application, and obtains the first online core quantity and the first online core utilization rate of the shared GPU core of the online task in the first stage from the online application.

[0207] The statistical unit sends the first number of offline cores, the first offline core utilization, the first number of online cores, and the first online core utilization to the communication unit, which then feeds back to the planning unit.

[0208] The planning unit determines a first idle GPU core number of the shared GPU core according to the first offline core number and the first online core number.

[0209] The planning unit plans a first expected core utilization rate of the offline task for the shared GPU core in the first phase according to the first number of offline cores, the first offline core utilization rate, the first number of online cores, the first online core utilization rate and the first number of idle GPU cores.

[0210] The planning unit calculates the first stage number of GPU cores to be allocated when the GPU core utilization of the offline task reaches the first expected core utilization according to the first offline core number, the first offline core utilization and the first expected core utilization.

[0211] The planning unit transmits the first-stage quantity to the communication unit, which then feeds it back to the core processing unit.

[0212] The core processing unit allocates a first-stage number of GPU cores from the shared GPU cores to the offline application.

[0213] In this embodiment, when the current stage of the online task is the second stage, the statistical unit obtains the second offline core number and the second offline core utilization of the shared GPU core of the offline task in the second stage from the offline application, and obtains the second online core number and the second online core utilization of the shared GPU core of the online task in the second stage from the online application.

[0214] The statistical unit sends the number of second offline cores, the utilization rate of the second offline cores, the number of second online cores, and the utilization rate of the second online cores to the communication unit, which then feeds back to the planning unit.

[0215] The planning unit determines the second idle GPU core number of the shared GPU core according to the second offline core number and the second online core number.

[0216] The planning unit plans a second expected core utilization of the shared GPU core for the offline task in the second phase according to the second offline core quantity, the second offline core utilization, the second online core quantity, the second online core utilization and the second idle GPU core quantity.

[0217] The planning unit calculates the second stage number of GPU cores to be allocated when the GPU core utilization of the offline task reaches the second expected core utilization according to the second offline core number, the second offline core utilization and the second expected core utilization.

[0218] The planning unit transmits the second-stage quantity to the communication unit, which then feeds it back to the core processing unit.

[0219] The core processing unit allocates a second-stage number of GPU cores from the shared GPU cores to the offline application.

[0220] In this embodiment, the first stage is the peak period of the online task's demand for the shared GPU kernel, and the second stage is the trough period of the online task's demand for the shared GPU kernel. According to the historical usage and utilization of the GPU resources by the online task at different stages, and the historical usage and utilization of the shared GPU kernel by the offline task, the utilization of the GPU kernel resources by the offline task at the peak and trough periods is planned, so as to accurately calculate the number of GPU kernels expected to be allocated when the offline task reaches the planned GPU kernel utilization, so as to control the GPU kernels allocated to the offline task at different stages, and realize the dynamic regulation of the GPU kernel resources of the offline task. In this way, it is not only possible to avoid the competition between the online task and the offline task for the GPU kernel resources at the peak and trough periods. In addition, the number of GPU kernels allocated to the offline task can be dynamically regulated to the offline task by the idle GPU kernel of the online task on the basis of meeting the demand of the online task for the GPU kernel, so as to realize the dynamic regulation of the number of GPU kernels of the offline task, thereby improving the utilization of the GPU kernel resources.

[0221] In the traditional solution, since the online reasoning task and the offline training task use the cores of different GPU cards, the GPU resources of the two GPU cards do not interfere with each other, and the online reasoning task has different requirements for the number of cores in the GPU card at different times, which leads to a waste of GPU resources. Therefore, this solution provides an application scenario of a resource allocation method, so that offline training tasks and online reasoning tasks share GPU resources.

[0222] In this application scenario, an offline training task and an online reasoning task share a core in a GPU card. If the offline training task and the online reasoning task are mixed and deployed on the same GPU card, the two tasks are likely to compete for the cores in the GPU card, resulting in the problem of reduced service quality of the online reasoning task. Therefore, in this embodiment, based on the demand for the number of cores in the GPU resource card for the online reasoning task, the number of GPU cores allocated to the offline training task is dynamically adjusted, so that in the time period when the demand for GPU cores for the online reasoning task is high, the GPU cores allocated to the offline training task are reduced, and in the time period when the demand for GPU cores for the online reasoning task is low, more idle GPU cores are allocated to the offline training task. In this way, the problems of insufficient number of GPU cores for offline training tasks and long waiting time in queues can be solved without affecting the service quality of online reasoning tasks, thereby improving the utilization of GPU resources.

[0223] The architecture diagram of the resource allocation method in this application scenario is as follows Fig.12 As shown in the figure. This resource allocation method is implemented based on the interception of kernels by the Compute Unified Device Architecture (CUDA). The interception tool intercepts the kernel requests sent by offline applications to the GPU card for offline training tasks. Kernel requests are resource allocation requests. Based on the interception of the GPU kernel request, the kernel sending is controlled to control the utilization rate of the kernels in the GPU card by offline training tasks, thereby avoiding the impact on online reasoning tasks. Both offline applications and online applications are applications running on the Compute Unified Device Architecture.

[0224] Fig.12 It is an architectural diagram of the resource allocation method on a node in a distributed cluster. The architecture mainly includes offline pods, an interception tool hook.so, and a resource scheduler deployed on the node. The resource scheduler is also called the colocation control component colocation agent. Among them, hook.so can be an interception library, which is injected into the pod of the offline training task submitted by the business through the webhook mechanism of kubernete, and takes effect through the preloading mechanism LD_PRELOAD or / etc / ld.so.preload. When the offline application initiates a kernel request, it will be intercepted by the interception tool. The resource scheduler is deployed on each node in the form of kubernetesdaemonset, which is responsible for managing the offline GPU usage of all GPU cards on the node. There will be multiple GPU cards on a node, and each GPU card will schedule at most one offline training task, that is, an offline training task can only use the kernel of one GPU card.

[0225] The interception tool includes the statistical unit stat-reporter and the kernel processing unit KernelHandler. The statistical unit is mainly responsible for periodically reporting the number of cores actually allocated to the offline training task within the cycle, and resetting the number of allocated cores to 0 at the end of the cycle to trigger the release of cores for a new cycle. The kernel processing unit is responsible for interacting with the resource scheduler to obtain the target number of cores to be allocated to the offline application fed back by the resource scheduler.

[0226] In addition, the core processing unit is also responsible for deciding whether to issue a core and how many cores to issue, etc. Each time a core is issued, the number of cores actually allocated increases by one until the number of cores issued reaches the target number of this cycle.

[0227] In the resource scheduler, each GPU card corresponds to a circular array with a length of 60 seconds. Each item in the circular array records the number of kernels actually allocated by the interception tool within this second, that is, the number of offline kernels, and the offline kernel utilization of the collected offline training tasks for the GPU card. The planning unit in the resource scheduler periodically obtains all offline kernel numbers and offline kernel utilizations from the circular array. Through the offline kernel numbers and offline kernel utilizations, it calculates how many kernels need to be sent to the offline training tasks so that the utilization of the offline training tasks for the GPU card can reach the expected kernel utilization, that is, calculates the target number of kernels that need to be sent. After the calculation is completed, the communication unit SocketServer of the resource scheduler synchronizes the target number to the kernel processing unit of the interception tool as the basis for the kernel sending of the next cycle.

[0228] The specific processing flow is as follows:

[0229] First, create a pod for the offline training task in the k8s cluster. The webhook in the cluster will intercept the pod create request and modify the pod spec, including injecting the init container and modifying the volume mount of the GPU container. The initcontainer contains the hook.so file, which is copied to an emptydir volume in the init container. The emptydir will be mounted in the GPU container at the same time and the LD_PRELOAD environment variable will be injected, with the address pointing to the hook.so file in the emptydir.

[0230] After the offline training task process is started, the interception tool will be loaded. The interception tool will establish a remote procedure call (RPC) connection with the resource scheduler, register the current pod and GPU information with the resource scheduler, and obtain the initialized target number of cores that can be delivered in this cycle from the resource scheduler;

[0231] When the offline training task calls cuLuanchKernel related functions, the interception tool will intercept these calls and enter the control flow:

[0232] Determine whether the number of cores issued in this cycle is less than the target number. If so, continue to issue cores and update the number of issued cores.

[0233] If the number of cores issued in this cycle reaches the target number, the kernel issuance in this cycle ends until the next cycle begins, and the "determine whether the number of cores issued in this cycle is less than the target number" and subsequent steps are executed again;

[0234] After the kernel distribution of this cycle is completed, a thread of the interception tool will report the current number of distributed kernels to the resource scheduler;

[0235] After receiving the reported number of issued cores, the resource scheduler adds it to the corresponding position of the ring array of the GPU card;

[0236] The resource scheduler collects the utilization of the corresponding GPU card every second, and records it at the corresponding position of the GPU card's ring array, and then enters the next round of data collection. At the same time, the resource scheduler recalculates how many cores need to be issued to achieve the expected utilization based on the number and utilization of the issued cores in the current ring array, as well as the expected utilization, that is, recalculates the target number, and sends the target number back to the interception tool side;

[0237] The interception tool receives the recalculated target number as the upper limit of the number of kernels to be sent in the next cycle.

[0238] In this embodiment, CUDA interception technology is used to intercept the kernel delivery timing of offline training tasks without perception. Through the utilization control algorithm, the kernel delivery rate is controlled, thereby achieving the utilization control effect of offline training tasks. By controlling the utilization of offline training tasks, the occurrence of computing resource competition between offline training tasks and online reasoning services is avoided, and the service quality of online reasoning is avoided, making the hybrid deployment more secure and reliable.

[0239] In addition, this hybrid technology can dynamically adjust GPU resources of different granularities for offline training tasks, maximize the secondary utilization of idle GPU computing power, improve GPU utilization, solve the problem of GPU resource shortage in offline training tasks, and save GPU resources and costs.

[0240] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0241] Based on the same inventive concept, the embodiment of the present application also provides a resource allocation device for implementing the resource allocation method involved above. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in one or more resource allocation device embodiments provided below can refer to the limitations on the resource allocation method above, and will not be repeated here.

[0242] In one embodiment, Fig.13 As shown, a resource allocation device 1300 is provided, including:

[0243] The request acquisition module 1302 is used to acquire a resource allocation request of an offline task for shared GPU resources, where the shared GPU resources are shared by offline tasks and online tasks.

[0244] The determination module 1304 is used to determine offline usage information of the shared GPU resources by the offline task.

[0245] The expected acquisition module 1306 is used to acquire expected usage information of the offline task on the shared GPU resources, where the expected usage information is related to the online usage information of the online task on the shared GPU resources.

[0246] The planning module 1308 is used to perform resource planning on the shared GPU resources according to the offline usage information and the expected usage information, and obtain GPU resource information to be allocated for the offline task.

[0247] The allocation module 1310 is used to allocate GPU resources to offline tasks from shared GPU resources according to the information of GPU resources to be allocated.

[0248] In this embodiment, offline tasks and online tasks share GPU resources. By obtaining the resource allocation request of the offline task for the shared GPU resources, the offline usage information of the offline task for the shared GPU resources is determined to obtain the historical usage information of the offline task for the shared GPU resources. The expected usage information of the offline task for the shared GPU resources is obtained to determine the GPU resource usage information planned for the offline task, and the planned usage information is related to the online usage information of the online task for the shared GPU resources, so that the use of the GUP resources by the online task is taken into account when planning the GPU resource usage information of the offline task, and the competition for the shared GUP resources when the online task and the offline task are executed can be avoided. According to the offline usage information and the expected usage information, the shared GPU resources are planned to estimate how many GUP resources need to be allocated to the offline task so that the usage information of the offline task for the shared GPU resources is consistent with the expected usage information. According to the GPU resource information to be allocated, the GPU resources are allocated to the offline task from the shared GPU resources, so that the use of the GPU resources by the offline task is controlled within the expected range, and the dynamic regulation of the GPU resources is realized. Moreover, when online tasks and offline tasks share GPU resources, the utilization of GPU resources can be improved by dynamically adjusting the allocation of GPU resources to offline tasks, which can effectively avoid the waste of GPU resources.

[0249] In one embodiment, the offline usage information includes offline resource usage and offline resource utilization, the expected usage information includes expected resource utilization, and the to-be-allocated GPU resource information includes a target number of to-be-allocated GPU resources;

[0250] The planning module 1308 is further used to perform resource planning on the shared GPU resources according to the offline resource usage and the offline resource utilization, and obtain the target number of GPU resources to be allocated when the resource utilization of the offline task reaches the expected resource utilization;

[0251] The allocation module 1310 is further configured to allocate a target number of GPU resources from the shared GPU resources to the offline task.

[0252] In this embodiment, resource planning is performed on shared GPU resources based on offline resource usage, offline resource utilization and expected resource utilization, so that the target number of GPU resources required for offline tasks to achieve the expected resource utilization can be accurately calculated based on the historical usage, historical utilization and expected utilization of shared GPU resources by offline resources, thereby controlling the GPU resources allocated to offline tasks within expectations and realizing dynamic regulation of GPU resources for offline tasks.

[0253] In one embodiment, the offline resource usage includes the periodic offline resource usage of the offline task in the historical period, the offline resource utilization includes the periodic offline resource utilization of the offline task in the historical period, the expected resource utilization includes the periodic expected resource utilization of the offline task in the planned period, and the target number of GPU resources to be allocated includes the periodic number of GPU resources to be allocated when the resource utilization of the offline task in the planned period reaches the expected resource utilization;

[0254] The allocation module 1310 is further configured to allocate a number of GPU resources of the period from the shared GPU resources to the offline task within the planning period.

[0255] In this embodiment, based on the usage and utilization rate of shared GPU resources by offline resources in historical cycles and the utilization rate planned to be achieved in the next cycle, the estimated number of GPU resources to be allocated when the offline tasks reach the planned resource utilization rate in the next cycle is accurately calculated, thereby periodically controlling the GPU resources allocated to offline tasks and realizing dynamic regulation of the GPU resources of offline tasks.

[0256] In one embodiment, the shared GPU resources include GPU cores; periodic offline resource usage, including the number of offline cores allocated to offline tasks in a historical period; periodic offline resource utilization, including the offline core utilization for offline tasks in a historical period; periodic expected resource utilization, including the expected core utilization for offline tasks in a planned period; and the number of cycles of GPU resources to be allocated, including the number of cycles of GPU cores to be allocated when the resource utilization of offline tasks in a planned period reaches the expected resource utilization.

[0257] In this embodiment, the number of GPU cores allocated in the next period is predicted by periodically calculating the number and utilization of GPU cores allocated, as well as the expected core utilization. This allows for periodic regulation of the allocation of GPU cores to offline tasks, making the allocation of GPU cores to offline tasks controllable.

[0258] In one embodiment, the allocation module 1310 is further used to determine the core allocation rate within the planned cycle according to the planned cycle and the number of GPU cores to be allocated within the cycle; within the planned cycle, GPU cores are allocated to offline tasks multiple times according to the core allocation rate until the number of allocated GPU cores reaches the number of cycles.

[0259] In this embodiment, the core allocation rate within the planned cycle is determined according to the planned cycle and the number of GPU cores to be allocated in the cycle, so that within the planned cycle, GPU cores are allocated to offline tasks multiple times according to the core allocation rate until the number of allocated GPU cores reaches the number of cycles, thereby controlling the rate at which GPU cores are issued within the cycle and making better use of the GPU cores.

[0260] In one embodiment, the expected acquisition module 1306 is also used to obtain online usage information of shared GPU resources by online tasks; determine the idle GPU resources of the shared GPU resources based on the offline usage information and the online usage information; and configure expected usage information for the shared GPU resources for the offline tasks based on the offline usage information, the online usage information and the idle GPU resources.

[0261] In this embodiment, historical usage information of shared GPU resources by online tasks is obtained, and idle GPU resources of shared GPU resources are calculated based on historical usage information of offline tasks and historical usage information of online tasks. Based on offline usage information, online usage information and idle GPU resources, usage information of shared GPU resources can be accurately planned for offline tasks, so as to control the usage of GPU resources by offline tasks within an expected range and avoid insufficient supply of GPU resources.

[0262] In one embodiment, the online usage information includes at least one of first online usage information of the online task in the first stage or second online usage information in the second stage, and the usage of the shared GPU resources by the online task in the first stage is greater than the usage of the shared GPU resources in the second stage; the offline usage information includes at least one of first offline usage information of the offline task in the first stage or second offline usage information in the second stage, the idle GPU resources include at least one of the first idle GPU resources in the first stage or the second idle GPU resources in the second stage, and the expected usage information includes at least one of the first expected usage information in the first stage or the second expected usage information in the second stage.

[0263] In this embodiment, the first stage is the peak period of the demand for shared GPU resources by online tasks, and the second stage is the trough period of the demand for shared GPU resources by online tasks. According to the usage of GPU resources by online tasks during the peak period and the trough period, the usage of GPU resources by offline tasks during the peak period and the trough period is planned, so as to avoid the competition for GPU resources between online tasks and offline tasks during the peak period and the trough period. In addition, by planning the GPU resources allocated to offline tasks, it is possible to dynamically adjust the idle GPU resources of online tasks to offline tasks on the basis of meeting the demand for GPU resources by online tasks, thereby improving the utilization rate of GPU resources.

[0264] In one embodiment, the request acquisition module 1302 is further used to intercept the resource allocation request sent to the resource scheduler through an interception tool, where the resource allocation request is used to request allocation of GPU resources for offline tasks from shared GPU resources;

[0265] The allocation module 1310 is used to feed back the information of GPU resources to be allocated to the interception tool through the resource scheduler; and allocate GPU resources to the offline task from the shared GPU resources according to the information of GPU resources to be allocated through the interception tool.

[0266] In this embodiment, the resource allocation request sent to the resource scheduler is intercepted by an interception tool, so as to realize the request timing of the resource allocation request for the offline task through the interception tool. The GPU resource information that can be allocated to the offline task this time is planned by the resource scheduler to realize the dynamic regulation of the GPU resources of the offline task. The information of the GPU resources to be allocated is fed back to the interception tool, and the GPU resources are allocated to the offline task from the shared GPU resources according to the information of the GPU resources to be allocated through the interception tool, so that the timing of the GPU resources being issued can be controlled through the interception tool.

[0267] In one embodiment, the interception tool includes a statistical unit and a kernel processing unit; the determination module 1304 is further used to obtain offline usage information of the shared GPU resources by the offline task through the statistical unit, and report the offline usage information to the resource scheduler;

[0268] The allocation module 1310 is also used to receive the to-be-allocated GPU resource information fed back by the resource scheduler through the core processing unit; and allocate GPU resources to offline tasks from the shared GPU resources through the core processing unit according to the to-be-allocated GPU resource information.

[0269] In this embodiment, the interception tool includes a statistical unit and a kernel processing unit. Through the statistical unit, offline usage information of shared GPU resources by offline tasks is obtained, and the offline usage information is reported to the resource scheduler to provide the resource scheduler with the latest GPU resource allocation information in a timely manner. Through the kernel processing unit, the to-be-allocated GPU resource information fed back by the resource scheduler is received, and through the kernel processing unit, according to the to-be-allocated GPU resource information, GPU resources are allocated to offline tasks from shared GPU resources, so that dynamic regulation of GPU resources for offline tasks in GPU resource sharing scenarios can be achieved through interaction between the kernel processing unit and the resource scheduler.

[0270] In one embodiment, the apparatus further comprises:

[0271] The update module is used to feed back the resource allocation result to the resource scheduler after allocating GPU resources to the offline task through the statistical unit; and update the offline usage information of the offline task on the shared GPU resource based on the resource allocation result through the resource scheduler.

[0272] In this embodiment, after allocating GPU resources to offline tasks, the statistical unit feeds back the resource allocation results to the resource scheduler. The resource scheduler timely updates the offline usage information of the offline tasks for the shared GPU resources based on the resource allocation results, thereby ensuring the accuracy of the data.

[0273] In one embodiment, the device is applied to a node in a distributed cluster, and the node deploys offline applications and online applications. The offline tasks include offline training tasks of the model, and the online tasks include online reasoning tasks of the model. The offline application is used to execute the offline training tasks of the model, and the online application is used to execute the online reasoning tasks of the model.

[0274] Each module in the resource allocation device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0275] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig.14 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store relevant data on resource allocation. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a resource allocation method is implemented.

[0276] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Fig.15As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a resource allocation method is implemented. The display unit of the computer device is used to form a visually visible image, and can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse, etc.

[0277] Those skilled in the art will understand that Fig.14 and Fig.15 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0278] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0279] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0280] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0281] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0282] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0283] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0284] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A resource allocation method, characterized in that: The method comprises: Obtaining a resource allocation request of an offline task for a shared GPU resource, where the shared GPU resource is shared by the offline task and the online task; Determining offline usage information of the shared GPU resource by the offline task; Acquire expected usage information of the shared GPU resource by the offline task, where the expected usage information is related to online usage information of the shared GPU resource by the online task; Perform resource planning on the shared GPU resources according to the offline usage information and the expected usage information to obtain GPU resource information to be allocated for the offline task; According to the information of the GPU resources to be allocated, GPU resources are allocated to the offline task from the shared GPU resources.

2. The method according to claim 1, characterized in that The offline usage information includes offline resource usage and offline resource utilization, the expected usage information includes expected resource utilization, and the to-be-allocated GPU resource information includes a target number of to-be-allocated GPU resources; The performing resource planning on the shared GPU resources according to the offline usage information and the expected usage information to obtain the GPU resource information to be allocated for the offline task includes: According to the offline resource usage, the offline resource utilization and the expected resource utilization, resource planning is performed on the shared GPU resources to obtain a target number of GPU resources to be allocated when the resource utilization of the offline task reaches the expected resource utilization; The allocating GPU resources to the offline task from the shared GPU resources according to the to-be-allocated GPU resource information includes: The target number of GPU resources is allocated to the offline task from the shared GPU resources.

3. The method according to claim 2, characterized in that The offline resource usage includes the periodic offline resource usage of the offline task in the historical period, the offline resource utilization includes the periodic offline resource utilization of the offline task in the historical period, the expected resource utilization includes the periodic expected resource utilization of the offline task in the planned period, and the target number of GPU resources to be allocated includes the periodic number of GPU resources to be allocated when the resource utilization of the offline task in the planned period reaches the expected resource utilization; The allocating the target number of GPU resources to the offline task from the shared GPU resources includes: In the planning cycle, the number of GPU resources in the cycle is allocated to the offline task from the shared GPU resources.

4. The method according to claim 3, characterized in that The shared GPU resources include shared GPU cores; the periodic offline resource usage includes the number of offline cores allocated to the offline task in the historical period; the periodic offline resource utilization includes the offline core utilization for the offline task in the historical period; The expected resource utilization rate of the cycle includes the expected kernel utilization rate for the offline task within the planned cycle; the number of cycles of GPU resources to be allocated includes the number of cycles of GPU kernels to be allocated when the kernel utilization rate of the offline task reaches the expected kernel utilization rate within the planned cycle.

5. The method according to claim 4, characterized in that The allocating the period number of GPU resources from the shared GPU resources to the offline task within the planning period includes: Determine a core allocation rate within the planned period according to the planned period and the number of GPU cores to be allocated during the period; In the planning cycle, GPU cores are allocated to the offline task multiple times according to the core allocation rate until the number of allocated GPU cores reaches the cycle number.

6. The method according to claim 1, characterized in that The obtaining expected usage information of the shared GPU resource by the offline task includes: Obtaining online usage information of the shared GPU resources by the online task; Determining idle GPU resources of the shared GPU resources according to the offline usage information and the online usage information; According to the offline usage information, the online usage information and the idle GPU resources, expected usage information of the offline task for the shared GPU resources is planned.

7. The method according to claim 6, characterized in that The online usage information includes at least one of first online usage information of the online task in the first stage or second online usage information of the online task in the second stage, and the usage of the shared GPU resource by the online task in the first stage is greater than the usage of the shared GPU resource in the second stage; The offline usage information includes at least one of the first offline usage information of the offline task in the first stage or the second offline usage information in the second stage, the idle GPU resources include at least one of the first idle GPU resources in the first stage or the second idle GPU resources in the second stage, the expected usage information includes at least one of the first expected usage information in the first stage or the second expected usage information in the second stage, and the to-be-allocated GPU resource information includes at least one of the first to-be-allocated GPU resource information in the first stage or the second to-be-allocated GPU resource information in the second stage.

8. The method according to claim 1, characterized in that The obtaining of the resource allocation request of the offline task for the shared GPU resources includes: By using an interception tool, a resource allocation request sent to a resource scheduler is intercepted, wherein the resource allocation request is used to request allocation of GPU resources for offline tasks from shared GPU resources; The allocating GPU resources to the offline task from the shared GPU resources according to the to-be-allocated GPU resource information includes: Feedback the to-be-allocated GPU resource information to the interception tool through the resource scheduler; By means of the interception tool, GPU resources are allocated to the offline task from the shared GPU resources according to the information of the GPU resources to be allocated.

9. The method according to claim 8, characterized in that The interception tool includes a statistical unit and a kernel processing unit; the determining the offline usage information of the shared GPU resource by the offline task includes: Obtaining, by means of the statistical unit, offline usage information of the shared GPU resource by the offline task, and reporting the offline usage information to the resource scheduler; The allocating GPU resources to the offline task from the shared GPU resources according to the to-be-allocated GPU resource information by the interception tool includes: Receiving, through the core processing unit, the to-be-allocated GPU resource information fed back by the resource scheduler; The core processing unit allocates GPU resources to the offline task from the shared GPU resources according to the information of the GPU resources to be allocated.

10. The method according to claim 9, characterized in that The method further comprises: After allocating GPU resources to the offline task, the statistical unit feeds back the resource allocation result to the resource scheduler; The resource scheduler is used to update offline usage information of the shared GPU resource by the offline task based on the resource allocation result.

11. The method according to any one of claims 1 to 10, characterized in that: The method is applied to a node in a distributed cluster, wherein the node deploys an offline application and an online application, wherein the offline task includes an offline training task of a model, and the online task includes an online reasoning task of the model, wherein the offline application is used to execute the offline training task of the model, and the online application is used to execute the online reasoning task of the model.

12. A resource allocation device, characterized in that: The device comprises: A request acquisition module, used to acquire a resource allocation request of an offline task for a shared GPU resource, where the shared GPU resource is shared by the offline task and the online task; A determination module, used to determine offline usage information of the shared GPU resource by the offline task; an expectation acquisition module, configured to acquire expected usage information of the shared GPU resource by the offline task, wherein the expected usage information is related to online usage information of the shared GPU resource by the online task; A planning module, configured to perform resource planning on the shared GPU resources according to the offline usage information and the expected usage information, and obtain GPU resource information to be allocated for the offline task; The allocation module is used to allocate GPU resources to the offline task from the shared GPU resources according to the information of the GPU resources to be allocated.

13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Cited By

  • GPU equipment isolation and mounting method and device, equipment and storage medium

    CN121143951A