Resource management method and device, storage medium, equipment and program product

By real-time detection of mixed management data of shared GPU clusters, calculating the amount of idle resources and adjusting offline task scheduling strategies, the resource waste caused by the management of online tasks and offline tasks is solved, and resource utilization and task execution efficiency are improved.

CN120276855APending Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510391204.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Under the traditional resource management model, purchasing and managing GPU resources on online and offline tasks respectively leads to low resource utilization, especially when online tasks are in a trough, GPU resources are idle, causing waste.

Method used

By real-time detection of mixed management data of target services in the shared GPU cluster, the amount of idle GPU resources for online tasks is calculated, and offline task scheduling strategies are dynamically adjusted to realize the dynamic sharing of resources between online tasks and offline tasks.

Benefits of technology

It effectively avoids the idleness and waste of online GPU resources, significantly improves GPU resource utilization and task execution efficiency, and reduces business operation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276855A_ABST
    Figure CN120276855A_ABST
Patent Text Reader

Abstract

The invention discloses a resource management method and device, a storage medium, equipment and a program product, which are applied to a resource management scene. The method comprises the following steps: detecting hybrid management data of a target service in a shared graphics processing unit (GPU) cluster in real time, wherein the hybrid management data comprises use information of a first GPU resource of an online task under the target service and resource demand information of an offline task under the target service; calculating idle GPU resource quantity in the first GPU resource of the online task based on the hybrid management data; according to the method, the offline task scheduling strategy is adjusted based on the idle GPU resource quantity and the resource demand information of the offline task under the target service, the offline task scheduling strategy indicates the second GPU resource used by the offline task under the target service, dynamic resource sharing of the online task and the offline task is achieved, and idleness and waste of the online GPU resource are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a resource management method, apparatus, storage medium, device, and program product. Background Art

[0002] With the rapid development of artificial intelligence (AI) and big data technologies, the graphics processing unit (GPU) has become a key resource for accelerating online and offline data processing tasks. Online tasks, such as real-time data analysis, image recognition, and natural language processing, etc., have real-time and high-performance requirements for GPU resources; while offline tasks, such as model training and large-scale data processing, etc., although have lower requirements for real-time performance, have large demand for computing resources and long duration.

[0003] However, in the traditional resource management mode, different GPU resources need to be purchased and managed separately for online tasks and offline tasks, which not only increases the cost, but also leads to low resource utilization. Especially in the case where there are peaks and valleys in online tasks, the GPU resource segments for online tasks are largely idle during valleys, resulting in resource waste. Summary of the Invention

[0004] Embodiments of this application provide a resource management method, apparatus, storage medium, device, and program product, which realize the dynamic sharing of resources between online tasks and offline tasks, and avoid the idle and waste of online GPU resources.

[0005] On the one hand, embodiments of this application provide a resource management method, and the method includes:

[0006] Real-time detect the co-location management data of the target service in the shared graphics processing unit (GPU) cluster, where the co-location management data includes the usage information of the first GPU resources of the online tasks under the target service, and the resource demand information of the offline tasks under the target service;

[0007] Based on the co-location management data, calculate the amount of idle GPU resources in the first GPU resources of the online tasks;

[0008] Based on the amount of idle GPU resources and the resource demand information of the offline tasks under the target service, adjust the offline task scheduling policy, where the offline task scheduling policy indicates the second GPU resources for the offline tasks under the target service.

[0009] On the other hand, embodiments of this application provide a resource management apparatus, and the apparatus includes:

[0010] A detection unit for real-time detecting the co-location management data of a target service in a shared Graphics Processing Unit (GPU) cluster, where the co-location management data includes the usage information of the first GPU resources of the online tasks under the target service and the resource requirement information of the offline tasks under the target service;

[0011] A calculation unit for calculating the amount of idle GPU resources in the first GPU resources of the online tasks based on the co-location management data;

[0012] An adjustment unit for adjusting the offline task scheduling policy based on the amount of idle GPU resources and the resource requirement information of the offline tasks under the target service, where the offline task scheduling policy indicates the second GPU resources for the offline tasks under the target service.

[0013] In some embodiments, the shared GPU cluster is configured in a three-layer cluster architecture, and the three-layer cluster architecture includes:

[0014] An upper-layer offline co-location cluster for receiving the offline tasks of each service;

[0015] An intermediate co-location management layer configured with a first co-location component, where the first co-location component is used for detecting the co-location management data, calculating the amount of idle GPU resources, and adjusting the offline task scheduling policy;

[0016] A lower-layer GPU cluster configured with a second co-location component, where the second co-location component is used for managing each GPU node in the shared GPU cluster.

[0017] In some embodiments, the co-location management data further includes the co-location validity period and co-location status of the online and offline co-location; the calculation unit is used for: regularly detecting whether the current time is within the co-location validity period; if the current time is within the co-location validity period, putting the offline tasks under the target service into the co-location processing queue; and calculating the amount of idle GPU resources in the first GPU resources of the online tasks based on the co-location management data.

[0018] In some embodiments, the hybrid management data further includes the quality of service indicators and detection rules of the online tasks; the adjustment unit is configured to: based on the quality of service indicators and detection rules of the online tasks, and the usage information of the first GPU resources, detect the online service quality of the online tasks in real time; according to the amount of idle GPU resources and the online service quality, determine the offline task scheduling strategy as issuing offline tasks or evicting offline tasks; if the decision is to issue offline tasks, then based on the amount of idle GPU resources and the resource requirement information of the offline tasks under the target service, determine the second GPU resources for the offline tasks under the target service, and issue the offline tasks under the target service to the second GPU resources; or if the decision is to evict offline tasks, then partially or fully reclaim the idle GPU resources occupied by the offline tasks, and reallocate the reclaimed idle GPU resources to the online tasks.

[0019] In some embodiments, the adjustment unit is configured to;

[0020] If the decision is to issue offline tasks, then through the first hybrid component of the intermediate hybrid management layer, based on the amount of idle GPU resources and the resource requirement information of the offline tasks under the target service, determine the second GPU resources for the offline tasks under the target service;

[0021] Through the first hybrid component, issue the first offline task scheduling strategy corresponding to the offline tasks under the target service to the upper-layer offline hybrid cluster, and allocate the offline tasks under the target service to the GPU nodes corresponding to the second GPU resources in the underlying GPU cluster.

[0022] In some embodiments, if the decision is to issue offline tasks, the adjustment unit is further configured to: after allocating the second GPU resources, detect whether there are still remaining GPU resources in the amount of idle GPU resources; if there are still remaining GPU resources in the amount of idle GPU resources, then through the first hybrid component, issue the second offline task scheduling strategy corresponding to the offline tasks under other services to the upper-layer offline hybrid cluster, and allocate the offline tasks under other services to the GPU nodes corresponding to the remaining GPU resources in the underlying GPU cluster; wherein, the priority of the offline tasks under other services using the amount of idle GPU resources is lower than the priority of the offline tasks under the target service using the amount of idle GPU resources.

[0023] In some embodiments, the adjustment unit is configured to: if the decision-making offline task scheduling policy is to evict offline tasks, evict the offline tasks in the order of a preset priority, where the preset priority order is: first evict the offline tasks under the other services, and then evict the offline tasks under the target service; reallocate the recovered idle GPU resources to the online tasks;

[0024] Wherein, if the eviction volume of evicting offline tasks is full eviction, set the hybrid state to the eviction state; or if the eviction volume of evicting offline tasks is partial eviction, set the hybrid state to the resource-constrained hybrid state.

[0025] In some embodiments, the adjustment unit is configured to evict the offline tasks under the other services, including: partially or fully recovering the remaining GPU resources occupied by the offline tasks under the other services, and reallocating the recovered remaining GPU resources to the online tasks;

[0026] The adjustment unit is configured to evict the offline tasks under the target service, including: partially or fully recovering the second GPU resources occupied by the offline tasks under the target service, and reallocating the recovered second GPU resources to the online tasks.

[0027] In some embodiments, the adjustment unit is further configured to: detect in real time whether the current time is the end time of the hybrid validity period; if the current time is not the end time of the hybrid validity period, repeat adjusting the offline task scheduling policy; or if the current time is the end time of the hybrid validity period, prohibit issuing offline tasks.

[0028] In some embodiments, the adjustment unit is further configured to: if the current time is not within the hybrid validity period, detect whether there are offline tasks under the target service in the underlying GPU cluster; if there are offline tasks under the target service in the underlying GPU cluster, evict the offline tasks under the target service from the underlying GPU cluster, and set the hybrid state to the non-hybrid state.

[0029] In some embodiments, the adjustment unit is further configured to: when evicting the offline tasks under the target service from the underlying GPU cluster, send an alarm message, where the alarm message is used to prompt the offline task eviction event.

[0030] In some embodiments, the adjustment unit is further configured to: perform quality of service control on the resources of each offline task through a second co-located component in the underlying GPU cluster, where the quality of service control includes adjusting the central processing unit (CPU) weight and network input / output priority of each offline task, and the offline tasks include at least one of the online tasks under the target service and the online tasks under other services; and re-adjust the offline task scheduling policy according to the result of the quality of service control.

[0031] On the other hand, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which is suitable for being loaded by a processor to execute the resource management method according to any one of the above embodiments.

[0032] On the other hand, an embodiment of the present application provides a computer device including a processor and a memory, where the memory stores a computer program, and the processor is configured to execute the resource management method according to any one of the above embodiments by calling the computer program stored in the memory.

[0033] On the other hand, an embodiment of the present application provides a computer program product including computer instructions, which implement the resource management method according to any one of the above embodiments when executed by a processor.

[0034] In the embodiments of the present application, the co-located management data of the target service in the shared graphics processing unit (GPU) cluster is detected in real time. The co-located management data includes the usage information of the first GPU resources of the online tasks under the target service and the resource demand information of the offline tasks under the target service. Based on the co-located management data, the amount of idle GPU resources in the first GPU resources of the online tasks is calculated. Based on the amount of idle GPU resources and the resource demand information of the offline tasks under the target service, the offline task scheduling policy is adjusted. The offline task scheduling policy indicates the second GPU resources for the offline tasks under the target service. By detecting the co-located management data of the target service in the shared GPU cluster in real time and accurately calculating the amount of idle GPU resources of the online tasks based on this data, and then dynamically adjusting the scheduling policy of the offline tasks, the embodiments of the present application effectively realize the dynamic sharing of resources between online tasks and offline tasks, avoid the idle and waste of online GPU resources, and significantly improve the GPU resource utilization rate and task execution efficiency of the target tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0036] Figure 1 This is a schematic diagram of the three - layer cluster architecture provided by the embodiments of the present application.

[0037] Figure 2 This is the first process schematic diagram of the resource management method provided by the embodiments of the present application.

[0038] Figure 3 This is the second process schematic diagram of the resource management method provided by the embodiments of the present application.

[0039] Figure 4 This is a schematic diagram of the structure of the resource management device provided by the embodiments of the present application.

[0040] Figure 5 This is a schematic diagram of the structure of the computer device provided by the embodiments of the present application. Detailed implementation manners

[0041] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0042] The embodiments of the present application provide a resource management method, device, storage medium, device and program product. Exemplarily, the resource management method of the embodiments of the present application can be executed by a computer device, where the computer device can be a terminal or a server and other devices. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart TV, a smart speaker, a wearable smart device, a personal computer (PC), a smart vehicle terminal and other devices, and the terminal can also include a client. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.

[0043] The embodiments of the present application can be applied to scenarios such as task scheduling and resource management.

[0044] First, some nouns or terms that appear in the process of describing the embodiments of the present application are explained as follows:

[0045] Kubernetes cluster: A resource scheduling and management platform for container orchestration. Users can submit tasks on cloud providers' Kubernetes-based platforms, and the tasks run in the form of containers. Its abbreviation is k8s.

[0046] Task instance (pod): The smallest scheduling unit in k8s. A task consists of multiple task instances.

[0047] Online-offline hybrid scheduling: When users submit online tasks to an online k8s cluster, they usually request more resources, and there are peak and valley phenomena during the task running process, which will cause the machine resources to be underutilized, especially during the valley period. During the resource idle period, the resource utilization rate is improved by filling in offline tasks, which is called online-offline hybrid scheduling.

[0048] GPU cluster: The nodes managed by the k8s cluster are GPU nodes, and this k8s cluster is called a GPU cluster.

[0049] Custom Resource Definition (CRD). CRD is an extension mechanism provided by Kubernetes that allows users to define and use custom resource types in the Kubernetes cluster. By defining CRD, users can create, read, update, and delete custom resource objects in the Kubernetes cluster.

[0050] Control group (cgroup): Used to limit, control, and separate the resources of process groups (such as CPU, memory, disk input / output, etc.).

[0051] Quality of Service (QoS): A network technology used to manage and ensure the quality and performance of different types in the network.

[0052] In the traditional resource management mode, it is necessary to purchase and manage different GPU resources for online tasks and offline tasks respectively, which not only increases the cost but also leads to low resource utilization. Especially when there are peaks and valleys in online tasks, a large amount of GPU resources for online tasks are idle during the valley period, resulting in resource waste.

[0053] Current GPU hybrid scheduling technologies mostly focus on the system level, unable to effectively utilize the idle GPU resources within the business, and lacking the ability to perform fine-grained resource allocation for offline tasks. In addition, due to the lack of an effective resource isolation and sharing mechanism, offline tasks between different businesses may interfere with each other, affecting the service quality and stability of online tasks.

[0054] Therefore, the embodiments of the present application provide a resource management method. By detecting the co-location management data of the target service in the shared GPU cluster in real time, and accurately calculating the idle GPU resource amount of the online tasks based on this data, and then dynamically adjusting the scheduling strategy of the offline tasks, the dynamic sharing of resources between the online tasks and the offline tasks is effectively realized, avoiding the idle and waste of online GPU resources, significantly improving the resource utilization rate and task execution efficiency. Instead of purchasing additional system-level offline GPU resources, the business operation cost is reduced, while the performance and stability of the business are ensured. The embodiments of the present application can adjust the resource allocation in real time according to the business requirements and resource status, and at the same time provide sufficient flexibility and fineness to meet various complex business scenarios.

[0055] The solution provided by the embodiments of the present application relates to technologies such as resource management, and is specifically described through the following embodiments. The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the priority order of the embodiments.

[0056] Please refer to Figure 1 , Figure 1 , which is the architecture schematic diagram of the three-layer cluster architecture provided by the embodiments of the present application. Among them, the resource management method provided by the embodiments of the present application is based on the three-layer cluster architecture 1, and the shared GPU cluster can be configured in the three-layer cluster architecture 1. The three-layer cluster architecture 1 may include:

[0057] The upper-layer offline co-location cluster 10, which is used to receive the offline tasks of each service;

[0058] The middle co-location management and control layer 20, which is configured with a first co-location component. The first co-location component is used to detect the co-location management data, calculate the idle GPU resource amount, and adjust the offline task scheduling strategy;

[0059] The lower-layer GPU cluster 30, which is configured with a second co-location component. The second co-location component is used to manage each GPU node in the shared GPU cluster.

[0060] Among them, the upper-layer offline co-location cluster 10 is connected to the offline training platform 2 and is used to receive the offline tasks of each service submitted by the user from the offline training platform 2. For example, business-1, business-2, business-3, etc. Each business includes multiple offline tasks, and the offline task is a task instance (pod) for offline training. For example, the upper-layer offline co-location cluster 10 is a single k8s cluster and is a part of the shared GPU cluster.

[0061] Among them, the intermediate hybrid management and control layer 20 is respectively connected to the upper-layer offline hybrid cluster 10 and the lower-layer GPU cluster 30. Among them, the first hybrid component configured in the intermediate hybrid management and control layer 20 can be used to be responsible for the entire life cycle of the online and offline hybrid of each service sharing the GPU cluster. A custom resource definition (CRD) of the online and offline hybrid of the shared GPU cluster is maintained for each service. The online and offline hybrid CRD records the hybrid management data of the target service. For example, the hybrid management data includes hybrid configuration information, the usage information of the first GPU resources of the online tasks under the target service, as well as the resource requirements of the offline tasks under the target service, the quality of service indicators and detection rules of the online tasks, etc. For example, the hybrid configuration information of the service contains some basic hybrid information of the service, such as recording the basic configuration information related to the online and offline hybrid of the service, such as including but not limited to: the hybrid validity period of the online and offline hybrid (the start time and end time of the hybrid validity period), the graceful eviction policy of the offline task, the target hybrid cluster corresponding to the service, the maximum hybrid resource limit, the resource reservation condition, etc. The hybrid configuration information is recorded in the form of CRD of the k8s cluster, and the real-time hybrid state of the service will also be recorded in the CRD in real time.

[0062] Among them, the target hybrid cluster corresponding to the service refers to a target k8s cluster or multiple target k8s clusters to which the GPU resources will be scheduled and allocated when the service performs the online and offline hybrid operation.

[0063] Among them, the maximum hybrid resource limit refers to the maximum amount of GPU resources that the system sets for the service when the service performs the online and offline hybrid. The maximum hybrid resource limit is used to prevent the offline tasks of a certain service from excessively occupying the GPU resources, thereby affecting the quality of service of other services or the online tasks of the service itself. The system will set the maximum hybrid resource limit according to service requirements, the total amount of GPU resources, and other relevant factors.

[0064] Among them, the resource reservation condition refers to the rules or conditions for the system to reserve a certain amount of GPU resources for the service when the service performs the online and offline hybrid. The condition can be configured based on various factors such as service requirements, the resource usage of online tasks, and the resource requirements of offline tasks to ensure the security of service operation.

[0065] Among them, the offline task graceful eviction strategy refers to the process of planned and controlled termination or migration of currently running offline tasks in a shared GPU cluster environment when the time interval between the current time and the end time of the co-allocation validity period is less than or equal to the time interval threshold and the demand for GPU resources by online tasks surges. This strategy ensures that the expected online GPU resources occupied by offline tasks are immediately released at the end time point of the co-allocation validity period to ensure the security of the online task service quality. For example, the time interval threshold can be a pre-set time range used to determine when to start the graceful eviction of offline tasks. This time interval threshold can be set according to business requirements, the tightness of GPU resources, and the sensitivity of online tasks to resources.

[0066] The first co-allocation component can also, based on the amount of idle GPU resources in the first GPU resources of online tasks under the co-allocation management data mining target service, perform online job service quality inspection and interference detection, and adjust the offline task scheduling strategy, etc. Among them, adjusting the offline task scheduling strategy can include: dynamically scheduling in real time the second GPU resources for offline tasks under the target service, distributing or recycling the second GPU resources used by offline tasks under the target service; after allocating the first GPU resources for online tasks and the second GPU resources for offline tasks under the target service, distributing the remaining GPU resources in the offline co-allocation of the target service to offline tasks under other services to more deeply exploit the cluster GPU resources and improve the GPU utilization rate of the service and the cluster as a whole.

[0067] Among them, the underlying GPU cluster 30 is configured with a second co-allocation component, which is used to manage each GPU node in the shared GPU cluster. For example, each GPU node in the shared GPU cluster can be configured with this second co-allocation component. This second co-allocation component can be used to perform quality of service (QoS) control on each resource dimension of offline resources at the control group (cgroup) level. This QoS control can include: controlling the CPU weight of offline tasks, controlling the network I / O priority of offline tasks, etc., to ensure that after the conflict between online and offline resources, the resource usage of offline tasks yields to online tasks to ensure the service quality of online tasks.

[0068] For example, the online inference platform 3 can be connected to this three-layer cluster architecture 1. For example, the online inference platform 3 is connected to the intermediate co-allocation control layer 20 and the underlying GPU cluster 30, and through the collaborative work of the intermediate co-allocation control layer 20 and the underlying GPU cluster 30, efficient processing of online tasks is achieved.

[0069] For example, the underlying GPU cluster 30 is multiple k8s clusters and is part of the shared GPU cluster. Each node in the underlying GPU cluster 30 (such as node-1, node-N) is used to deploy the online GPU nodes and offline GPU nodes for each service. For example, the online GPU node under node-1 is the first GPU resource for the online tasks under service-1; the offline GPU node under node-1 is the second GPU resource for the offline tasks under service-1.

[0070] For example, the quota management platform 4 and the monitoring platform 5 can be connected to the first hybrid deployment component configured in the intermediate hybrid deployment control layer 20. The quota management platform 4 is responsible for managing and allocating the resource quotas of each service in the shared GPU cluster, which includes but is not limited to GPU resources, memory resources, storage resources, etc. By connecting the quota management platform 4 to the first hybrid deployment component, dynamic management and real-time adjustment of resource quotas can be achieved.

[0071] The monitoring platform 5 is responsible for real-time detection, performance analysis and feedback of each component and job in the shared GPU cluster, as well as visual display of resource usage.

[0072] For example, the three-layer cluster architecture 1 can be deployed in a computer device to provide efficient GPU resource sharing and management capabilities.

[0073] The resource management method provided by the embodiments of this application is based on the three-layer cluster architecture 1. By real-time detecting the GPU resource amount of the target service and the GPU resource amount of the online tasks under the target service, it calculates in real time the offline GPU resources that can be mined by the target service (that is, calculates the idle GPU resource amount in the first GPU resource of the online tasks, and this idle GPU resource amount is used for the offline tasks under the target service). In the case of a non-intrusive k8s system, it real-time controls the dynamic scheduling and allocation of the offline tasks of the target service, enabling the GPU resources of the target service to effectively rotate between the online tasks and offline tasks in the target service, achieving a more fine-grained allocation of offline GPU resources; it real-time detects the online service quality of the online tasks, and when the online service quality deteriorates, it automatically and timely reclaims the GPU resources occupied by the offline tasks to ensure the stability and performance of the target service. The embodiments of this application can also provide a flexible resource isolation and sharing mechanism, deeply mining the remaining GPU resources of the target service. On the premise of ensuring the online service quality, it takes into account the effective and fair allocation of the GPU resources of the offline tasks, improves the GPU resource utilization rate of the service itself, enables the offline tasks under the target service to not need to purchase other offline GPU resources at the system level additionally, avoids waste of service GPU resources, and reduces the service operation cost.

[0074] An embodiment of the present application provides a resource management method. This method can be executed by a terminal or a server, or jointly executed by a terminal and a server. For the purpose of illustration, the embodiment of the present application takes the case where the resource management method is executed by the server.

[0075] Please refer to Figures 2 to 3 , Figure 2 and Figure 3 which is a schematic flowchart of the resource management method provided by the embodiment of the present application. The method may include the following steps:

[0076] Step 110, real-time detect the hybrid management data of the target service in the shared Graphics Processing Unit (GPU) cluster. The hybrid management data includes the usage information of the first GPU resources of the online tasks under the target service and the resource demand information of the offline tasks under the target service.

[0077] This step 110 involves real-time detecting the hybrid management data of a specific service (target service) in the shared GPU cluster. The hybrid management data at least includes two key pieces of information: one is the detailed information of the GPU resources (the first GPU resources) currently being used by the online tasks under the target service; the other is the GPU resource information required by the offline tasks under the target service. This step 110 obtains the latest resource usage and demand, providing basic data for subsequent resource allocation.

[0078] For example, a polling mechanism can be adopted to collect data from each GPU node or on the online / offline hybrid at regular intervals (such as every 1 minute). Alternatively, when a change in the GPU resource usage or the offline task status is detected, data collection is triggered.

[0079] For example, the usage information of the first GPU resources of the online tasks may include the GPU utilization rate, memory occupancy, the number of currently running processes, etc. of the offline tasks. For example, the resource demand information of the offline tasks may include the expected number of GPUs for the offline tasks, the estimated running time, the required computing power, etc.

[0080] Step 120, based on the hybrid management data, calculate the amount of idle GPU resources in the first GPU resources of the online tasks.

[0081] For example, after obtaining the hybrid management data, the amount of GPU resources not fully utilized by the current online tasks, that is, the amount of idle GPU resources, can be calculated according to the usage information of the first GPU resources of the online tasks. This can be achieved by comparing the actual usage of the online tasks with the total amount of GPU resources allocated to them. This step 120 can determine how much GPU resources can be reallocated to the offline tasks.

[0082] Step 130: Based on the amount of idle GPU resources and the resource requirement information of offline tasks under the target service, adjust the offline task scheduling policy, where the offline task scheduling policy indicates the second GPU resources for offline tasks under the target service.

[0083] Based on the amount of idle GPU resources obtained in the above steps and the resource requirement information of offline tasks under the target service, adjust the scheduling policy of offline tasks. Specifically, it determines which offline tasks can be assigned to which GPU resources (the second GPU resources) to meet their computing requirements while maximizing the utilization of idle GPU resources. This step 130 improves the overall utilization rate of the GPU cluster and the task execution efficiency by dynamically adjusting resource allocation.

[0084] For example, different priorities can be set for offline tasks according to their urgency and resource requirements. Offline tasks with higher priorities can obtain GPU resources first.

[0085] For example, the number of GPUs assigned to offline tasks can be dynamically adjusted according to their estimated running time and required computing power.

[0086] For example, during the execution of offline tasks, the intermediate hybrid management layer can continuously detect the usage of GPU resources. If it is found that a certain offline task occupies too much GPU resources and affects online tasks, its resource allocation should be adjusted in a timely manner or its execution should be suspended to ensure the stability and performance of online tasks.

[0087] In some embodiments, the shared GPU cluster is configured in a three-layer cluster architecture, and the three-layer cluster architecture includes:

[0088] The upper-layer offline hybrid cluster is used to receive offline tasks of each service;

[0089] The intermediate hybrid management layer is configured with a first hybrid component, and the first hybrid component is used to detect hybrid management data, calculate the amount of idle GPU resources, and adjust the offline task scheduling policy;

[0090] The lower-layer GPU cluster is configured with a second hybrid component, and the second hybrid component is used to manage each GPU node in the shared GPU cluster.

[0091] For example, based on Figure 1 The three-layer cluster architecture 1 shown, the upper-layer offline hybrid cluster 10, the intermediate hybrid management layer 20, and the lower-layer GPU cluster 30 are described respectively.

[0092] For example, the upper-layer offline hybrid scheduling cluster 10 (such as a single k8s cluster) serves as the entry point of the entire three-layer cluster architecture 1 and is responsible for receiving offline training tasks from different services (such as Service-1, Service-2, etc.). These tasks are submitted through the offline training platform 2 and exist in the form of task instances (pods). By centrally managing offline tasks through the upper-layer offline hybrid scheduling cluster 10, resource scheduling and priority allocation can be more conveniently performed, while ensuring that offline tasks do not affect the execution of online tasks in the underlying GPU cluster.

[0093] For example, the middle hybrid scheduling control layer 20 is configured with a first hybrid scheduling component, which is responsible for the hybrid management of the shared GPU cluster for online and offline tasks. It mainly includes the following functions:

[0094] Detect hybrid management data: Real-time detect and collect the hybrid management data of each service, including the first GPU resource usage of online tasks and the resource requirement information of offline tasks.

[0095] Calculate the amount of idle GPU resources: Based on the collected hybrid management data, dynamically calculate the amount of underutilized GPU resources in online tasks, that is, the amount of idle GPU resources.

[0096] Adjust the offline task scheduling strategy: According to the amount of idle GPU resources and the resource requirement information of offline tasks, dynamically adjust the scheduling strategy of offline tasks. It can include deciding when, how, and to which offline services to allocate additional GPU resources, or how to reclaim the GPU resources occupied by offline tasks to ensure the efficient execution of offline tasks without affecting the stability and performance of online tasks.

[0097] Online and offline hybrid CRD: Maintain a custom resource definition (CRD) for each service to record the hybrid configuration information, hybrid status, etc. of the service, which is convenient for the first hybrid scheduling component to manage and schedule. For example, the hybrid configuration information of the service contains some basic information about service hybrid, such as recording the basic configuration information related to the online and offline hybrid of the service, including but not limited to: the hybrid validity period of the online and offline hybrid (the start time and end time of the hybrid validity period), the graceful eviction policy for offline tasks, the target hybrid cluster corresponding to the service, the maximum hybrid resource limit, resource reservation conditions, etc. By recording the hybrid configuration information in the form of CRD of the k8s cluster, the real-time hybrid status of the service will also be recorded in this CRD in real time.

[0098] Resource isolation and sharing mechanism: Provide a flexible resource isolation and sharing mechanism to ensure that online tasks are given priority while also fairly and effectively allocating the GPU resources required for offline tasks.

[0099] Access Quota Management Platform 4 and Monitoring Platform 5: Through the access quota management platform 4, dynamic management and real-time adjustment of resource quotas are achieved; through the access monitoring platform 5, the cluster performance is detected in real time, and problems are discovered and solved in a timely manner.

[0100] For example, the underlying GPU cluster 30 is configured with a second hybrid component, and the second hybrid component is used to manage each GPU node in the shared GPU cluster. It can mainly include the following functions:

[0101] Quality of Service (QoS) Control: At the control group (cgroup) level, QoS control is performed on the CPU weight, network I / O priority, etc. of offline tasks to ensure that offline tasks can actively yield resources to online tasks in case of resource conflicts.

[0102] GPU Node Management: Responsible for deploying and managing the online GPU nodes and offline GPU nodes of each service to ensure their efficient and stable operation.

[0103] In some embodiments, the hybrid management data further includes the hybrid validity period and hybrid status of the online-offline hybrid; based on the hybrid management data, the amount of idle GPU resources in the first GPU resources of the online task is calculated, including: regularly detecting whether the current time is within the hybrid validity period; if the current time is within the hybrid validity period, the offline tasks under the target service are placed in the hybrid processing queue; based on the hybrid management data, the amount of idle GPU resources in the first GPU resources of the online task is calculated.

[0104] In some embodiments, the method further includes: if the current time is not within the hybrid validity period, detecting whether there are offline tasks under the target service in the underlying GPU cluster; if there are offline tasks under the target service in the underlying GPU cluster, evicting the offline tasks under the target service from the underlying GPU cluster and setting the hybrid status to non-hybrid status.

[0105] In some embodiments, the method further includes: when evicting the offline tasks under the target service from the underlying GPU cluster, sending an alarm message, which is used to prompt the offline task eviction event.

[0106] For example, in the three-layer cluster architecture 1, the hybrid management data not only includes the usage information of the first GPU resources of the online task and the resource requirements information of the offline task, but also includes the hybrid validity period and hybrid status of the online-offline hybrid.

[0107] For example, the hybrid validity period defines the valid time range of the hybrid operation. Only within the validity period will the hybrid operation be executed, including calculating the amount of idle GPU resources, adjusting the offline task scheduling strategy, etc.

[0108] For example, the co-allocation status reflects the status of the current co-allocation operation in offline co-allocation, such as the Inactive status, Evict status, and PoorRunning status (resource-constrained co-allocation status). This co-allocation status helps the system detect the execution of co-allocation operations and intervene when necessary.

[0109] The Inactive status means that during the Inactive validity period, offline tasks need to be taken offline, and the shared GPU cluster does not allow offline tasks to run in offline co-allocation.

[0110] The Evict status means the status when all offline tasks need to be evicted during the co-allocation validity period.

[0111] The PoorRunning status (resource-constrained co-allocation status) means the status when some offline tasks need to be evicted during the co-allocation validity period. For example, if the total demand of the offline service is 100 GPUs, but the current available idle GPU resources for offline tasks is only 80 GPUs, the offline tasks will also be dispatched based on the offline task adjustment policy, but the resource requirements of the offline tasks are not fully met. Therefore, the co-allocation status is the PoorRunning status.

[0112] For example, the first co-allocation component of the middle-layer co-allocation control layer 20 will continuously sense the GPU in offline co-allocation CRD configurations of each service and detect relevant information such as the creation, modification, and deletion operations of each service configuration to obtain co-allocation management data. These co-allocation management data will be maintained in the cache for subsequent quick access.

[0113] Regularly pull the co-allocation management data from the cache and periodically detect whether the current time is within the co-allocation validity period. This step ensures that co-allocation processing tasks are correctly processed within the validity period, avoiding resource allocation errors caused by time expiration. The co-allocation processing tasks are online tasks and offline tasks under the target task.

[0114] If the current time is within the co-allocation validity period, the first co-allocation component of the middle-layer co-allocation control layer 20 will put the offline tasks of the target service into the co-allocation processing queue. Subsequently, another thread will continuously detect the resource quota situation, offline resource situation, and online resource situation of the target service, etc., in order to comprehensively calculate the available idle GPU resources that can be exploited, that is, the idle GPU resources in the first GPU resources of the online tasks that are not fully utilized. This idle GPU resource volume can then be used to schedule offline tasks to improve the overall utilization rate of GPU resources.

[0115] If the current time is not within the co-allocation validity period, the first co-allocation component of the intermediate co-allocation management layer 20 will re-detect whether there are offline tasks under the target service in the underlying GPU cluster. If there are, an emergency eviction operation will be immediately performed to ensure that the service quality of online tasks is not affected.

[0116] When evicting offline tasks, an alarm message will be sent to prompt the offline task eviction event. This step enhances the maintainability and fault troubleshooting ability of the system, enabling operation and maintenance personnel to promptly discover and handle abnormal situations. After the eviction is completed, the co-allocation status will be set to the non-co-allocation (Inactive) state to indicate that the current service is no longer performing co-allocation processing.

[0117] The embodiments of the present application are not only applicable to the co-allocation processing of a single service, but also capable of supporting the co-allocation processing of multiple services simultaneously. By maintaining the co-allocation management data and status information of each service, unified management and efficient scheduling of multiple service resources are achieved. In addition, as the business develops and requirements change, new co-allocation configuration fields can be added, resource calculation algorithms can be optimized, or alarm strategies can be adjusted, etc., to meet the requirements in different scenarios.

[0118] In some embodiments, the co-allocation management data further includes the service quality indicators and detection rules of online tasks;

[0119] Based on the amount of idle GPU resources and the resource requirement information of offline tasks, adjust the offline task scheduling strategy, including:

[0120] Based on the service quality indicators and detection rules of online tasks, and the usage information of the first GPU resource, real-time detect the online service quality of online tasks;

[0121] According to the amount of idle GPU resources and the online service quality, decide that the offline task scheduling strategy is to issue offline tasks or evict offline tasks;

[0122] If the decision on the offline task scheduling strategy is to issue offline tasks, then based on the amount of idle GPU resources and the resource requirement information of offline tasks under the target service, determine the second GPU resource for the offline tasks under the target service to use, and issue the offline tasks under the target service to the second GPU resource; or

[0123] If the decision on the offline task scheduling strategy is to evict offline tasks, then partially or fully recycle the idle GPU resources occupied by the offline tasks, and re-allocate the recycled idle GPU resources to online tasks.

[0124] For example, the hybrid management data not only includes the basic configuration information of offline tasks and online tasks, but also includes the quality of service indicators of online tasks (such as response time, throughput, error rate, etc.) and detection rules. These indicators and rules are important bases for real-time detection of the health status of online tasks.

[0125] For example, based on the quality of service indicators and detection rules of online tasks in the hybrid management data, as well as the usage information of the first GPU resource, the online service quality of online tasks is detected in real time. This step is the key to ensuring the stable operation of online tasks. By promptly discovering problems and taking corresponding measures, the decline in service quality can be effectively avoided.

[0126] After obtaining the amount of idle GPU resources and the evaluation result of online service quality, the system will comprehensively consider these two factors to determine the offline task scheduling strategy. If the online service quality is good and the amount of idle GPU resources is sufficient, it will be decided to issue offline tasks; if the online service quality is poor or the amount of idle GPU resources is insufficient, it will be decided to partially evict offline tasks (i.e., compress the amount of offline GPU resources) or fully evict offline tasks (i.e., evict the amount of offline GPU resources) to release resources for online tasks.

[0127] If the decision is to issue offline tasks, based on the amount of idle GPU resources and the resource requirement information of offline tasks under the target service, the second GPU resource for offline tasks under the target service will be determined, and offline tasks will be issued to this resource. This step aims to make full use of GPU resources and improve resource utilization.

[0128] If the decision is to evict offline tasks, the idle GPU resources occupied by offline tasks will be partially or fully recycled, and the recycled resources will be reallocated to online tasks. This step aims to ensure the service quality of online tasks and avoid performance degradation or interruption caused by insufficient resources.

[0129] In some embodiments, if the decision on the offline task scheduling strategy is to issue offline tasks, then based on the amount of idle GPU resources and the resource requirement information of offline tasks under the target service, the second GPU resource for offline tasks under the target service is determined, and the offline tasks under the target service are issued to the second GPU resource, including:

[0130] If the decision on the offline task scheduling strategy is to issue offline tasks, then through the first hybrid component of the intermediate hybrid management layer, based on the amount of idle GPU resources and the resource requirement information of offline tasks under the target service, the second GPU resource for offline tasks under the target service is determined;

[0131] The first offline task scheduling strategy corresponding to the offline tasks under the target service is issued to the upper-layer offline hybrid cluster through the first hybrid component, and the offline tasks under the target service are allocated to the GPU nodes corresponding to the second GPU resource in the underlying GPU cluster.

[0132] For example, when it is decided to issue an offline task based on the amount of idle GPU resources, the online service quality, and the resource requirement information of the offline task, the next resource allocation process will be entered.

[0133] Through the first hybrid component of the intermediate hybrid management and control layer 20, the required amount of GPU resources will be accurately calculated based on the current amount of idle GPU resources and the resource requirement information of the offline task under the target service, and the second GPU resources for the offline task under the target service will be determined. This step ensures the reasonable allocation of resources and avoids the situation of resource waste or shortage.

[0134] After determining the second GPU resources, the first hybrid component will generate a first offline task scheduling policy for the offline task under the target service. Then, the first hybrid component will send the first offline task scheduling policy to the upper-layer offline hybrid cluster 10. After receiving the first offline task scheduling policy, the upper-layer offline hybrid cluster 10 will forward the offline task under the target service to the underlying GPU cluster 30. In the underlying GPU cluster 30, the offline task under the target service will be allocated to the GPU nodes corresponding to the second GPU resources according to the first offline task scheduling policy.

[0135] In some embodiments, if the decision on the offline task scheduling policy is to issue an offline task, the method further includes:

[0136] After allocating the second GPU resources, detecting whether there are still remaining GPU resources in the amount of idle GPU resources;

[0137] If there are still remaining GPU resources in the amount of idle GPU resources, the second offline task scheduling policy corresponding to the offline task under other services will be sent to the upper-layer offline hybrid cluster through the first hybrid component, and the offline task under other services will be allocated to the GPU nodes corresponding to the remaining GPU resources in the underlying GPU cluster;

[0138] Among them, the priority of the offline task under other services using the amount of idle GPU resources is lower than the priority of the offline task under the target service using the amount of idle GPU resources.

[0139] For example, when it is decided to issue an offline task and the second GPU resources for the offline task under the target service are determined based on the amount of idle GPU resources and the resource requirement information of the offline task under the target service, a preliminary allocation of resources will be carried out. After allocating the second GPU resources, the resource allocation process will not be stopped immediately, but further detection will be carried out to determine whether there are still remaining GPU resources in the amount of idle GPU resources. This step is the key to ensuring the full utilization of GPU resources.

[0140] If it is detected that there are still remaining GPU resources, the first hybrid component of the intermediate hybrid control layer 20 will send the second offline task scheduling policy corresponding to the offline tasks under other services to the upper-layer offline hybrid cluster 10. This policy will specify how the offline tasks of other services should use these remaining GPU resources.

[0141] After receiving the second offline task scheduling policy, the upper-layer offline hybrid cluster 10 will forward the offline tasks under other services to the underlying GPU cluster 30. In the underlying GPU cluster 30, the second hybrid component will allocate the offline tasks under other services to the GPU nodes corresponding to the remaining GPU resources according to the policy. During the allocation process, it will be clearly set that the priority of the offline tasks under other services using the idle GPU resources is lower than that of the offline tasks under the target service. This means that when allocating resources, the needs of the offline tasks under the target service will be satisfied first, and then the offline tasks of other services will be considered.

[0142] In some embodiments, if the decision-making offline task scheduling policy is to evict offline tasks, some or all of the idle GPU resources occupied by the offline tasks will be recycled, and the recycled idle GPU resources will be reallocated to online tasks, including:

[0143] If the decision-making offline task scheduling policy is to evict offline tasks, the offline tasks will be evicted in the preset priority order, and the preset priority order is: first evict the offline tasks under other services, and then evict the offline tasks under the target service;

[0144] The recycled idle GPU resources will be reallocated to online tasks;

[0145] Among them, if the eviction volume of evicting offline tasks is full eviction, the hybrid state will be set to the eviction state; or if the eviction volume of evicting offline tasks is partial eviction, the hybrid state will be set to the resource-constrained hybrid state.

[0146] In some embodiments, evicting the offline tasks under other services includes: recycling some or all of the remaining GPU resources occupied by the offline tasks under other services, and reallocating the recycled remaining GPU resources to online tasks;

[0147] Evicting the offline tasks under the target service includes: recycling some or all of the second GPU resources occupied by the offline tasks under the target service, and reallocating the recycled second GPU resources to online tasks.

[0148] For example, offline tasks are evicted in the order of preset priorities. Offline tasks under other services are evicted first, and then those under the target service are evicted. This order ensures that the resource requirements of the target service are prioritized when resources are scarce. Then, the GPU resources occupied by the evicted offline tasks are reclaimed and redistributed to online tasks to ensure the service quality of the online tasks.

[0149] For example, if the eviction volume of offline tasks is full eviction, the hybrid deployment state is set to the eviction state (Evict). This means that all offline tasks are completely evicted, and the resources are completely reclaimed and redistributed to online tasks.

[0150] For example, if the eviction volume of offline tasks is partial eviction, the hybrid deployment state is set to the resource-constrained hybrid deployment state (PoorRunning). This means that some offline tasks are evicted, some resources are reclaimed and redistributed to online tasks, and the system is in a state of resource tension but still running.

[0151] Among them, the process of evicting offline tasks under other services includes: partially or fully reclaiming the remaining GPU resources occupied by offline tasks under other services. The reclaimed remaining GPU resources are redistributed to online tasks to ensure the service quality of the online tasks.

[0152] Among them, the process of evicting offline tasks under the target service includes: partially or fully reclaiming the second GPU resources occupied by offline tasks under the target service. The reclaimed second GPU resources are redistributed to online tasks to ensure the service quality of the online tasks.

[0153] For example, when implementing the eviction of offline tasks, in order to ensure the online service quality, the minimum number of GPU cards that need to be released is calculated. For example, the resource requirements of the online service, the service quality threshold, and the resource occupancy of offline tasks can be comprehensively considered to calculate the minimum number of GPU cards that need to be released. Then, the corresponding offline tasks are selected for eviction according to the calculated minimum number of GPU cards that need to be released. During the eviction process, it is necessary to ensure that the execution progress of the offline tasks is properly saved and can be resumed when resources are sufficient in the future.

[0154] In some embodiments, before deciding to issue a new offline task, it is checked whether there is a certain time interval between the last issued offline task or evicted offline task and the current operation. This is to avoid the "resource glitch" phenomenon caused by frequent resource adjustments, that is, the resource usage fluctuates violently in a short period of time. In addition, in addition to the time interval, it is also checked whether at least one dimension of the data in each dimension (such as the amount of GPU resources that can be mined, the online resources, and the offline resources) has changed. These changes can include the release of new GPU resources, changes in the resource requirements of online tasks, or adjustments to the resource occupancy of offline tasks. Only when at least one dimension of the data changes will the expansion of offline resources be considered to avoid unnecessary resource adjustments.

[0155] For example, when it is determined that the offline GPU jobs of the target service can be expanded, and the current amount of idle GPU resources is sufficient to support the new offline task, and at the same time it will not seriously affect the online service quality, the system will decide to issue a new offline task. This helps to make full use of idle resources and improve resource utilization.

[0156] For example, when it is determined that the current amount of idle GPU resources is insufficient, or issuing a new offline task will have a greater impact on the online service quality, the system will decide to evict some or all of the existing offline tasks. This can ensure that online tasks obtain sufficient resources and maintain their service quality.

[0157] In some embodiments, the method further includes:

[0158] Real-time detect whether the current time is the end time of the co-location validity period;

[0159] If the current time is not the end time of the co-location validity period, then repeat the adjustment of the offline task scheduling policy; or

[0160] If the current time is the end time of the co-location validity period, then prohibit issuing offline tasks.

[0161] For example, real-time detect the relationship between the current time and the end time of the co-location validity period. The co-location validity period refers to a preset time period during which online and offline tasks are allowed to share GPU resources. Then, based on the real-time detection result, it is judged whether the current time has reached the end time of the co-location validity period.

[0162] If the current time is not the end time of the co-location validity period, the previous scheduling policy will continue to be executed, that is, according to the amount of idle GPU resources and the online service quality, dynamically adjust the scheduling of offline tasks (including issuing new offline tasks or evicting existing offline tasks) to ensure that resources are efficiently utilized within the co-location validity period.

[0163] If the current time reaches the end time of the co-allocation validity period, specific measures will be taken, namely, prohibiting the issuance of new offline tasks. This is to ensure that after the end of the co-allocation validity period, the GPU resources can be promptly returned to the online tasks, avoiding any impact on the stable operation of the online services. Additionally, existing offline GPU jobs can be evicted in advance to ensure the timely release and return of the co-allocation resources. After the co-allocation resources are returned (i.e., all offline tasks have been evicted and the GPU resources have been fully released to the online tasks), the co-allocation status will be set to "Inactive". This indicates that there are no active co-allocation operations currently, and the GPU resources are fully utilized by the online tasks.

[0164] After that, the entire process starting from step 110 will be repeatedly executed to continuously detect and adjust the allocation and usage of GPU resources, ensuring the effective utilization of resources and the stable operation of tasks.

[0165] In some embodiments, the method further includes:

[0166] Through the second co-allocation component in the underlying GPU cluster, perform quality of service (QoS) control on the resources of each offline task at the control group level. The QoS control includes adjusting the central processing unit (CPU) weight and network input / output priority of each offline task. Each offline task includes at least one of the online tasks under the target service and the online tasks under other services;

[0167] According to the results of the QoS control, readjust the offline task scheduling strategy.

[0168] Among them, to ensure that the service quality of online tasks is not affected by offline tasks and at the same time improve the overall utilization rate of GPU resources, the second co-allocation component in the underlying GPU cluster can perform quality of service (QoS) control on the resources of each offline task at the control group (cgroup) level. This control covers multiple resource dimensions, including the adjustment of the central processing unit (CPU) weight and network input / output (I / O) priority, etc. This control mechanism helps to ensure that offline tasks can yield resources for online tasks in case of resource conflicts, thus guaranteeing the service quality of online tasks.

[0169] In some embodiments, perform QoS control on the resources of offline tasks at the cgroup level. Specifically, it includes: setting the CPU weight of offline tasks lower than that of online tasks to limit the occupancy of CPU resources by offline tasks; setting the network I / O priority of offline tasks lower than that of online tasks to reduce the competition for network resources by offline tasks; among them, the QoS control strategy is adjusted according to the resource requirements of online tasks.

[0170] For example, to limit the CPU resource occupancy of offline tasks, the second hybrid scheduling component will set the CPU weight of offline tasks lower than that of online tasks, which means that in case of resource contention, online tasks will obtain more CPU time slices. Similarly, the second hybrid scheduling component will also set the network I / O priority of offline tasks lower than that of online tasks to reduce the competition of offline tasks for network resources and ensure that the network performance of online tasks is not affected.

[0171] Among them, the QoS control policy is not fixed, but is adjusted according to the resource requirements of online tasks. For example, when online tasks require more resources, the second hybrid scheduling component can further reduce the CPU weight and network I / O priority of offline tasks to ensure that online tasks obtain sufficient resources.

[0172] In some embodiments, the method further includes: when there is a resource conflict, detecting the usage information of the first GPU resources of online tasks, where the usage information includes the resource usage ratio; if the resource usage ratio of online tasks exceeds a preset threshold, dynamically reducing the CPU weight and network I / O priority of offline tasks; if the resource usage ratio of online tasks is lower than the preset threshold, allowing the CPU weight and network I / O priority of offline tasks to return to the default value.

[0173] When there is a resource conflict, the second hybrid scheduling component will detect the usage information of the first GPU resources (which may include resources such as CPU, memory, and network) of online tasks, especially the resource usage ratio.

[0174] If the resource usage ratio of online tasks exceeds a preset threshold, indicating that the current resources of online tasks are tense, the second hybrid scheduling component will dynamically reduce the CPU weight and network I / O priority of offline tasks to release more resources for online tasks.

[0175] If the resource usage ratio of online tasks is lower than the preset threshold, it indicates that the current resources are relatively abundant, and the second hybrid scheduling component allows the CPU weight and network I / O priority of offline tasks to return to the default value to balance resource utilization and task performance.

[0176] In some embodiments, the QoS control further includes: restricting the memory usage of offline tasks to prevent offline tasks from occupying too much memory resources and affecting the operation of online tasks; when detecting that the memory resources are tense, preferentially releasing the memory resources occupied by offline tasks to ensure the memory requirements of online tasks.

[0177] For example, to prevent offline tasks from occupying too much memory resources and affecting the operation of online tasks, the second hybrid scheduling component will restrict the memory usage of offline tasks. This helps to ensure that online tasks can still operate normally when the memory resources are tense.

[0178] When detecting that the memory resources are tight, the second hybrid component will preferentially release the memory resources occupied by offline tasks to ensure the memory requirements of online tasks. This mechanism helps to ensure the stable operation of critical tasks in the case of limited memory resources.

[0179] In some embodiments, the method further includes: recording the historical data of resource usage of offline tasks in different resource dimensions; optimizing the QoS control policy according to the historical data of resource usage to improve the efficiency and fairness of resource allocation.

[0180] For example, the second hybrid component will record the historical data of resource usage of offline tasks in different resource dimensions. These historical data help to analyze the resource usage patterns and behavioral characteristics of offline tasks.

[0181] Based on the historical data of resource usage, the second hybrid component can optimize the QoS control policy to improve the efficiency and fairness of resource allocation. For example, by adjusting the CPU weights of offline tasks and the settings of network I / O priorities, the resource allocation can be made more in line with the actual task requirements to improve the overall system performance.

[0182] To better illustrate the resource management method provided by the embodiments of the present application, please refer to Figure 3 , the process of the resource management method provided by the embodiments of the present application can be summarized as the following steps:

[0183] S1, Detect the hybrid management data of the target service in the shared Graphics Processing Unit (GPU) cluster in real time.

[0184] S2, Maintain the hybrid management data in the cache.

[0185] S3, Determine whether the current time is within the hybrid validity period; if not, execute step S4; if so, execute step S7.

[0186] S4, If the current time is not within the hybrid validity period, detect whether there are offline tasks under the target service in the underlying GPU cluster; if so, execute step S5; if not, return to execute step S1.

[0187] S5, If there are offline tasks under the target service in the underlying GPU cluster, evict the offline tasks under the target service from the underlying GPU cluster.

[0188] S6, Set the hybrid status to the Inactive state. For example, the setting format is "status:phase:Inactive". Then return to execute step S2 and maintain this hybrid status in the cache.

[0189] S7, if the current time is within the co-allocation validity period, put the offline tasks under the target service into the co-allocation processing queue.

[0190] S8, calculate the amount of idle GPU resources in the first GPU resources of the online tasks based on the co-allocation management data. For example, the amount of idle GPU resources that can be mined can be calculated based on the co-allocation management data, the online resource controller, and the offline resource controller.

[0191] S9, based on the quality of service indicators and detection rules of the online tasks, as well as the usage information of the first GPU resources, detect the online service quality of the online tasks in real time.

[0192] S10, based on the amount of idle GPU resources and the online service quality, decide whether the offline task scheduling policy is to issue offline tasks or evict offline tasks.

[0193] S11, if the decision of the offline task scheduling policy is to issue offline tasks, determine the second GPU resources for the offline tasks under the target service through the first co-allocation component of the intermediate co-allocation control layer based on the amount of idle GPU resources and the resource requirement information of the offline tasks under the target service.

[0194] S12, through the first co-allocation component, issue the first offline task scheduling policy corresponding to the offline tasks under the target service to the upper-layer offline co-allocation cluster, and allocate the offline tasks under the target service to the GPU nodes corresponding to the second GPU resources in the underlying GPU cluster.

[0195] S13, after allocating the second GPU resources, detect whether there are still remaining GPU resources in the amount of idle GPU resources; if so, execute step S14; if not, return to execute step S10.

[0196] S14, if there are still remaining GPU resources in the amount of idle GPU resources, issue the second offline task scheduling policy corresponding to the offline tasks under other services to the upper-layer offline co-allocation cluster through the first co-allocation component, and allocate the offline tasks under other services to the GPU nodes corresponding to the remaining GPU resources in the underlying GPU cluster.

[0197] S15. If the decision on the offline task scheduling policy is to evict offline tasks, then partially or fully reclaim the idle GPU resources occupied by the offline tasks, and reallocate the reclaimed idle GPU resources to the online tasks. Among them, if the decision on the offline task scheduling policy is to evict offline tasks, then evict the offline tasks in the order of the preset priority. The preset priority order is: first evict the offline tasks under other services, and then evict the offline tasks under the target service; reallocate the reclaimed idle GPU resources to the online tasks; among them, if the eviction volume of evicting offline tasks is full eviction, then set the hybrid state to the eviction state (Evict state); or if the eviction volume of evicting offline tasks is partial eviction, then set the hybrid state to the resource-constrained hybrid state (PoorRunning state).

[0198] S16. Real-time detect whether the current time is the end time of the hybrid validity period; if so, execute step S17; if not, return to execute step S7.

[0199] S17. If the current time is the end time of the hybrid validity period, then prohibit the issuance of offline tasks.

[0200] S18. Set the hybrid state to the Evict state. For example, the setting format is "status:phase:Evict".

[0201] S19. Evict the offline tasks currently running in the in-offline hybrid, and set the hybrid state to the Inactive state. For example, the setting format is "status:phase:Inactive". Then return to execute step S3.

[0202] For the specific implementation manners corresponding to each step in the above overall process in the embodiments of the present application, reference may be made to the embodiment content corresponding to the above first process schematic diagram. For the sake of brevity, it will not be elaborated here.

[0203] The application scenario of the embodiments of the present application is in a shared GPU cluster, where multiple services share the GPU single-cluster and / or multi-cluster resource pools. Each service itself has an online GPU resource quota in the resource pool of the shared GPU cluster. At the same time, the service has GPU-accelerated online tasks (such as online GPU training tasks) and online tasks (such as offline GPU training tasks). It is necessary to use the available idle GPU resource amount of the online GPU resources of the target service for the offline tasks of the target service. The offline tasks under the target service do not need to compete with other services for system-level offline GPU resources. On the premise of saving business operation costs, it can further ensure the quality of online and offline services. It requires an in-offline hybrid technology at the service level to detect the resource usage information and online service quality of the online GPU training service of the target service in real time, mine and calculate the idle GPU resources available for the offline GPU training tasks under the target service in real time, and manage and distribute the offline GPU training tasks of the same service. And perform online service interference detection and health checks in real time to realize the real-time dynamic recycling of offline GPU training tasks and release GPU resources. In addition, by combining the in-offline hybrid technology at the system level, the remaining GPU resources of the overall resources of the online and offline tasks under the target service can be further mined and utilized; deeply mine the GPU resources of the service and the cluster to improve the GPU utilization rate of the service or the cluster.

[0204] In the in-offline hybrid scenario of the shared GPU cluster in the embodiments of the present application, the GPU offline hybrid resources are more finely managed, supporting the effective and safe transfer of GPU resources under the same target service between the GPU AI online tasks and offline tasks of the target service, avoiding resource competition between offline tasks and offline tasks under other services in the system, and the service quality is interfered; while improving the resource service quality of the offline tasks under the target service, it further improves the GPU resource utilization rate of the service and saves business operation costs. The embodiments of the present application not only improve the GPU resource utilization rate of the target service, but also further mine the remaining GPU resources after the in-offline hybrid of the service-level GPU, distribute the remaining GPU resources to the offline tasks under other services, and perform fine-grained division at the offline task level, giving priority to ensuring the service quality of the online tasks under the target service, followed by ensuring the service quality of the offline tasks under the target service, and finally providing service guarantee for the offline tasks under other services. On the premise of ensuring resource fairness, further mine the GPU resources at the cluster level or service level to improve the overall GPU resource utilization rate and save the company's computing power costs.

[0205] The embodiment of the present application provides a GPU hybrid deployment solution based on a shared GPU cluster in an offline hybrid deployment scenario. In the offline hybrid deployment scenario of the shared GPU cluster, through a new and more fine-grained GPU offline hybrid deployment technology, not only can the GPU resource utilization rate of the business itself be improved, the GPU resource guarantee for the business itself in the offline state can be ensured, and the business operation cost can be saved, but also the remaining GPU resources of the business can be further explored, realizing flexible resource isolation and sharing between the business and the cluster, improving the overall GPU resource utilization rate of the cluster, and saving the company's computing power cost.

[0206] Any combination of the above technical solutions can form an optional embodiment of the present application, which will not be elaborated here one by one.

[0207] The embodiment of the present application detects the hybrid management data of the target business in the shared Graphics Processing Unit (GPU) cluster in real time. The hybrid management data includes the usage information of the first GPU resources of the online tasks under the target business and the resource requirement information of the offline tasks under the target business. Based on the hybrid management data, the amount of idle GPU resources in the first GPU resources of the online tasks is calculated. Based on the amount of idle GPU resources and the resource requirement information of the offline tasks under the target business, the offline task scheduling policy is adjusted. The offline task scheduling policy indicates the second GPU resources for the offline tasks under the target business. By detecting the hybrid management data of the target business in the shared GPU cluster in real time and accurately calculating the amount of idle GPU resources of the online tasks based on this data, the embodiment of the present application dynamically adjusts the scheduling policy of the offline tasks, effectively realizing the dynamic resource sharing between the online tasks and the offline tasks, avoiding the idle and waste of the online GPU resources, and significantly improving the GPU resource utilization rate and task execution efficiency of the target tasks.

[0208] To facilitate the better implementation of the resource management method of the embodiment of the present application, the embodiment of the present application also provides a resource management device. Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the resource management device provided by the embodiment of the present application. Among them, the resource management device 200 may include:

[0209] A detection unit 210, configured to detect the hybrid management data of the target business in the shared Graphics Processing Unit (GPU) cluster in real time. The hybrid management data includes the usage information of the first GPU resources of the online tasks under the target business and the resource requirement information of the offline tasks under the target business;

[0210] A calculation unit 220, configured to calculate the amount of idle GPU resources in the first GPU resources of the online tasks based on the hybrid management data;

[0211] An adjustment unit 220, configured to adjust an offline task scheduling policy based on the amount of idle GPU resources and the resource requirement information of the offline tasks under a target service, where the offline task scheduling policy indicates second GPU resources for the offline tasks under the target service.

[0212] In some embodiments, the shared GPU cluster is configured in a three-tier cluster architecture, and the three-tier cluster architecture includes:

[0213] An upper-layer offline hybrid deployment cluster, configured to receive offline tasks of each service;

[0214] An intermediate hybrid deployment management layer, configured with a first hybrid deployment component, where the first hybrid deployment component is configured to detect hybrid deployment management data, calculate the amount of idle GPU resources, and adjust the offline task scheduling policy;

[0215] A lower-layer GPU cluster, configured with a second hybrid deployment component, where the second hybrid deployment component is configured to manage each GPU node in the shared GPU cluster.

[0216] In some embodiments, the hybrid deployment management data further includes the hybrid deployment validity period and the hybrid deployment status in the offline hybrid deployment; a calculation unit 220, configured to: periodically detect whether the current time is within the hybrid deployment validity period; if the current time is within the hybrid deployment validity period, place the offline tasks under the target service into a hybrid deployment processing queue; and calculate the amount of idle GPU resources in the first GPU resources of the online tasks based on the hybrid deployment management data.

[0217] In some embodiments, the hybrid deployment management data further includes the quality of service indicators and detection rules of the online tasks; an adjustment unit 220, configured to: detect the online service quality of the online tasks in real time based on the quality of service indicators and detection rules of the online tasks and the usage information of the first GPU resources; make a decision on the offline task scheduling policy to issue or evict the offline tasks according to the amount of idle GPU resources and the online service quality; if the decision on the offline task scheduling policy is to issue the offline tasks, determine second GPU resources for the offline tasks under the target service based on the amount of idle GPU resources and the resource requirement information of the offline tasks under the target service, and issue the offline tasks under the target service to the second GPU resources; or if the decision on the offline task scheduling policy is to evict the offline tasks, partially or fully recycle the idle GPU resources occupied by the offline tasks and reallocate the recycled idle GPU resources to the online tasks.

[0218] In some embodiments, the adjustment unit 220 is configured to;

[0219] If the decision on the offline task scheduling policy is to issue the offline tasks, determine second GPU resources for the offline tasks under the target service based on the amount of idle GPU resources and the resource requirement information of the offline tasks under the target service through the first hybrid deployment component in the intermediate hybrid deployment management layer;

[0220] The first hybrid scheduling component is used to send the first offline task scheduling policy corresponding to the offline tasks of the target service to the upper-layer offline hybrid cluster, and allocate the offline tasks of the target service to the GPU nodes corresponding to the second GPU resources in the underlying GPU cluster.

[0221] In some embodiments, if the decision on the offline task scheduling policy is to send the offline tasks, the adjustment unit 220 is further configured to: after allocating the second GPU resources, detect whether there are any remaining GPU resources in the idle GPU resource amount; if there are remaining GPU resources in the idle GPU resource amount, send the second offline task scheduling policy corresponding to the offline tasks of other services to the upper-layer offline hybrid cluster through the first hybrid scheduling component, and allocate the offline tasks of other services to the GPU nodes corresponding to the remaining GPU resources in the underlying GPU cluster; wherein, the priority of using the idle GPU resource amount for the offline tasks of other services is lower than the priority of using the idle GPU resource amount for the offline tasks of the target service.

[0222] In some embodiments, the adjustment unit 220 is configured to: if the decision on the offline task scheduling policy is to evict the offline tasks, evict the offline tasks in the preset priority order, and the preset priority order is: first evict the offline tasks of other services, and then evict the offline tasks of the target service; re-allocate the recycled idle GPU resources to the online tasks;

[0223] Wherein, if the eviction amount of the evicted offline tasks is a full eviction, the hybrid state is set to the eviction state; or if the eviction amount of the evicted offline tasks is a partial eviction, the hybrid state is set to the resource-constrained hybrid state.

[0224] In some embodiments, the adjustment unit 220 is configured to evict the offline tasks of other services, including: partially or fully recycling the remaining GPU resources occupied by the offline tasks of other services, and re-allocating the recycled remaining GPU resources to the online tasks;

[0225] The adjustment unit 220 is configured to evict the offline tasks of the target service, including: partially or fully recycling the second GPU resources occupied by the offline tasks of the target service, and re-allocating the recycled second GPU resources to the online tasks.

[0226] In some embodiments, the adjustment unit 220 is further configured to: detect in real time whether the current time is the end time of the hybrid validity period; if the current time is not the end time of the hybrid validity period, repeat the adjustment of the offline task scheduling policy; or if the current time is the end time of the hybrid validity period, prohibit the sending of offline tasks.

[0227] In some embodiments, the adjustment unit 220 is further configured to: if the current time is not within the co-allocation validity period, detect whether there are offline tasks under the target service in the underlying GPU cluster; if there are offline tasks under the target service in the underlying GPU cluster, evict the offline tasks under the target service from the underlying GPU cluster, and set the co-allocation state to a non-co-allocation state.

[0228] In some embodiments, the adjustment unit 220 is further configured to: when evicting the offline tasks under the target service from the underlying GPU cluster, send an alarm message, which is used to prompt the offline task eviction event.

[0229] In some embodiments, the adjustment unit 220 is further configured to: through the second co-allocation component in the underlying GPU cluster, perform quality of service control on the resources of each offline task, and the quality of service control includes adjusting the central processing unit (CPU) weight and network input / output priority of each offline task, and each offline task includes at least one of the online tasks under the target service and the online tasks under other services; according to the results of the quality of service control, readjust the offline task scheduling policy.

[0230] It should be noted that the functions of the modules in the resource management device 200 in the embodiments of the present application can be correspondingly referred to the specific implementation manners of any embodiments in the above method embodiments, and will not be elaborated here.

[0231] Each unit in the above device can be implemented in whole or in part by software, hardware, and their combination. Each of the above units can be embedded in the processor of the computer device in a hardware form or be independent of it, or can be stored in the memory of the computer device in a software form, so that the processor can call and execute the operations corresponding to each of the above units.

[0232] For example, the resource management device 200 can be integrated in a terminal or server with a memory and a processor installed and having computing capabilities, or the resource management device 200 is the terminal or server.

[0233] In some embodiments, the present application further provides a computer device, including a memory and a processor, and a computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0234] Figure 5 For the schematic structural diagram of the computer device provided by the embodiments of the present application, as Figure 5As shown in the figure, the computer device 300 may include: a communication interface 301, a memory 302, a processor 303, and a communication bus 304. The communication interface 301, the memory 302, and the processor 303 communicate with each other through the communication bus 304. The communication interface 301 is used for the device 300 to communicate with external devices. The memory 302 can be used to store software programs and modules. The processor 303 runs the software programs and modules stored in the memory 302, such as the software programs for the corresponding operations in the foregoing method embodiments.

[0235] In some embodiments, the processor 303 may call the software programs and modules stored in the memory 302 to perform the following operations: real-time detect the hybrid management data of the target service in the shared Graphics Processing Unit (GPU) cluster, where the hybrid management data includes the usage information of the first GPU resources of the online tasks under the target service and the resource demand information of the offline tasks under the target service; based on the hybrid management data, calculate the amount of idle GPU resources in the first GPU resources of the online tasks; based on the amount of idle GPU resources and the resource demand information of the offline tasks under the target service, adjust the offline task scheduling policy, where the offline task scheduling policy indicates the second GPU resources for the offline tasks under the target service to use.

[0236] In some embodiments, the computer device 300 may be integrated, for example, in a terminal or a server that has a storage device and is equipped with a processor and has computing capabilities, or the computer device 300 is the terminal or the server.

[0237] The present application also provides a computer-readable storage medium for storing a computer program. The computer-readable storage medium can be applied to the computer device, and the computer program enables the computer device to execute the corresponding processes in the foregoing various methods in the embodiments of the present application. For the sake of brevity, details are not described herein again.

[0238] The present application also provides a computer program product. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the corresponding processes in the foregoing various methods in the embodiments of the present application. For the sake of brevity, details are not described herein again.

[0239] The present application also provides a computer program. The computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the corresponding processes in the foregoing various methods in the embodiments of the present application. For the sake of brevity, details are not described herein again.

[0240] It should be understood that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0241] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.

[0242] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0243] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0244] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0245] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0246] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0247] In addition, the functional units in the embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0248] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer or a server) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0249] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A resource management method, characterized in that, The method includes: Real-time detecting the hybrid management data of the target service in the shared Graphics Processing Unit (GPU) cluster, where the hybrid management data includes the usage information of the first GPU resources of the online tasks under the target service and the resource requirement information of the offline tasks under the target service; Calculating the amount of idle GPU resources in the first GPU resources of the online tasks based on the hybrid management data; Adjusting the offline task scheduling policy based on the amount of idle GPU resources and the resource requirement information of the offline tasks under the target service, where the offline task scheduling policy indicates the second GPU resources for the offline tasks under the target service.

2. The resource management method according to claim 1, wherein The shared GPU cluster is configured in a three-layer cluster architecture, and the three-layer cluster architecture includes: The upper-layer offline hybrid cluster for receiving the offline tasks of each service; The middle hybrid management and control layer configured with a first hybrid component for detecting the hybrid management data, calculating the amount of idle GPU resources, and adjusting the offline task scheduling policy; The lower-layer GPU cluster configured with a second hybrid component for managing each GPU node in the shared GPU cluster.

3. The resource management method according to claim 2, wherein The hybrid management data further includes the hybrid validity period and hybrid status of the online and offline hybrid; The calculating the amount of idle GPU resources in the first GPU resources of the online tasks based on the hybrid management data includes: Regularly detecting whether the current time is within the hybrid validity period; If the current time is within the hybrid validity period, putting the offline tasks under the target service into the hybrid processing queue; Calculating the amount of idle GPU resources in the first GPU resources of the online tasks based on the hybrid management data.

4. The resource management method according to claim 3, wherein The hybrid management data further includes the Quality of Service (QoS) indicators and detection rules of the online tasks; The adjusting the offline task scheduling policy based on the amount of idle GPU resources and the resource requirement information of the offline tasks includes: Real-time detecting the online QoS of the online tasks based on the QoS indicators and detection rules of the online tasks and the usage information of the first GPU resources; Deciding whether the offline task scheduling policy is to issue offline tasks or evict offline tasks according to the amount of idle GPU resources and the online QoS; If it is decided that the offline task scheduling policy is to issue offline tasks, determining the second GPU resources for the offline tasks under the target service based on the amount of idle GPU resources and the resource requirement information of the offline tasks under the target service, and issuing the offline tasks under the target service to the second GPU resources; or If it is decided that the offline task scheduling policy is to evict offline tasks, partially or fully reclaiming the idle GPU resources occupied by the offline tasks and reallocating the reclaimed idle GPU resources to the online tasks.

5. The resource management method according to claim 4, wherein If the decision-making offline task scheduling policy is to issue an offline task, based on the amount of idle GPU resources and the resource requirement information of the offline task under the target service, determine the second GPU resources for the offline task under the target service, and issue the offline task under the target service to the second GPU resources, including: If the decision-making offline task scheduling policy is to issue an offline task, the first hybrid component of the intermediate hybrid management layer determines the second GPU resources for the offline task under the target service based on the amount of idle GPU resources and the resource requirement information of the offline task under the target service; The first hybrid component issues the first offline task scheduling policy corresponding to the offline task under the target service to the upper-layer offline hybrid cluster, and allocates the offline task under the target service to the GPU node corresponding to the second GPU resources in the underlying GPU cluster.

6. The resource management method according to claim 5, wherein If the decision-making offline task scheduling policy is to issue an offline task, the method further includes: After allocating the second GPU resources, detect whether there are still remaining GPU resources in the amount of idle GPU resources; If there are still remaining GPU resources in the amount of idle GPU resources, the first hybrid component issues the second offline task scheduling policy corresponding to the offline tasks under other services to the upper-layer offline hybrid cluster, and allocates the offline tasks under other services to the GPU nodes corresponding to the remaining GPU resources in the underlying GPU cluster; Among them, the priority of the offline tasks under other services using the amount of idle GPU resources is lower than the priority of the offline tasks under the target service using the amount of idle GPU resources.

7. The resource management method according to claim 6, wherein If the decision-making offline task scheduling policy is to evict offline tasks, partially or fully reclaim the idle GPU resources occupied by the offline tasks, and reallocate the reclaimed idle GPU resources to the online tasks, including: If the decision-making offline task scheduling policy is to evict offline tasks, evict the offline tasks in the preset priority order, and the preset priority order is: first evict the offline tasks under other services, and then evict the offline tasks under the target service; Reallocate the reclaimed idle GPU resources to the online tasks; Among them, if the eviction volume of evicting offline tasks is full eviction, set the hybrid state to the eviction state; or if the eviction volume of evicting offline tasks is partial eviction, set the hybrid state to the resource-constrained hybrid state.

8. The resource management method according to claim 7, wherein The eviction of the offline tasks under other services includes: partially or fully reclaiming the remaining GPU resources occupied by the offline tasks under other services, and reallocating the reclaimed remaining GPU resources to the online tasks; The eviction of the offline tasks under the target service includes: partially or fully reclaiming the second GPU resources occupied by the offline tasks under the target service, and reallocating the reclaimed second GPU resources to the online tasks.

9. The resource management method according to any one of claims 4 to 8, characterized in that The method further includes: Real-time detect whether the current time is the end time of the hybrid validity period; If the current time is not the end time of the co-allocation validity period, repeat the adjustment of the offline task scheduling policy; or If the current time is the end time of the co-allocation validity period, prohibit the issuance of offline tasks.

10. The resource management method according to claim 3, characterized in that The method further includes: If the current time is not within the co-allocation validity period, detect whether there are offline tasks under the target service in the underlying GPU cluster; If there are offline tasks under the target service in the underlying GPU cluster, evict the offline tasks under the target service from the underlying GPU cluster and set the co-allocation state to a non-co-allocation state.

11. The resource management method according to claim 10, wherein, The method further includes: When evicting the offline tasks under the target service from the underlying GPU cluster, issue an alarm message for prompting the offline task eviction event.

12. The resource management method according to claim 2, wherein The method further includes: Through the second co-allocation component in the underlying GPU cluster, perform quality of service control on the resources of each offline task. The quality of service control includes adjusting the central processing unit (CPU) weight and network input / output priority of each offline task. Each offline task includes at least one of the online tasks under the target service and the online tasks under other services; According to the result of the quality of service control, readjust the offline task scheduling policy.

13. A resource management device, characterized in that, The apparatus includes: A detection unit, configured to detect in real time the co-allocation management data of the target service in the shared graphics processing unit (GPU) cluster. The co-allocation management data includes the usage information of the first GPU resources of the online tasks under the target service and the resource requirement information of the offline tasks under the target service; A calculation unit, configured to calculate the amount of idle GPU resources in the first GPU resources of the online tasks based on the co-allocation management data; An adjustment unit, configured to adjust the offline task scheduling policy based on the amount of idle GPU resources and the resource requirement information of the offline tasks under the target service. The offline task scheduling policy indicates the second GPU resources for use by the offline tasks under the target service.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the resource management method according to any one of claims 1-12.

15. A computer device, characterized in that, The computer device includes a processor and a memory. The memory stores a computer program, and the processor is configured to execute the resource management method according to any one of claims 1-12 by calling the computer program stored in the memory.

16. A computer program product comprising computer instructions, characterized in that, The computer instructions, when executed by a processor, implement the resource management method according to any one of claims 1-12.

Citation Information

Cited By

  • Task scheduling method and device, chip, equipment and storage medium

    CN121210071A