A GPU resource management method, system and computer readable medium
By obtaining GPU resource information on computer devices and generating registration information, adding it to resource clusters for management, the problem of idle GPU resources is solved, and resource utilization and flexibility of computing devices are improved.
Patent Information
- Application Number
- CN202411426986.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-10-12
AI Technical Summary
In the prior art, GPU resources on computer devices cannot be integrated and scheduled on demand, resulting in idle GPU resources and affecting resource utilization.
By obtaining GPU resource information on computer devices, generating registration information according to user needs, adding some GPU resources to the resource cluster for management, and using lightweight K3S clusters and Docker containers for virtualization management, ensuring that the cluster cannot know the status of the unjoined GPU resources.
It realizes on-demand integration and scheduling of GPU resources, improves resource utilization, meets diversified computing needs, and avoids invasive impact on the computer equipment environment.
Smart Images

Figure CN119311417B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of resource management, and in particular to a GPU resource management method, system and computer-readable medium. Background Art
[0002] A graphics processing unit (GPU) is a computing processor specifically designed for processing graphics and image rendering. It is mainly used to accelerate the display of computer graphics and videos and is widely used in computer-intensive tasks. Different computer tasks usually have different requirements for the quantity and model of GPU resources. Therefore, it is necessary to manage GPU resources in the computing environment to achieve the integration, allocation and scheduling of GPU resources to ensure that GPU resources can meet actual application needs. The service provider purchases physical servers and deploys them in the data center. Through the resource management platform, it monitors the usage of GPU resources and allocates and schedules resources. Users can rent the required GPU resources through the resource management platform to perform tasks.
[0003] A user or enterprise may have multiple physical computer devices. When a user needs to perform a task, if the GPU resources on each physical computer device cannot meet the task requirements, the GPU resources on different physical machines can be integrated into a cluster for unified management and scheduling. In one scenario, if a physical server has eight GPU cards, and the user on the physical server uses four GPU cards to perform a task, in order to prevent the remaining four GPU cards from being idle all the time, resulting in a waste of resources, the remaining four cards can be added to the cluster and reallocated to other users. However, when the GPU resources are integrated and scheduled by building a cluster, the working node of the cluster is the entire physical server, and all GPU cards of the physical server will be deployed to the cluster. The GPU resources that the user does not want to join the cluster are also managed. When other users use the GPU resources on the server, they can view the information of all GPU resources on the server, so that when the GPU resources are scheduled according to the needs of the users in the cluster, the four cards that the users on the physical server do not want to allocate are occupied. Summary of the invention
[0004] In view of this, the present application provides a GPU resource management method, system and computer-readable medium for implementing the addition of some GPU resources in a specified computer device to a cluster for management.
[0005] To solve the above problems, the technical solutions provided by this application are as follows:
[0006] On the one hand, the present application provides a GPU resource management method, comprising:
[0007] The computer device obtains information of M GPU resources on the computer device, where the information of the GPU resources includes an identity identifier, a type, and a status of the GPU resource, where the identity identifier is used to uniquely identify the GPU resource;
[0008] The computer device determines, according to user requirements, the identity identifiers of N GPU resources, wherein the user requirements include the N GPU resources that need to be uniformly managed as specified by the user according to the type, number and status of the GPU resources, wherein N is less than M;
[0009] The computer device generates first registration information according to the identity identifiers of the N GPU resources, and sends the first registration information to the resource cluster, where the first registration information is used to add the N GPU resources to the resource cluster for management.
[0010] In a possible implementation, the resource cluster includes a lightweight kubernets K3S cluster, the computer device generates first registration information according to the identity identifiers of the N GPU resources, and after sending the first registration information to the resource cluster, the method further includes:
[0011] The computer device creates a Docker container, which is used to mount the N GPU resources and manage the N GPU resources together with the resource K3S cluster.
[0012] In a possible implementation, after the computer device creates the Docker container, when the user demand changes, the method further includes:
[0013] The computer device deletes the Docker container;
[0014] The computer device determines the identity identifiers of K GPU resources according to the changed user requirements, wherein the changed user requirements include the K GPU resources that need to be uniformly managed and that are re-specified by the user according to the type, number, and status of the GPU resources, wherein K is less than M;
[0015] The computer device generates second registration information according to the identity identifiers of the K GPU resources, and sends the second registration information to the K3S resource cluster, where the second registration information is used to add the K GPU resources to the K3S resource cluster.
[0016] In a possible implementation, after the computer device generates first registration information according to the identity identifiers of the N GPU resources and sends the first registration information to the resource cluster, the method further includes:
[0017] When the computer device needs to use GPU resources locally to perform a task, a designated target GPU resource is determined to perform the task, where the target GPU resource is other GPU resources on the computer device except the N GPU resources.
[0018] In a possible implementation, after the computer device creates a Docker container, the method further includes:
[0019] The server of the K3S resource cluster obtains user configuration information, where the configuration information is used to identify the number and type of GPU resources required to execute the task;
[0020] The server configures a first target GPU resource according to the configuration information, where the first target GPU resource is a GPU resource in the K3S resource cluster that satisfies the configuration information;
[0021] The server schedules the first target GPU resources to execute the task.
[0022] In a possible implementation, the server scheduling the first target GPU resource to execute the task includes:
[0023] The server analyzes the demand of the task and the load of the first target GPU resource, wherein the demand includes the demand for the video memory of the GPU resource;
[0024] The server selects a second target GPU resource from the first target GPU resources according to the demand and the load of the first target GPU resource;
[0025] The server schedules the second target GPU resources to execute the task.
[0026] In another aspect, the present application provides a GPU resource management system, the system comprising a computer device, the computer device comprising:
[0027] An acquisition module, used for acquiring information of M graphics processing unit (GPU) resources on a computer device, wherein the information of the GPU resources includes an identity identifier, a type and a status of the GPU resources, and the identity identifier is used for uniquely identifying the GPU resources;
[0028] a determination module, configured to determine the identity identifiers of N GPU resources according to user requirements, wherein the user requirements include the N GPU resources that need to be uniformly managed as specified by the user according to the type, number and status of the GPU resources, wherein N is less than M;
[0029] A generation module is used to generate first registration information according to the identity identifiers of the N GPU resources, and send the first registration information to the resource cluster, where the first registration information is used to add the N GPU resources to the resource cluster.
[0030] In a possible implementation, the resource cluster includes a lightweight kubernets K3S cluster, and the computer device further includes:
[0031] A creation module is used to create a Docker container, where the Docker container is used to mount the N GPU resources and manage the N GPU resources together with the K3S cluster.
[0032] In a possible implementation, the system further includes a server, and the server includes:
[0033] An acquisition module, used to acquire user configuration information, where the configuration information is used to identify the number and type of GPU resources required to execute a task;
[0034] A configuration module, configured to configure a first target GPU resource according to the configuration information, where the first target GPU resource is a GPU resource in the K3S cluster that satisfies the configuration information;
[0035] A scheduling module is used to schedule the first target GPU resources to execute the task.
[0036] On the other hand, the present application provides a computer-readable medium, wherein the computer-readable medium is used to store a computer program, and when the computer program is executed by a computer device, the method is implemented.
[0037] It can be seen that this application has the following beneficial effects:
[0038] The computer device first obtains the identity identifiers, types and status information of all GPU resources installed on the device, and then, through the user's analysis of the types, number and status of GPU resources, specifies some GPU resources on the computer device to be added to the cluster for unified management. After determining the identity identifier corresponding to the specified GPU resource, registration information is generated according to the identity identifier, and a request is sent to the resource cluster, requesting that the GPU resource corresponding to the identity identifier be added to the resource cluster. Thus, a method of adding some GPU resources in the computer device to the cluster for management is realized, so that the cluster cannot know the GPU resources that have not been added, and will not affect the GPUs in the computer device that have not been added to the cluster. In this way, the idle GPU resources that meet the needs on different computer devices can be integrated and scheduled according to user needs, which meets the diverse computing needs and improves the utilization rate of GPU resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0040] Figure 1 A flowchart of a GPU resource management method provided in an embodiment of the present application;
[0041] Figure 2 A schematic diagram of the structure of a K3S cluster provided in an embodiment of the present application;
[0042] Figure 3 A schematic diagram of a scheduling structure of a K3S cluster provided in an embodiment of the present application;
[0043] Figure 4 A schematic diagram of a computer device provided in an embodiment of the present application;
[0044] Figure 5 A schematic diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0046] As described in the background technology, in one scenario, a computer device has multiple GPU cards. If you want to reserve some GPU cards for executing local personal tasks and add other GPUs on the computer device to a cluster for unified management, when other users use the GPU cards on the computer device, competition will arise for the GPU cards that the individual wants to reserve.
[0047] The present application provides a GPU resource management method, in which a computer device obtains information about all GPU resources installed on the device, and then specifies, based on user needs, some GPU resources on the computer device to be added to a resource cluster for management, generates registration information based on an identity identifier of the part of the GPU resources, and requests that the corresponding GPU resources be added to the resource cluster for management, so that the part of the GPU resources that meet the needs on different computer devices can be selected and integrated and scheduled according to user needs, thereby meeting diverse computing needs and improving the utilization rate of GPU resources.
[0048] The method provided in the present application is applied to a computer device having multiple GPU resources, which may be a terminal device or a server, wherein the server is an independent physical server, and the terminal device includes but is not limited to a desktop computer having multiple GPU resources.
[0049] The solution provided in the embodiments of the present application relates to technologies in the field of resource management, which are specifically described through the following embodiments.
[0050] See also Figure 1 As shown, it is a flowchart of a GPU resource management method provided in an embodiment of the present application. The GPU resource management method can be executed by a computer device. In this embodiment, a computer device that can be installed with multiple GPU cards is used as an example for explanation.
[0051] S101: A computer device obtains information of M GPU resources on the computer device.
[0052] The information of the GPU resource includes an identifier, type, and status of the GPU resource, and the identifier is used to uniquely identify the GPU resource.
[0053] The GPU resource is a GPU card installed on a computer device. The type of GPU resource can be type A, type B, type C, etc., and of course it can also be other types, without any restrictive explanation here. The identity identifier of the GPU resource is unique, and the identity identifier can be determined according to the location where it is installed in the computer device or the identification method of the computer device. The status of the GPU resource includes the real-time status and performance data of the GPU resource, such as GPU resource utilization, video memory usage, temperature, availability, and other detailed information.
[0054] A user can install multiple GPU resources on a computer device, and the types, identifiers or status information of these GPU resources can be detected and obtained by the computer device after they are installed on the computer device. The number of GPU resources installed on the computer device is M, where M is an integer greater than 1, which means that the method provided by the present application needs to be performed on a computer device with multiple GPU resources installed.
[0055] S102: The computer device determines the identity identifiers of N GPU resources according to user requirements.
[0056] The user demand includes N GPU resources that need to be uniformly managed as specified by the user according to the type, number and state of the GPU resources, where N is less than M and greater than 0.
[0057] The user can check the working status of the GPU resource through the command line or related applications, and view the information of all GPU resources installed on the computer device. For example, the user can view detailed information such as the identity identifier, status, and type of the GPU resource through the command.
[0058] The user analyzes the GPU resources viewed, whether they are available, and the number of different GPU resource types to determine which GPU resources to designate for unified management.
[0059] For example, a computer device has eight available GPU resources, including four type A GPU resources and four type B GPU resources. A user wants to reserve four type B GPU resources for executing his local tasks and designate four type A GPU resources for unified management. Then, the identity identifiers of the four type A GPU resources are determined.
[0060] For another example, a computer device has eight GPU resources, including four type A GPU resources and four type B GPU resources. The four type A GPU resources are executing computing tasks, and a user wants to reserve two type B GPU resources for executing other personal computing tasks and designate the remaining two type B GPU resources for unified management, then determine the identity identifiers of the two type B GPU resources.
[0061] Of course, there are other implementations of determining the identity identifiers of GPU resources that need to be uniformly managed according to actual needs, which all fall within the scope of protection of this application.
[0062] S103: The computer device generates first registration information according to the identity identifiers of the N GPU resources, and sends the first registration information to the resource cluster.
[0063] The first registration information is used to add N GPU resources to the resource cluster for management.
[0064] According to the GPU resources that need to be uniformly managed as specified by the user, the corresponding identity identifier can be determined, and then registration information can be generated based on the identity identifier, and the registration information can be sent to the resource cluster, and the GPU resources corresponding to the identity identifier can be added to the resource cluster for unified management, so as to provide GPU resources for other users, thereby achieving the purpose of incorporating some GPU resources on computer devices into the resource cluster for unified management.
[0065] For example, the designated GPU resources that need to be managed are four type A GPU resources, and the identity identifiers of the four GPU resources are determined to be 1, 2, 3, and 4, respectively. Then, registration information is generated based on the four identity identifiers and relevant parameters such as the network address of the computer device, and a registration request is sent to the resource cluster to add the GPU resources with identity identifiers 1-4 in the computer device to the resource cluster.
[0066] In summary, by obtaining information about all GPU resources installed on the device through a computer device, and by having the user specify some GPU resources on the computer device based on the information about the GPU resources, determining an identity identifier, and then generating registration information based on the identity identifier, and requesting that the corresponding GPU resources be added to a resource cluster, some GPU resources in the computer device can be specified for management.
[0067] In this way, some GPUs on computer devices are added to resource clusters for management. GPU resources on different computer devices can be integrated according to user needs, and unified management can be performed through resource clusters to improve the utilization rate of GPU resources.
[0068] When specifying part of the GPU resources to join the resource cluster, it is necessary to use virtualization technology to manage the GPU resources on the computer device by mapping them to the resource cluster.
[0069] In one possible implementation, the resource cluster includes a K3S cluster. The computer device generates first registration information based on the identity identifiers of N GPU resources and sends the first registration information to the resource cluster. Then, the computer device creates a Docker container for mounting the specified GPU resources and manages the N GPU resources together with the K3S cluster.
[0070] Lightweight kubernetes (K3S) is a lightweight Kubernetes version that removes unnecessary components compared to Kebernetes, has low resource usage and can be started quickly, making it more suitable for edge computing and resource-constrained environments. A K3S cluster is a group of computing nodes used to run containerized applications. The cluster includes at least one control node and one or more worker nodes. In this application, GPU resources mounted in computer devices through Docker containers are deployed on the worker nodes.
[0071] Figure 2 A schematic diagram of the structure of a K3S cluster provided for the implementation of this application is shown in the figure. Taking the K3S cluster having a control node and a working node as an example, the K3S cluster includes: a control node, i.e., a server node, and a working node, i.e., an agent node. The server node and the agent node in the K3S cluster both include the following components:
[0072] The node agent, in one possible implementation Kubelet, manages the lifecycle of the container.
[0073] A tunnel proxy, in one possible implementation, is used to establish a secure communication channel between the inside and outside of the cluster.
[0074] A network proxy, in one possible implementation Kube Proxy, is used to handle network traffic and load balance service requests.
[0075] The allocation component, in one possible implementation Flannel, is used to allocate a subnet to each host to enable communication between containers.
[0076] The container daemon, in one possible implementation, is Containerd, a high-performance container runtime responsible for managing containers in a container group.
[0077] A container group, in one possible implementation, is a Pod, the smallest scheduling unit in a cluster. Multiple containers can run in a Pod, and a container includes multiple GPU resources.
[0078] The server node can receive requests from different applications or models and schedule GPU resources in the K3S cluster. The server node also includes:
[0079] The scheduler, in one possible implementation, is responsible for scheduling the Pod to the appropriate node.
[0080] The control manager, in one possible implementation, is the Controller Manager, which is responsible for managing the cluster state and ensuring that the actual state is consistent with the desired state.
[0081] A supervisor, in one possible implementation, is a process management tool that is used to control and monitor the starting, stopping, and restarting of processes.
[0082] Define and manage multi-container Docker applications on agent nodes by using the container orchestration tool (Docker Compose).
[0083] Docker is a lightweight virtualization technology that allows developers to package applications and dependent packages into a portable image and flexibly migrate and deploy them between physical machines. When the Docker process is run through the K3S cluster, an independent process space is formed, and there is no need to configure the environment of the computer device.
[0084] A computer device can create a Docker container based on an image packaged by a Docker application. In this application, only one Docker container is deployed on a computer device, and the GPU resources specified on the computer device that need to be uniformly managed are uniformly mounted in the Docker container, and a mapping relationship is established with the cluster to achieve management of GPU resources.
[0085] By mounting the specified GPU resources to the same Docker container, users in the cluster can use the GPU resources to perform tasks without having to configure the computer device environment, thus avoiding interference with tasks being performed by other GPUs on the computer device and achieving non-intrusive impact on the computer device environment. In addition, adding computer devices to the K3S cluster can make them unrestricted by computer devices and manage them uniformly. And because the K3S cluster is lightweight and can be deployed faster, it allows the construction of K3S clusters on home devices, which is more suitable for individual user needs.
[0086] In actual application scenarios, user needs may change. When user needs change, it is necessary to recycle the GPU resources originally added to the cluster and add the GPU resources that meet the changed user needs back to the cluster for management.
[0087] In a possible implementation, after the computer device creates the Docker container, when the user's requirements change, the following steps S104-S106 may also be performed.
[0088] S104: The computer device deletes the Docker container.
[0089] In the computer device, a command is used to delete the Docker container that mounts the GPU resources that originally need to be managed, thereby releasing the GPU resources occupied by the corresponding Docker process.
[0090] S105: The computer device determines the identity identifiers of K GPU resources according to the changed user requirements.
[0091] The changed user demand includes K GPU resources that need to be uniformly managed and that are re-specified by the user according to the type, number and state of the GPU resources, where K is less than M and K is greater than 0.
[0092] For example, a computer device has eight available GPU resources, four A-type GPU resources and four B-type GPU resources. The user requirement before the change is to reserve four A-type GPU resources for executing local tasks and specify four B-type GPU resources to be added to the cluster for unified management, then the identity identifiers of the four B-type GPU resources are determined. The user requirement after the change is to reserve more GPU resources for executing local tasks and specify two B-type GPU resources to be added to the cluster, then the identity identifiers of the two B-type GPU resources are determined.
[0093] S106: The computer device generates second registration information according to the identity identifiers of the K GPU resources, and sends the second registration information to the K3S cluster.
[0094] The second registration information is used to add K GPU resources to the K3S cluster.
[0095] For example, the designated GPU resources that need to be managed uniformly are two B-type GPU resources, and the identity identifiers of the two GPU resources are determined to be 1 and 3 respectively. Then, registration information is generated based on the two identity identifiers and relevant parameters such as the network address of the computer device, and a registration request is sent to the K3S cluster to add the GPU resources with identity identifiers 1 and 3 in the computer device to the K3S cluster.
[0096] By deleting the Docker container, GPU resources can be removed from the cluster at any time, and specified GPU resources can be re-added to the cluster according to user needs without affecting the environment on the computer device, meeting the user's frequently changing computing needs.
[0097] After adding the GPU resources that need to be centrally managed to the resource cluster, if the user on the computer device needs to perform a personal task on the computer device, the GPU resources need to be specified, and the GPU resources that are not added to the resource cluster are selected to run the task.
[0098] In one possible implementation, after first registration information is generated based on the identity identifiers of N GPU resources and sent to the resource cluster, when the computer device needs to use GPU resources locally to perform a task, a designated target GPU resource is determined to perform the task, where the target GPU resource is other GPU resources on the computer device except the N GPU resources.
[0099] For example, a computer device has GPU resources with identifiers 1-8, of which type A GPU resources with identifiers 1-4 have been added to the resource cluster, and there are four type B GPU resources remaining on the computer device. Now a user on the computer device wants to perform a deep learning training task locally and requires two type B GPU resources for training. The user can select GPU resources with identifiers 5 and 6 from the remaining four type B GPU resources and specify these two GPU resources to perform the task through commands.
[0100] When personal tasks need to be performed on a computer device, by specifying GPU resources, only GPU resources that are not added to the cluster can be used to perform personal tasks, without causing competition for GPU resources added to the cluster, thereby improving the security and stability of tasks in the cluster.
[0101] The K3S cluster mentioned in the above embodiment needs to initialize the server node when it is created. First, install and deploy the K3S cluster server on a server, and deploy the necessary GPU management plug-in to support the detection and tag management of GPU resources. Then, after the computer device generates registration information, the server receives the registration information, detects the corresponding GPU resource according to the identity identifier in the registration information, registers the GPU resource to the resource library of the K3S cluster, and includes it in the scheduling range. After that, the K3S cluster manages and schedules the GPU resources added to the cluster through the server.
[0102] Figure 3 A schematic diagram of the scheduling structure of a K3S cluster provided for an embodiment of the present application. The control plane is responsible for the state, scheduling strategy and resource allocation of the cluster. The container daemon (Container), container (Docker), node agent (Kubelet), network agent (Kube proxy) and all or part of the GPU on the computer device are deployed in the working node. The components of the control plane are deployed on the server nodes, including distributed key-value storage, cluster scheduler, cluster control manager, cloud control manager and cluster application programming interface. The following is an introduction to each component:
[0103] Distributed key-value storage, in one possible implementation, is Etcd, which is used to store cluster status data and configuration to ensure data persistence.
[0104] The cluster scheduler, in one possible implementation, is Kube Scheduler, which is used to run the control manager, manage and optimize the allocation of cluster resources to ensure efficient operation.
[0105] The cluster control manager, in one possible case Kube Controller Manage, monitors the cluster status and performs appropriate adjustments.
[0106] The cluster application programming interface, in one possible case, is the Kube API Server, which is used to receive and process requests from user applications and is the management center of the cluster.
[0107] In a possible implementation, after the computer device creates a Docker container, the following steps S107-S109 may be further performed on the server of the K3S cluster:
[0108] S107: The server of the K3S cluster obtains user configuration information.
[0109] The configuration information is used to identify the number and type of GPU resources required to execute the task.
[0110] In the K3S cluster, the server obtains the user's configuration information to enable the interaction between the user's application and the GPU resources in the computer device. The user generates configuration information based on the number and type of GPU resources required to perform the task, and the server obtains the configuration information required by the user to perform the task through the interface or configuration file.
[0111] Specifically, users can define detailed information such as the number and type of required GPU resources in a resource configuration file according to task requirements. The server submits the configuration file to the application programming interface through the node agent, and the application programming interface receives and parses the configuration file.
[0112] S108: The server configures the first target GPU resources according to the configuration information.
[0113] Among them, the first target GPU resource is the GPU resource in the K3S resource cluster that meets the configuration information.
[0114] The server applies for GPU resources that match the requirements in the K3S cluster according to the number and type of GPU resources required to execute the task in the configuration information, and obtains the first target GPU resources that meet the user's needs.
[0115] Specifically, after the application programming interface parses the configuration file, the cluster scheduler selects GPU resources that meet the requirements on different computer devices for configuration according to the number and type of GPU resources requested in the configuration file.
[0116] S109: The server schedules the first target GPU resources to execute the task.
[0117] After the server allocates the first target GPU resource that meets the needs to the user, the GPU resource is used for scheduling to execute the user's task.
[0118] Specifically, after the server is configured with GPU resources, the node agent identifies the container where the GPU resources are located, and the scheduler schedules the task to the corresponding container to perform the computing task using the GPU resources.
[0119] The server ensures the consistency between user needs and GPU resource allocation and meets user needs by configuring and scheduling corresponding GPU resources according to user configuration information.
[0120] When the server schedules GPU resources according to the configuration information, it can adjust the scheduling strategy of GPU resources according to different tasks when user needs have not changed, so as to select the optimal GPU resources to ensure the operation of high-performance tasks.
[0121] In a possible implementation, the server schedules the first target GPU resource to execute the task, and may also execute steps S110 - S112 .
[0122] S110: The server analyzes task requirements and the load of the first target GPU resources.
[0123] Among them, the task requirements include the demand for GPU resources and video memory.
[0124] The server uses a monitoring tool to monitor the usage of the first target GPU resources in real time and detect the peak value and trend of the video memory usage of the first target GPU resources to analyze the video memory demand of the task and the load of the first target GPU resources.
[0125] S111: The server selects a second target GPU resource from the first target GPU resources according to demand and load of the first target GPU resources.
[0126] The server matches the appropriate GPU resources according to the current GPU resource load and the task's video memory requirements, and adjusts the scheduling strategy in a timely manner to optimize the use of GPU resources. For example, when it is detected that the GPU resources with high video memory in the first target GPU resources are idle, they can be reassigned to tasks that require high video memory GPU resources.
[0127] S112: The server schedules the second target GPU resources to execute the task.
[0128] The server schedules the corresponding GPU resources to execute tasks according to the adjusted scheduling strategy.
[0129] By analyzing the task requirements and the load of GPU resources, we can adjust the scheduling strategy of GPU resources, allocate the best GPU resources for different tasks, and schedule the tasks to be executed on the container where the GPU is located. When the demand does not change, critical tasks can use stable GPU resources, realize reasonable allocation of GPU resources, ensure the efficient operation of GPU resources, and improve the utilization rate of GPU resources.
[0130] Based on the above embodiments, the present application embodiment provides a GPU resource management system, which includes a computer device, referring to Figure 4 FIG. 4 is a schematic diagram of a computer device provided in an embodiment of the present application. The computer device 400 includes:
[0131] An acquisition module 401 is used to acquire information of M graphics processing unit (GPU) resources on a computer device, where the information of the GPU resources includes an identity identifier, a type, and a status of the GPU resources, where the identity identifier is used to uniquely identify the GPU resources;
[0132] A determination module 402 is used to determine the identity identifiers of N GPU resources according to user requirements, where the user requirements include the N GPU resources that need to be uniformly managed as specified by the user according to the type, number and status of the GPU resources, where N is less than M;
[0133] The generating module 403 is used to generate first registration information according to the identity identifiers of the N GPU resources, and send the first registration information to the resource cluster, where the first registration information is used to add the N GPU resources to the resource cluster.
[0134] In a possible implementation, the resource cluster includes a K3S cluster, and the computer device further includes:
[0135] The creation module 404 is used to create a Docker container, which is used to mount N GPU resources and manage the N GPU resources together with the K3S cluster.
[0136] In a possible implementation, the computer device further includes:
[0137] Deletion module 405, used to delete the Docker container when user requirements change;
[0138] The determination module 402 is further used to determine the identity identifiers of K GPU resources according to the changed user requirements, where the changed user requirements include the K GPU resources that the user re-specifies according to the type, number and status of the GPU resources and needs to be uniformly managed, where K is less than M;
[0139] The generation module 403 is also used to generate second registration information according to the identity identifiers of the K GPU resources, and send the second registration information to the K3S cluster, where the second registration information is used to add the K GPU resources to the K3S cluster.
[0140] In a possible implementation, the determination module 403 is further used to determine a designated target GPU resource to perform a task when a task needs to be performed locally using a GPU resource, where the target GPU resource is another GPU resource on the computer device except for the N GPU resources.
[0141] In a possible implementation, the system further includes a server, referring to Figure 5 FIG. 5 is a schematic diagram of a server provided in an embodiment of the present application. The server 500 includes:
[0142] An acquisition module 501 is used to acquire user configuration information, where the configuration information is used to identify the number and type of GPU resources required to execute a task;
[0143] A configuration module 502 is used to configure a first target GPU resource according to the configuration information, where the first target GPU resource is a GPU resource in the K3S cluster that meets the configuration information;
[0144] The scheduling module 503 is used to schedule the first target GPU resource to execute the task.
[0145] In a possible implementation, the scheduling module 503 includes:
[0146] An analysis module 5031 is used to analyze the demand of the task and the load of the first target GPU resource, where the demand includes the demand for the video memory of the GPU resource;
[0147] A selection module 5032, configured to select a second target GPU resource from the first target GPU resource according to demand and load of the first target GPU resource;
[0148] The scheduling submodule 5033 is used to schedule the second target GPU resources to execute tasks.
[0149] On the basis of the above embodiments, an embodiment of the present application further provides a computer-readable medium, on which a computer program is stored, and the computer program is processed to execute the method provided in the above embodiments.
[0150] On the basis of the above embodiments, the embodiments of the present application further provide a computer program product of a computer program, which, when executed on a computer device, enables the computer device to execute the method provided in the above embodiments.
[0151] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system or device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0152] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A GPU resource management method, characterized in that: The method comprises: The computer device obtains information of M graphics processing unit (GPU) resources on the computer device, where the information of the GPU resources includes an identity identifier, a type, and a status of the GPU resource, where the identity identifier is used to uniquely identify the GPU resource; The computer device determines, according to user requirements, the identity identifiers of N GPU resources, wherein the user requirements include the N GPU resources that need to be uniformly managed as specified by the user according to the type, number and status of the GPU resources, wherein N is less than M; The computer device generates first registration information according to the identity identifiers of the N GPU resources, and sends the first registration information to the resource cluster, where the first registration information is used to add the N GPU resources to the resource cluster for management; The resource cluster includes a lightweight K3S cluster; the computer device creates a Docker container, which is used to mount the N GPU resources and manage the N GPU resources together with the K3S cluster.
2. The method according to claim 1, characterized in that: After the computer device creates the Docker container, when the user demand changes, the method further includes: The computer device deletes the Docker container; The computer device determines the identity identifiers of K GPU resources according to the changed user requirements, wherein the changed user requirements include the K GPU resources that need to be uniformly managed and that are re-specified by the user according to the type, number, and status of the GPU resources, wherein K is less than M; The computer device generates second registration information according to the identity identifiers of the K GPU resources, and sends the second registration information to the K3S cluster, where the second registration information is used to add the K GPU resources to the K3S cluster.
3. The method according to claim 1, characterized in that After the computer device generates first registration information according to the identity identifiers of the N GPU resources and sends the first registration information to the resource cluster, the method further includes: When the computer device needs to use GPU resources locally to perform a task, a designated target GPU resource is determined to perform the task, where the target GPU resource is other GPU resources on the computer device except the N GPU resources.
4. The method according to claim 1, characterized in that: After the computer device creates the Docker container, the method further includes: The server of the K3S cluster obtains user configuration information, where the configuration information is used to identify the number and type of GPU resources required to execute the task; The server configures a first target GPU resource according to the configuration information, where the first target GPU resource is a GPU resource in the K3S cluster that satisfies the configuration information; The server schedules the first target GPU resources to execute the task.
5. The method according to claim 4, characterized in that: The server scheduling the first target GPU resource to execute the task includes: The server analyzes the demand of the task and the load of the first target GPU resource, wherein the demand includes the demand for the video memory of the GPU resource; The server selects a second target GPU resource from the first target GPU resources according to the demand and the load of the first target GPU resource; The server schedules the second target GPU resources to execute the task.
6. A GPU resource management system, characterized in that: The system includes a computer device, wherein the computer device includes: An acquisition module, used for acquiring information of M graphics processing unit (GPU) resources on a computer device, wherein the information of the GPU resources includes an identity identifier, a type and a status of the GPU resources, and the identity identifier is used for uniquely identifying the GPU resources; a determination module, configured to determine the identity identifiers of N GPU resources according to user requirements, wherein the user requirements include the N GPU resources that need to be uniformly managed as specified by the user according to the type, number and status of the GPU resources, wherein N is less than M; A generating module, configured to generate first registration information according to the identity identifiers of the N GPU resources, and send the first registration information to a resource cluster, wherein the first registration information is used to add the N GPU resources to the resource cluster; the resource cluster includes a lightweight kubernets K3S cluster; A creation module is used to create a Docker container, where the Docker container is used to mount the N GPU resources and manage the N GPU resources together with the K3S cluster.
7. The system according to claim 6, characterized in that: The system further comprises a server, wherein the server comprises: An acquisition module, used to acquire user configuration information, where the configuration information is used to identify the number and type of GPU resources required to execute a task; A configuration module, configured to configure a first target GPU resource according to the configuration information, where the first target GPU resource is a GPU resource in the K3S cluster that satisfies the configuration information; A scheduling module is used to schedule the first target GPU resources to execute the task.
8. A computer-readable medium, characterized in that The computer-readable medium is used to store a computer program, and when the computer program is executed by a computer device, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Resource scheduling device, resource scheduling system and resource scheduling method
CN108804217A
GPU resource-oriented task scheduling method, device and system
CN109992422A