GPU resource scheduling method, system, server and device
By mounting and uninstalling GPU resources on CPU devices, the resource waste caused by idle GPU devices is solved, and efficient resource utilization and cost reduction are achieved.
Patent Information
- Application Number
- PCT/CN2024/126181
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-10-21
- Publication Date
- 2025-06-26
AI Technical Summary
In the prior art, GPU devices are often idle when the user does not run the GPU application, resulting in wasted GPU resources.
Provides a GPU resource scheduling method that allows mounting and uninstalling GPU resources on a CPU device. Users mount GPU resources when they need to run GPU applications, and uninstall resources when the application is over or no longer needed.
By flexibly managing GPU resources, the idle time of GPU devices is reduced, the cost of users using GPUs is reduced, and the waste of GPU resources is effectively reduced.
Smart Images

Figure CN2024126181_26062025_PF_FP_ABST
Abstract
Description
GPU resource scheduling method, system, server and device
[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on December 18, 2023, with application number 202311746733.5 and application name “GPU resource scheduling method, system, server and device,” the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0002] The present disclosure relates to computer technology, and in particular to a GPU resource scheduling method, system, server, and device. Background Art
[0003] The main functions of a graphics processing unit (GPU) include graphics rendering, image processing, and computational acceleration. In fields such as gaming, animation, and visual effects, GPUs are essential components for achieving high-quality graphics and images. Furthermore, in fields such as scientific computing and deep learning, GPUs can also be used as computational accelerators, significantly improving computational speed and efficiency. With the rapid development of artificial intelligence (AI) technology, the demand for GPUs in this field is increasing. However, GPUs are relatively expensive, making them very costly for the average user.
[0004] A GPU can only function properly when installed on a device with a central processing unit (CPU) and a hard disk. To run an application that requires a GPU (called a GPU application), users must purchase or assemble a GPU device and then run the GPU application on it.
[0005] However, in actual applications, some users only use GPU devices when running GPU applications in specific task scenarios, and usually do not need to run GPU applications for a long time. When users do not run GPU applications, the GPU device is idle, resulting in a waste of GPU resources.
[0006] Summary of the Invention
[0007] The present disclosure provides a GPU resource scheduling method, system, server, and device to solve the problem of GPU resource waste in existing solutions.
[0008] In a first aspect, the present disclosure provides a GPU resource scheduling method, applied to a CPU device, the method comprising:
[0009] Before running the GPU application, mount the GPU resource;
[0010] Run the GPU application, and when a GPU task needs to be executed, send GPU task information to the GPU resource and receive the GPU task execution result returned by the GPU resource;
[0011] When the GPU unloading condition is met, the GPU resources are unloaded.
[0012] In a second aspect, the present disclosure provides a GPU resource scheduling method, applied to a GPU resource scheduling server, the method comprising:
[0013] In response to a GPU resource mounting request from a CPU device, allocating mountable GPU resources to the CPU device;
[0014] Send configuration information of the GPU resources allocated to the CPU device to the CPU device, where the configuration information is used by the CPU device to mount the GPU resources before running the GPU application, use the GPU resources during the running of the GPU application, and unload the GPU resources when a GPU unloading condition is met.
[0015] In a third aspect, the present disclosure provides a GPU resource scheduling method, which is applied to GPU resources. The method includes:
[0016] Receive GPU task information sent by a CPU device, wherein the GPU task is information about a GPU task that needs to be executed when the CPU device runs a GPU application, and the GPU resource is mounted on the CPU device and will be unloaded by the CPU device when a GPU unloading condition is met;
[0017] Execute the GPU task according to the GPU task information and obtain a GPU task execution result;
[0018] Return the GPU task execution result to the CPU device.
[0019] In a fourth aspect, the present disclosure provides a GPU resource scheduling system, comprising: a CPU device, GPU resources, and a GPU resource scheduling server.
[0020] The GPU resource scheduling server is used to: allocate GPU resources to the CPU device in response to a GPU resource mounting request of the CPU device, and send configuration information of the GPU resources allocated to the CPU device to the CPU device;
[0021] The CPU device is used to obtain configuration information of the GPU resources allocated to the CPU device, and mount the GPU resources according to the configuration information of the GPU resources.
[0022] The CPU device is further configured to: run a GPU application and send GPU task information to the GPU resource when a GPU task needs to be executed;
[0023] The GPU resource is used to: receive GPU task information sent by the CPU device, execute the GPU task according to the GPU task information, obtain a GPU task execution result, and return the GPU task execution result to the CPU device;
[0024] The CPU device is further configured to: receive a GPU task execution result returned by the GPU resource;
[0025] The CPU device is further configured to: unload the GPU resources.
[0026] In a fifth aspect, the present disclosure provides a CPU device, comprising:
[0027] at least one CPU; and a memory communicatively connected to the at least one CPU;
[0028] The memory stores instructions that can be executed by the at least one CPU, and the instructions are executed by the at least one CPU to enable the CPU device to execute the method described in the first aspect.
[0029] In a sixth aspect, the present disclosure provides a GPU device, comprising:
[0030] at least one GPU, at least one CPU, and a memory communicatively connected to the at least one GPU;
[0031] The memory stores instructions that can be executed by the at least one GPU, and the instructions are executed by the at least one GPU to enable the GPU device to execute the method described in the second aspect.
[0032] In a seventh aspect, the present disclosure provides a GPU resource scheduling server, comprising:
[0033] at least one processor; and a memory communicatively coupled to the at least one processor;
[0034] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the GPU resource scheduling server to execute the method described in the second aspect.
[0035] In an eighth aspect, an embodiment of the present disclosure provides a computer storage medium, wherein a computer program is stored on the computer storage medium, and when the computer program is executed by a processor, the method described in the first aspect, the second aspect or the third aspect is implemented.
[0036] In a ninth aspect, the present disclosure provides a computer program product, comprising a computer program, wherein the computer program is executed by a processor to implement the method described in the first aspect, the second aspect or the third aspect.
[0037] The GPU resource scheduling method, system, server and device provided by the present disclosure support flexible mounting and unmounting of GPU resources on CPU devices. Before a user needs to run a GPU application, the user mounts the GPU resource on the CPU device and runs the GPU application on the CPU device. During the running of the GPU application, when a GPU task needs to be executed, the CPU device sends GPU task information to the mounted GPU resource, so that the corresponding GPU resource executes the GPU task and returns the GPU task execution result to the CPU device, and the CPU device can obtain the GPU task execution result. When the GPU unloading conditions are met, the GPU resource can be unloaded from the CPU device to release the mounted GPU resource and reduce the waste of GPU resources. By mounting the GPU resource on the CPU device when the GPU resource is needed and unloading the mounted GPU resource when it is not needed, the cost of using the GPU for the user can be greatly reduced and the waste of GPU resources can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0039] FIG1 is an architecture diagram of a GPU resource scheduling system provided by an exemplary embodiment of the present disclosure;
[0040] FIG2 is a flowchart of a GPU resource scheduling method provided by an exemplary embodiment of the present disclosure;
[0041] FIG3 is a framework diagram of GPU resource scheduling provided by an exemplary embodiment of the present disclosure;
[0042] FIG4 is a flow chart of a GPU resource scheduling method provided by another exemplary embodiment of the present disclosure;
[0043] FIG5 is a flow chart of a GPU resource scheduling method provided by another exemplary embodiment of the present disclosure;
[0044] FIG6 is an interactive flow chart of GPU resource scheduling provided by an exemplary embodiment of the present disclosure;
[0045] FIG7 is a flow chart of a user using a GPU according to an exemplary embodiment of the present disclosure;
[0046] FIG8 is a schematic structural diagram of a CPU device provided by an exemplary embodiment of the present disclosure;
[0047] FIG9 is a schematic structural diagram of a GPU device provided by an exemplary embodiment of the present disclosure;
[0048] FIG10 is a schematic structural diagram of a GPU resource scheduling server provided by an exemplary embodiment of the present disclosure.
[0049] The above drawings illustrate specific embodiments of the present disclosure, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the present disclosure in any way, but rather to illustrate the concepts of the present disclosure to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0050] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0051] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and corresponding operation entrances must be provided for users to choose to authorize or refuse.
[0052] First, the terms involved in this disclosure are explained:
[0053] CPU device: refers to a computer device with a CPU installed (without a GPU installed) and CPU computing capabilities.
[0054] GPU device: refers to a computer device that has both a CPU and a GPU installed and has GPU computing capabilities.
[0055] GPU resources: refers to device resources with GPU computing capabilities, also known as GPU devices.
[0056] In response to the problem in the prior art that users need to purchase or assemble their own GPU devices and run GPU applications on the GPU devices, and when the user is not running the GPU application, the GPU device is idle, resulting in a waste of GPU resources. This embodiment provides a GPU resource scheduling method that supports flexible mounting and unmounting of GPU resources on a CPU device. Before a user needs to run a GPU application, they mount the GPU resources on their CPU device and can run the GPU application on the CPU device. When a GPU task needs to be executed during the execution of the GPU application, the CPU device sends GPU task information to the mounted GPU resources, causing the corresponding GPU resources to execute the GPU task and return the GPU task execution result to the CPU device, so that the CPU device can obtain the GPU task execution result. When the GPU resources are no longer needed, the GPU resources can be unloaded from the CPU device to release the mounted GPU resources and reduce the waste of GPU resources. By mounting GPU resources on the CPU device when needed and unmounting them when not needed, the cost of using the GPU can be greatly reduced and the waste of GPU resources can be reduced.
[0057] FIG1 is an architecture diagram of a GPU resource scheduling system provided by an exemplary embodiment of the present disclosure. As shown in FIG1 , the system architecture includes: a CPU device, GPU resources, and a GPU resource scheduling server. A CPU device is a computer device with a CPU installed and used by a user. GPU resources refer to GPU devices provided by a GPU service platform, which provides GPU devices to users as mountable resources. The GPU service platform allocates available GPU resources to CPU devices (or users) that request to mount GPU resources through the GPU resource scheduling server.
[0058] Based on the system architecture shown in Figure 1, the general process of GPU resource scheduling is as follows:
[0059] 1) When a user needs to run a GPU application, the user sends a GPU resource mount request to the GPU resource scheduling server through the CPU device used;
[0060] 2) The GPU resource scheduling server allocates available GPU resources to the CPU device (or user) and returns the configuration information of the allocated GPU resources to the CPU device;
[0061] 3) The CPU device mounts the GPU resources allocated to it;
[0062] 4) CPU devices run GPU applications;
[0063] 5) When the CPU device needs to execute a GPU task, it sends GPU task information to the mounted GPU resource;
[0064] 6) The GPU resource executes the GPU task based on the GPU task information, obtains the GPU task execution result, and returns the GPU task execution result to the CPU device;
[0065] 7) When the GPU unloading conditions are met (such as no need to run GPU applications), the CPU device unloads the mounted GPU resources.
[0066] The following detailed description of the technical solution of the present disclosure and how the technical solution of the present disclosure solves the above-mentioned technical problems is provided with specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The embodiments of the present disclosure will be described below in conjunction with the accompanying drawings.
[0067] FIG2 is a flowchart of a GPU resource scheduling method provided by an exemplary embodiment of the present disclosure. The execution subject of this embodiment is the CPU device in the aforementioned system architecture, which can be a mobile terminal, personal computer, artificial intelligence (AI) device, Internet of Things (IoT) device, server device, etc. equipped with a CPU. This embodiment is not specifically limited here. As shown in FIG2, the specific steps of this method are as follows:
[0068] Step S201: Before running the GPU application, mount the GPU resources.
[0069] In this embodiment, the CPU device mounts GPU resources before running the GPU application. The mounted GPU resources can be GPU cloud servers, which give GPU resources the advantages of regional deployment and flexible application. In addition, in some scenarios, the GPU resources can also be local GPU devices. For example, the user is an organization with a large number of CPU devices and a small number of GPU devices. In order to enable a large number of CPU devices to share the use of a small number of GPU devices, the CPU device can temporarily occupy and use a GPU device by applying for and mounting it. Of course, there are other situations where local GPU devices are used as GPU resources, which will not be explained here one by one.
[0070] Specifically, in response to the GPU resource mount instruction, the CPU device obtains configuration information of the GPU resources allocated to the CPU device and mounts the GPU resources based on the GPU resource configuration information. The GPU resource mount instruction can be a command sent by a user to the CPU device via a command, a visual interface, or the like, instructing the GPU device to mount the GPU resources. In an optional embodiment, the GPU resource mount instruction can be a GPU application execution instruction, which triggers the CPU device to mount the GPU resources when the user starts the GPU application.
[0071] Optionally, the CPU device obtains configuration information of GPU resources allocated to the CPU device, which can be specifically implemented in the following manner:
[0072] The CPU device sends a GPU resource mount request to the GPU resource scheduling server. The GPU resource mount request is used to request the GPU resource scheduling server to allocate available GPU resources to the CPU device and return the configuration information of the GPU resources allocated to the CPU device to the CPU device. The CPU device receives the configuration information of the GPU resources allocated to the CPU device.
[0073] Optionally, the user can obtain the configuration information of the GPU resources allocated to the user by the GPU resource scheduling server through the front-end interface provided by the GPU server platform, and input the configuration information of the GPU resources to the CPU device. The CPU device can receive the configuration information of the GPU resources input by the user as the configuration information of the GPU resources allocated to the CPU device.
[0074] Furthermore, the CPU device mounts the GPU resources according to the configuration information of the GPU resources. Specifically, this can be achieved in the following manner: the CPU device mounts the software package that the GPU application depends on to run to a specified directory according to the configuration information of the GPU resources, where the software package includes the implementation method of the GPU interface; and configures the network address of the GPU resources.
[0075] The GPU resource configuration information includes the network address of the GPU resource, i.e., the Internet Protocol Address (IP address). Based on the GPU resource's IP address, the CPU device and the mounted GPU resource can communicate with each other over a network. Specifically, the communication can be achieved via a public network or a proprietary network, which is not specifically limited here. In another embodiment, the GPU resource configuration information may not include the network address of the GPU resource. After the GPU resource is mounted, the CPU device obtains the network address of the mounted GPU resource from the GPU resource scheduling server when starting to run the GPU application.
[0076] The configuration information of the GPU resources may also include relevant information about the software package that the GPU application depends on when running on the CPU device. The CPU device can automatically obtain and install the software package based on the relevant information of the software package to support the operation of the GPU application. The software package is pre-developed and packaged by relevant technical personnel. The software package includes but is not limited to the GPU software library that the GPU resource depends on when running GPU tasks. The GPU software library contains alternative implementation methods for the GPU interface corresponding to each GPU task. In another optional embodiment, the configuration information of the GPU resource may not include relevant information about the software package. The relevant information about the software package can be publicly displayed on the front-end page. The CPU device of any user can obtain and install the software package based on the relevant information of the disclosed software package, providing a software foundation for mounting and using the GPU resource. The relevant information about the software package may include information such as a software package download link and a designated installation directory. The CPU device can automatically download the software package through the download link and install it to the designated directory.
[0077] It should be noted that the alternative implementation method of the GPU interface in this software package is a modification of the original implementation method of the GPU interface. When the CPU device issues a GPU interface call, the alternative implementation method of the GPU interface sends the task information corresponding to the GPU interface to the mounted GPU resource through the alternative implementation method, causing the GPU resource to execute the original implementation method of the GPU interface according to the GPU task information and return the GPU task execution result to the CPU device. In other words, the CPU device does not execute the actual (original) implementation method of the GPU interface. Instead, it sends the GPU task information corresponding to the GPU interface to the GPU resource, which then executes the actual implementation method of the GPU interface and returns the GPU task execution result to the CPU device.
[0078] The GPU task information includes input data and interface identification information of the GPU interface. Based on the interface identification information of the GPU interface, the GPU resource can determine the implementation method of the GPU interface.
[0079] Optionally, when the CPU device mounts the GPU resource based on the GPU resource configuration information, the CPU device installs a specific software package that interfaces with the GPU resource. This specific software package includes software packages that GPU applications rely on to run. The software packages that GPU applications rely on to run are described above. The specific software package that interfaces with the GPU resource is also responsible for intercepting GPU tasks generated by the GPU application, sending GPU task information to the GPU resource, and returning GPU task execution results returned by the GPU resource to the GPU application running on the CPU device. This specific software package is developed and packaged by relevant technical personnel.
[0080] Step S202: Run the GPU application, and when a GPU task needs to be executed, send GPU task information to the GPU resource, and receive the GPU task execution result returned by the GPU resource.
[0081] After mounting the GPU resources, the CPU device runs the GPU application. During the GPU application's execution, the GPU application internally calls one or more GPU interfaces, triggering the corresponding GPU tasks. For example, the GPU application typically calls interfaces for allocating storage space and loading data.
[0082] In this step, while running a GPU application, the CPU intercepts the GPU application's call request to the GPU interface and obtains the GPU task information corresponding to the GPU interface. Furthermore, the CPU sends the GPU task information to the GPU resource, causing the GPU resource to execute the GPU task based on the GPU task information and return the GPU task execution result to the CPU. The CPU receives the GPU task execution result corresponding to the GPU interface returned by the GPU resource.
[0083] Optionally, the GPU task information may include input data and interface identification information of the GPU interface. Based on the interface identification information of the GPU interface, the GPU resources may determine the implementation method of the GPU interface, thereby determining the processing logic to be executed by the GPU task. The implementation method of the GPU interface is used to execute the corresponding GPU task.
[0084] Optionally, the GPU task information may include input data and an operation identifier of the GPU interface. Based on the operation identifier of the GPU interface, the GPU resource may determine which operation the GPU interface corresponds to for the GPU task, thereby executing an implementation program for the corresponding operation.
[0085] When running a GPU application, the CPU device can intercept alternative implementation methods for each GPU interface. When a GPU application calls a GPU interface, the CPU device intercepts the alternative implementation method of the called GPU interface and, by executing the alternative implementation method, sends the GPU task information corresponding to the GPU interface to the mounted GPU resource, receives the GPU task execution results returned by the GPU resource, and then returns the GPU task execution results to the GPU application via a callback interface.
[0086] Exemplarily, FIG3 is a framework diagram of GPU resource scheduling provided by this embodiment. FIG3 takes the GPU resource as an example of a GPU cloud server for exemplary explanation. As shown in FIG3, a GPU application is run on a CPU device, and a GPU interface is internally called when the GPU application is run. A specific software package is installed on the CPU device, and a specific software service is installed on the GPU cloud server. The CPU device realizes network communication with the specific software service of the mounted GPU cloud server through the specific software package. The specific software package on the CPU device can intercept the GPU interface called by the GPU application, generate GPU task information corresponding to the GPU interface, and send the GPU task information to the GPU cloud server. The GPU cloud server receives the GPU task information sent by the CPU device through the installed specific software service. The GPU cloud server executes the GPU task according to the GPU task information to obtain the GPU task execution result, and returns the GPU task execution result to the specific software package running on the CPU device through the GPU software service.
[0087] In this embodiment, the CPU device and the mounted GPU resources communicate through a network rather than a traditional hardware plug-in card interconnection. The interconnection method is more flexible, so that the GPU device can flexibly mount and unmount GPU resources.
[0088] Step S203: Unload GPU resources when the GPU unloading condition is met.
[0089] In this embodiment, when the GPU resources are no longer needed, the CPU device can uninstall the mounted GPU resources at any time, thereby releasing the mounted GPU resources for use by other CPU devices.
[0090] Optionally, the CPU device can determine that GPU resources are no longer needed and unload the GPU resources based on pre-configured GPU offloading conditions when the GPU offloading conditions are met. The GPU offloading conditions can include the CPU device no longer needing to run GPU applications and / or the CPU device not running GPU applications for a preset duration. The GPU offloading conditions can be configured and adjusted based on actual usage scenarios and are not specifically limited here. For example, after all GPU applications on the CPU device have completed execution, the CPU device unloads GPU resources.
[0091] For example, when a user determines that the CPU device is temporarily no longer needed to run GPU applications, the user sends a GPU resource unloading instruction to the CPU device. The instruction can be issued via a visual interface or a command line, which is not specifically limited here. In response to receiving the GPU resource unloading instruction, the CPU device determines that the CPU resources are temporarily no longer needed and unloads the mounted GPU resources.
[0092] Specifically, when a CPU device unloads GPU resources, it sends a GPU resource unloading request to a GPU resource scheduling server. This GPU resource unloading request is used to request that the GPU resource scheduling server release the GPU resources allocated to the CPU device. Upon receiving the GPU resource unloading request from the CPU device, the GPU resource scheduling server releases the GPU resources allocated to the CPU device, for example, by marking the GPU resources as allocable or idle. The released GPU resources can be allocated to other CPU devices or reassigned to the CPU device.
[0093] Exemplarily, the GPU resource offload request may include identification information of the GPU resource to be offloaded. The GPU resource scheduling server releases the corresponding GPU resource based on the identification information of the GPU resource to be offloaded. Alternatively, the GPU resource offload request may not include identification information of the GPU resource, and the GPU resource scheduling server releases all GPU resources mounted on the CPU.
[0094] It should be noted that a CPU device can mount one or more GPU resources. When mounting multiple GPU resources, the software packages that the CPU device relies on to run GPU applications and the specific software packages that interface with the GPU resources do not need to be reinstalled. Furthermore, when uninstalling GPU resources, the software packages that the CPU device relies on to run GPU applications and the specific software packages that interface with the GPU resources can be deleted or retained, eliminating the need to reinstall the software packages the next time the GPU resources are mounted. However, when remounting the GPU resources, the network address of the mounted GPU resources must be reconfigured.
[0095] According to the solution of this embodiment, before the user needs to run the GPU application, he can mount the GPU resources through the CPU device he owns, and then run the GPU application on the CPU device without making any modifications to the GPU application. During the process of running the GPU application, when it is necessary to execute a GPU task, the CPU device sends the GPU task information to the mounted GPU resource, so that the corresponding GPU resource executes the GPU task and returns the GPU task execution result to the CPU device, and the CPU device can obtain the GPU task execution result. When it is not necessary to run the GPU application, the GPU resources can be uninstalled from the CPU device to release the mounted GPU resources and reduce the waste of GPU resources. By flexibly mounting and uninstalling GPU resources on the CPU device, mounting GPU resources on the CPU device when needed and uninstalling the mounted GPU resources when not needed, the cost of using the GPU for users can be greatly reduced and the waste of GPU resources can be reduced.
[0096] FIG4 is a flowchart of a GPU resource scheduling method provided by another exemplary embodiment of the present disclosure. The execution subject of this embodiment is the GPU resource scheduling server in the aforementioned system architecture, which can be a cloud server or a user's local server, etc., and this embodiment is not specifically limited here. As shown in FIG4, the specific steps of this method are as follows:
[0097] Step S401: In response to a GPU resource mounting request from a CPU device, allocate mountable GPU resources to the CPU device.
[0098] In this embodiment, when a CPU device needs to mount GPU resources, it sends a GPU resource mount request to a GPU resource scheduling server. The GPU resource scheduling server receives the GPU resource mount requests sent by each CPU device and allocates mountable GPU resources to the CPU device based on the GPU resource mount requests.
[0099] Optionally, the GPU resource scheduling server can maintain a GPU resource information table that records the status information (including but not limited to idle and allocated), configuration information, and specification information of each GPU resource. When the GPU resource scheduling server allocates a mountable GPU resource to a CPU device, it selects an idle GPU resource based on the GPU resource information table and allocates it to the CPU device. In addition, if a GPU resource has been mounted, the GPU resource information table can also record information about the CPU device to which the GPU resource is mounted, including but not limited to the identification information and network address of the mounted CPU device.
[0100] In some optional embodiments, the GPU resource mount request sent by the CPU device may also include GPU resource specification requirements. Based on the GPU resource specification requirements in the GPU resource mount request, the GPU resource scheduling server allocates a mountable GPU resource that meets the requirements to the CPU device. The GPU service platform can provide a variety of GPU resources with different specifications, each with different hardware configurations and computing capabilities. Users can choose which GPU resource specifications to mount on the CPU device based on their needs.
[0101] Step S402: Send configuration information of the GPU resources allocated to the CPU device to the CPU device. The configuration information is used by the CPU device to mount the GPU resources before running the GPU application, use the GPU resources during the running of the GPU application, and unload the GPU resources when the GPU unloading conditions are met.
[0102] After allocating mountable GPU resources to the CPU, the GPU resource scheduling server sends the configuration information of the GPU resources allocated to the CPU device to the CPU device. The CPU device mounts the corresponding GPU resources according to the configuration information of the GPU resources allocated to it.
[0103] The GPU resource configuration information includes the GPU resource's network address, also known as the Internet Protocol Address (IP address). Based on the GPU resource's IP address, the CPU device and the mounted GPU resource can establish network communication and interconnection. Specifically, interconnection can be achieved through the public network or a proprietary network, which is not specifically limited here.
[0104] The GPU resource configuration information may also include information about software packages that the CPU device relies on to run GPU applications. Based on this information, the CPU device can automatically obtain and install the software packages to support the GPU application. These software packages are pre-developed and packaged by relevant technical personnel. These software packages include, but are not limited to, the GPU software libraries that the GPU resources rely on to run GPU tasks, including alternative implementations of the GPU interfaces corresponding to each GPU task.
[0105] Optionally, the GPU resource scheduling server provides the CPU device with a software package that the GPU application depends on. The CPU device can automatically download and install the software package from the GPU resource scheduling server based on the download information of the software package to support the operation of the GPU application.
[0106] When the GPU application is not needed, the CPU device can uninstall the mounted GPU resources at any time to release the mounted GPU resources for use by other CPU devices. In an optional embodiment, after step S402, the following steps may also be included:
[0107] Step S403: In response to the GPU resource unloading request of the CPU device, the GPU resources allocated to the CPU device are released.
[0108] Specifically, when the CPU device unloads GPU resources, it sends a GPU resource unloading request to the GPU resource scheduling server.
[0109] After receiving the GPU resource offload request from the CPU device, the GPU resource scheduling server releases the GPU resources allocated to the CPU device, for example, marking the GPU resources as allocable or idle. The released GPU resources can be allocated to other CPU devices or allocated to the CPU device again.
[0110] Exemplarily, the GPU resource scheduling server may update the status information of the GPU resources in the GPU resource information table to make the status of the released GPU resources more idle, so that the released GPU resources can be allocated to other CPU devices, or allocated to the CPU device again.
[0111] In some optional embodiments, the GPU resource offload request may include identification information of the GPU resource to be offloaded. The GPU resource scheduling server releases the corresponding GPU resource based on the identification information of the GPU resource to be offloaded. Alternatively, the GPU resource offload request may not include any identification information of the GPU resource, and the GPU resource scheduling server releases all GPU resources mounted on the CPU.
[0112] The solution of this embodiment is based on the scheduling of GPU resources by the GPU resource scheduling server. When the CPU device needs to run a GPU application, it allocates mountable GPU resources to the CPU device, so that the CPU device can run the GPU application after mounting the GPU resources. In the process of running the GPU application, when it is necessary to execute a GPU task, the CPU device sends GPU task information to the mounted GPU resource, so that the corresponding GPU resource executes the GPU task and returns the GPU task execution result to the CPU device, and the CPU device can obtain the GPU task execution result. When the GPU application is not needed to run, the GPU resource scheduling server responds to the GPU resource unloading request of the CPU device and releases the GPU resources mounted on the CPU device, thereby reducing the waste of GPU resources. By flexibly mounting and unmounting GPU resources on the CPU device, mounting GPU resources on the CPU device when needed and unmounting the mounted GPU resources when not needed, the cost of using the GPU for users can be greatly reduced and the waste of GPU resources can be reduced.
[0113] Figure 5 is a flowchart of a GPU resource scheduling method provided by another exemplary embodiment of the present disclosure. The execution subject of this embodiment is any GPU resource in the aforementioned system architecture, which can be a GPU cloud server or a user's local GPU device, etc. This embodiment is not specifically limited here. As shown in Figure 5, the specific steps of this method are as follows:
[0114] Step S501: Receive GPU task information sent by a CPU device, wherein the GPU task is information about a GPU task that needs to be executed when the CPU device runs a GPU application. GPU resources are mounted on the CPU device and will be unloaded by the CPU device when a GPU unloading condition is met.
[0115] The CPU device refers to the CPU device that mounts the GPU resource. When the GPU unloading conditions are met, such as when the GPU resource is not needed (such as when the GPU application does not need to be run), the CPU device will unload the GPU resource.
[0116] When running a GPU application, the GPU application internally calls one or more GPU interfaces, triggering corresponding GPU tasks. The CPU device, which has GPU resources mounted on it, intercepts the GPU interfaces called by the GPU application, obtains the GPU task information corresponding to the GPU interfaces, and sends the GPU task information to the GPU resource. In this step, the GPU resource receives the GPU task information sent by the CPU device.
[0117] Step S502: Execute the GPU task according to the GPU task information to obtain a GPU task execution result.
[0118] In this embodiment, the GPU resource executes the corresponding GPU task according to the GPU task information sent by the CPU device, and obtains the GPU task execution result.
[0119] Optionally, the GPU task information may include input data and interface identification information of the GPU interface. Based on the interface identification information of the GPU interface, the GPU resources may determine the implementation method of the GPU interface, thereby determining the processing logic to be executed by the GPU task. The implementation method of the GPU interface is used to execute the corresponding GPU task.
[0120] In this step, the CPU resources run the implementation method of the GPU interface according to the input data and interface identification information of the GPU interface, complete the execution of the GPU task corresponding to the GPU interface, and obtain the execution result of the GPU task corresponding to the GPU interface.
[0121] Optionally, the GPU task information may include input data and an operation identifier of the GPU interface. Based on the operation identifier of the GPU interface, the GPU resource may determine which operation the GPU interface corresponds to for the GPU task, thereby executing an implementation program for the corresponding operation.
[0122] In this step, the CPU resource runs the implementation program corresponding to the operation identifier according to the input data and operation identifier of the GPU interface, completes the execution of the GPU task corresponding to the GPU interface, and obtains the execution result of the GPU task corresponding to the GPU interface.
[0123] Step S503: Return the GPU task execution result to the CPU device.
[0124] After obtaining the GPU task execution result, the GPU resource responds and returns the GPU task execution result to the CPU device.
[0125] In an optional embodiment, the GPU resource can be provided with a GPU software service that is connected to the CPU device. When the GPU resource is allocated / mounted to the CPU device, the GPU software service on the GPU resource will automatically start, thereby enabling data connection with the CPU device.
[0126] According to the solution of this embodiment, before the user needs to run the GPU application, he can mount the GPU resources through the CPU device he owns, and then run the GPU application on the CPU device without making any modifications to the GPU application. During the process of running the GPU application, when it is necessary to execute a GPU task, the CPU device sends the GPU task information to the mounted GPU resource. The GPU resource executes the GPU task and returns the GPU task execution result to the CPU device, and the CPU device can obtain the GPU task execution result. When the GPU application is not needed to run, the CPU device will unload the mounted GPU resources, and the GPU resources will be released, reducing the waste of GPU resources. By flexibly mounting and unloading GPU resources on the CPU device, mounting GPU resources on the CPU device when needed, and unloading the mounted GPU resources when not needed, the cost of using the GPU for users can be greatly reduced, and the waste of GPU resources can be reduced.
[0127] The present disclosure provides a GPU resource scheduling system, as shown in Figure 1. The system architecture includes: a CPU device, GPU resources, and a GPU resource scheduling server. A CPU device is a computer device with a CPU installed and used by a user. GPU resources refer to GPU devices provided by a GPU service platform, which provides GPU devices to users as mountable resources. The GPU service platform allocates available GPU resources to CPU devices (or users) requesting GPU resources through the GPU resource scheduling server.
[0128] The GPU resource scheduling server is used to: allocate GPU resources to the CPU device in response to the GPU resource mounting request of the CPU device, and send configuration information of the GPU resources allocated to the CPU device to the CPU device.
[0129] The CPU device is used to obtain the configuration information of the GPU resources allocated to the CPU device before running the GPU application, and mount the GPU resources according to the configuration information of the GPU resources.
[0130] The CPU device is also used to send GPU task information to the GPU resource when a GPU task needs to be executed during the running of the GPU application.
[0131] The GPU resource is used to: receive GPU task information sent by the CPU device, execute GPU tasks according to the GPU task information, obtain GPU task execution results, and return the GPU task execution results to the CPU device.
[0132] The CPU device is also used to receive the GPU task execution results returned by the GPU resource.
[0133] In this embodiment, the specific functions of the CPU device, GPU resources, and GPU resource scheduling server and the technical effects that can be achieved can be found in the aforementioned embodiments and will not be repeated here.
[0134] FIG6 is an interactive flow chart of GPU resource scheduling provided by this embodiment. As shown in FIG6 , based on the architecture of the GPU resource scheduling system shown in FIG1 , the interactive flow of GPU resource scheduling is as follows:
[0135] Step S601: The CPU device sends a GPU resource mounting request to the GPU resource scheduling server.
[0136] Step S602: The GPU resource scheduling server allocates mountable GPU resources to the CPU device in response to the GPU resource mounting request of the CPU device.
[0137] Step S603: The GPU resource scheduling server sends configuration information of the GPU resources allocated to the CPU device to the CPU device.
[0138] Step S604: The CPU device mounts the GPU resource according to the configuration information of the GPU resource.
[0139] Step S605: The CPU device runs the GPU application.
[0140] Step S606: When a GPU task needs to be executed, the CPU device sends GPU task information to the GPU resource.
[0141] Step S607: The GPU resource executes the GPU task according to the GPU task information and obtains a GPU task execution result.
[0142] Step S608: The GPU resource returns the GPU task execution result to the CPU device.
[0143] Step S609: After the GPU application is completed, the CPU device sends a GPU resource unloading request to the GPU resource scheduling server.
[0144] Step S610: The GPU resource scheduling server releases the GPU resources allocated to the CPU device.
[0145] In this embodiment, the specific functions of the CPU device, GPU resources, and GPU resource scheduling server and the technical effects that can be achieved can be found in the aforementioned embodiments and will not be repeated here.
[0146] Based on the GPU resource scheduling method and system provided in this disclosure, as shown in Figure 7, the process for a user to use a GPU is as follows: the user obtains a CPU device, deploys a GPU application on the CPU device, mounts the GPU resources on the GPU device, and runs the GPU application. During the GPU application's execution, the GPU interface is called, triggering a GPU task. The CPU device automatically intercepts the GPU task and sends the GPU task information to the mounted GPU resource for execution. After the GPU application completes, the CPU device unloads the GPU resources. The CPU device can then continue to execute other user applications.
[0147] The GPU resource scheduling method and system disclosed in the present invention realize the flexible mounting and unmounting of GPU resources on the CPU device. Before the user needs to run a GPU application, he mounts the GPU resources on the CPU device he owns, and then the GPU application can be run on the CPU device. In the process of running the GPU application, when it is necessary to execute a GPU task, the CPU device sends GPU task information to the mounted GPU resource, so that the corresponding GPU resource executes the GPU task and returns the GPU task execution result to the CPU device, and the CPU device can obtain the GPU task execution result. When it is not necessary to run the GPU application, the GPU resources can be unloaded from the CPU device to release the mounted GPU resources and reduce the waste of GPU resources. By mounting GPU resources on the CPU device when needed and unmounting the mounted GPU resources when not needed, the cost of using the GPU for the user can be greatly reduced and the waste of GPU resources can be reduced.
[0148] GPUs have powerful parallel computing capabilities and are used not only for graphics processing but also in computationally intensive tasks such as artificial intelligence. The GPU resource scheduling method and system provided herein can be applied to applications requiring extensive graphics processing (such as rendering), such as gaming, three-dimensional (3D) design, and graphics rendering. They can also be used in AI to accelerate data analysis and processing, deep learning training and inference, and the development of large language models.
[0149] For example, a user might develop a gaming application using a GPU-based programming language. The gaming application then needs to run on a GPU device. The user then mounts GPU resources on the CPU device that will run the gaming application, and runs the gaming application on the CPU device with the mounted GPU resources. If the gaming application goes offline or is not needed for an extended period, the CPU device can offload the GPU resources to avoid wasting GPU resources.
[0150] For example, taking deep learning algorithms or large language models as an example, users can use GPU-based programming languages to develop deep learning algorithms or large language models. Deep learning algorithms and large language models based on GPU programming languages need to be run on GPU devices. Users mount GPU resources on the CPU device that needs to run deep learning algorithms or large language models based on GPU programming languages, run deep learning algorithms or large language models on the CPU device with mounted GPU resources, and achieve parallel computing acceleration through the GPU. When the deep learning algorithm or large language model does not need to be run for a long period of time, the CPU device can unload GPU resources to avoid wasting GPU resources.
[0151] FIG8 is a schematic diagram of the structure of a CPU device provided in an embodiment of the present disclosure. As shown in FIG8 , the CPU device of this embodiment may include: at least one CPU 801; and a memory 802 in communication with the at least one CPU 801. The memory 802 stores instructions that can be executed by the at least one CPU 801, and the instructions are executed by the at least one CPU 801 to cause the CPU device to execute a method flow as executed by the CPU device in any of the above embodiments. The CPU device shown in FIG8 can be a cloud server or a local CPU device, and is not specifically limited here.
[0152] Optionally, the memory 802 may be independent or integrated with the CPU 801 .
[0153] The implementation principle and technical effects of the CPU device provided in this embodiment can be found in the aforementioned embodiments and will not be repeated here.
[0154] FIG9 is a schematic diagram of the structure of a GPU device provided in an embodiment of the present disclosure. As shown in FIG9 , the GPU device of this embodiment may include: at least one GPU 901, at least one CPU 902; and a memory 903 communicatively connected to the at least one GPU 901 and the at least one CPU 902. The memory 903 stores instructions executable by the at least one GPU 901 and the at least one CPU 902. When executed, the instructions cause the GPU device to execute the method flow for GPU resource execution as described in any of the above embodiments. The GPU device shown in FIG9 can be a cloud server or a local GPU device, and is not specifically limited here.
[0155] The implementation principle and technical effects of the GPU device provided in this embodiment can be found in the aforementioned embodiments and will not be repeated here.
[0156] FIG10 is a schematic diagram of the structure of a GPU resource scheduling server provided in an embodiment of the present disclosure. As shown in FIG10 , the GPU resource scheduling server includes: a memory 1001 and a processor 1002. The memory 1001 is used to store computer-executable instructions and can be configured to store various other data to support operations on the GPU resource scheduling server. The processor 1002 is in communication with the memory 1001 and is used to execute the computer-executable instructions stored in the memory 1001 to implement the method flow executed by the GPU resource scheduling server in any of the above-mentioned method embodiments. The specific functions and technical effects that can be achieved can be referred to in the aforementioned embodiments and will not be repeated here.
[0157] Optionally, as shown in Figure 10 , the GPU resource scheduling server further includes other components such as a firewall 1003, a load balancer 1004, a communication component 1005, and a power supply component 1006. Figure 10 only schematically illustrates some components, which does not mean that the GPU resource scheduling server only includes the components shown in Figure 10 .
[0158] The embodiments of the present disclosure further provide a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method executed by the CPU device in any of the aforementioned embodiments is implemented. The specific functions and technical effects that can be achieved are not repeated here.
[0159] The embodiments of the present disclosure further provide a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, the method for GPU resource execution in any of the aforementioned embodiments is implemented. The specific functions and technical effects that can be achieved are not further described here.
[0160] The embodiments of the present disclosure further provide a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method executed by the GPU resource scheduling server in any of the aforementioned embodiments is implemented. The specific functions and technical effects that can be achieved are not repeated here.
[0161] The present disclosure also provides a computer program product, including a computer program. When executed by a processor, the computer program implements the method performed by the CPU device in any of the aforementioned embodiments. The computer program is stored in a readable storage medium. At least one processor of an electronic device can read the computer program from the readable storage medium. The at least one processor executes the computer program, causing the electronic device to perform the method performed by the CPU device in any of the aforementioned method embodiments. The specific functions and technical effects achieved are not further described here.
[0162] The present disclosure also provides a computer program product, including a computer program. When executed by a processor, the computer program implements the method for GPU resource execution described in any of the aforementioned embodiments. The computer program is stored in a readable storage medium. At least one processor of an electronic device can read the computer program from the readable storage medium. The at least one processor executes the computer program, causing the electronic device to perform the method for GPU resource execution described in any of the aforementioned method embodiments. The specific functions and technical effects achieved are not further described herein.
[0163] The present disclosure also provides a computer program product, including a computer program. When executed by a processor, the computer program implements the method performed by the GPU resource scheduling server in any of the aforementioned embodiments. The computer program is stored in a readable storage medium. At least one processor of an electronic device can read the computer program from the readable storage medium. The at least one processor executes the computer program, causing the electronic device to perform the method performed by the GPU resource scheduling server in any of the aforementioned method embodiments. The specific functions and technical effects achieved are not further described here.
[0164] The present disclosure provides a chip comprising: a processing module and a communication interface. The processing module is capable of executing the technical solution of the CPU device in the aforementioned method embodiments. Optionally, the chip further comprises a storage module (e.g., a memory) configured to store instructions, and the processing module configured to execute the instructions stored in the storage module. Execution of the instructions stored in the storage module causes the processing module to execute the technical solution of the CPU device in any of the aforementioned method embodiments.
[0165] The present disclosure provides a chip comprising: a processing module and a communication interface. The processing module is capable of executing the technical solution for GPU resources in the aforementioned method embodiments. Optionally, the chip further comprises a storage module (e.g., a memory) configured to store instructions, and the processing module configured to execute the instructions stored in the storage module. Execution of the instructions stored in the storage module causes the processing module to execute the technical solution for GPU resources in any of the aforementioned method embodiments.
[0166] The present disclosure provides a chip comprising: a processing module and a communication interface. The processing module is capable of executing the technical solution of the GPU resource scheduling server described in the aforementioned method embodiments. Optionally, the chip further comprises a storage module (e.g., a memory) configured to store instructions, and the processing module configured to execute the instructions stored in the storage module. Execution of the instructions stored in the storage module causes the processing module to execute the technical solution of the GPU resource scheduling server described in any of the aforementioned method embodiments.
[0167] The integrated modules implemented in the form of software function modules can be stored in a computer-readable storage medium. The software function modules stored in a storage medium include several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform some of the steps of the methods of various embodiments of the present disclosure.
[0168] It should be understood that the above-mentioned processor can be a processing unit (Central Processing Unit, referred to as CPU), or it can be other general-purpose processors, digital signal processors (Digital Signal Processor, referred to as DSP), application specific integrated circuits (Application Specific Integrated Circuit, referred to as ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The memory may include high-speed random access memory (Random Access Memory, referred to as RAM), and may also include non-volatile storage, such as at least one disk storage, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disk, etc.
[0169] The above storage may be an object storage service (OSS).
[0170] The above-mentioned memory can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0171] The above-mentioned communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as a mobile hotspot (WiFi), a second-generation mobile communication system (2G), a third-generation mobile communication system (3G), a fourth-generation mobile communication system (4G) / Long Term Evolution (LTE), a fifth-generation mobile communication system (5G), or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth technology and other technologies.
[0172] The power supply assembly provides power to various components of the device in which the power supply assembly is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply assembly is located.
[0173] The storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0174] An exemplary storage medium is coupled to a processor, such that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an application-specific integrated circuit. Of course, the processor and storage medium can also exist as discrete components in an electronic device or a host control device.
[0175] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0176] The order of the above-mentioned embodiments of the present disclosure is for description only and does not represent the advantages and disadvantages of the embodiments. In addition, in some of the processes described in the above-mentioned embodiments and the accompanying drawings, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or in parallel. They are only used to distinguish between different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to different types. The meaning of "multiple" is more than two, unless otherwise clearly and specifically defined.
[0177] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present disclosure.
[0178] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein.
[0179] The above are only preferred embodiments of the present disclosure and are not intended to limit the patent scope of the present disclosure. Any equivalent structure or equivalent process transformation made using the contents of the present disclosure and the drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present disclosure.
Claims
1. A method for scheduling resources of a graphics processor (GPU), wherein: Applied to a central processing unit (CPU) device, the method comprises: Before running the GPU application, mounting the GPU resource; Run the GPU application, and when a GPU task needs to be executed, send GPU task information to the GPU resource, and receive the GPU task execution result returned by the GPU resource; When the GPU unloading condition is met, the GPU resources are unloaded.
2. The method according to claim 1, wherein: Before running the GPU application, mounting the GPU resource includes: In response to the GPU resource mounting instruction, obtaining configuration information of the GPU resources allocated to the CPU device; Mount the GPU resource according to the configuration information of the GPU resource.
3. The method according to claim 2, wherein: The obtaining of configuration information of GPU resources allocated to the CPU device includes: Sending a GPU resource mount request to a GPU resource scheduling server, wherein the GPU resource mount request is used to request the GPU resource scheduling server to allocate available GPU resources to the CPU device and return configuration information of the GPU resources allocated to the CPU device; Receive configuration information of the GPU resources allocated to the CPU device.
4. The method according to claim 2, wherein: The step of mounting the GPU resource according to the configuration information of the GPU resource includes: Mounting a software package that the GPU application depends on to a specified directory according to the configuration information of the GPU resources, wherein the software package includes an implementation method of the GPU interface; And configure the network address of the GPU resource.
5. The method according to any one of claims 1 to 4, wherein: When the GPU task needs to be executed, sending GPU task information to the GPU resource and receiving the GPU task execution result returned by the GPU resource include: In the process of running the GPU application, intercepting the calling request of the GPU application to the GPU interface, and obtaining GPU task information corresponding to the GPU interface, wherein the GPU task information includes input data and interface identification information of the GPU interface; Sending the GPU task information to the GPU resource; Receive the execution result of the GPU task corresponding to the GPU interface returned by the GPU resource.
6. The method according to any one of claims 1 to 4, wherein: When the GPU unloading condition is met, unloading the GPU resources includes: In response to receiving the GPU resource unloading instruction, a GPU resource unloading request is sent to the GPU resource scheduling server, wherein the GPU resource unloading request is used to request the GPU resource scheduling server to release the GPU resources allocated to the CPU device.
7. A GPU resource scheduling method, wherein: Applied to a GPU resource scheduling server, the method includes: In response to a GPU resource mounting request of a CPU device, allocating mountable GPU resources to the CPU device; Sending configuration information of the GPU resources allocated to the CPU device to the CPU device, wherein the configuration information is used by the CPU device to mount the GPU resources before running the GPU application, to use the GPU resources during the running of the GPU application, and to unload the GPU resources when a GPU unloading condition is met.
8. The method according to claim 7, wherein: Also includes: The CPU device is provided with a software package that the GPU application depends on for running.
9. The method according to claim 7 or 8, wherein: Also includes: In response to a GPU resource unloading request of a CPU device, the GPU resources allocated to the CPU device are released.
10. A GPU resource scheduling method, wherein: Applied to GPU resources, the method includes: Receive GPU task information sent by a CPU device, wherein the GPU task is information about a GPU task that needs to be executed when the CPU device runs a GPU application, and the GPU resource is mounted on the CPU device and will be unloaded by the CPU device when a GPU unloading condition is met; Execute the GPU task according to the GPU task information to obtain a GPU task execution result; Return the GPU task execution result to the CPU device.
11. The method according to claim 10, wherein: The receiving of GPU task information sent by the CPU device includes: Receive GPU task information corresponding to the GPU interface sent by the CPU device, where the GPU task information corresponding to the GPU interface includes input data and interface identification information of the GPU interface.
12. The method according to claim 11, wherein: The executing the GPU task according to the GPU task information to obtain the GPU task execution result includes: According to the input data and interface identification information of the GPU interface, the implementation method of the GPU interface is run to complete the execution of the GPU task corresponding to the GPU interface, and obtain the execution result of the GPU task corresponding to the GPU interface.
13. A GPU resource scheduling system, wherein: include: CPU devices, GPU resources and GPU resource scheduling servers, The GPU resource scheduling server is used to: in response to a GPU resource mounting request of a CPU device, allocate GPU resources to the CPU device, and send configuration information of the GPU resources allocated to the CPU device to the CPU device; The CPU device is used to: obtain configuration information of the GPU resource allocated to the CPU device, and mount the GPU resource according to the configuration information of the GPU resource; The CPU device is further used to: run the GPU application, and send GPU task information to the GPU resource when a GPU task needs to be executed; The GPU resource is used to: receive GPU task information sent by the CPU device, execute the GPU task according to the GPU task information, obtain the GPU task execution result, and return the GPU task execution result to the CPU device. fruit; The CPU device is also used to: receive the GPU task execution result returned by the GPU resource; The CPU device is also used to: unload the GPU resources.
14. A CPU device, wherein: include: At least one CPU; as well as a memory communicatively coupled to the at least one CPU; The memory stores instructions executable by the at least one CPU, and the instructions are executed by the at least one CPU so that the CPU device executes the method according to any one of claims 1 to 6.
15. A GPU device, wherein: include: At least one GPU, at least one CPU, and a memory communicatively coupled to the at least one GPU; The memory stores instructions that can be executed by the at least one GPU and the at least one CPU, and the instructions are executed by the at least one GPU and the at least one CPU so that the GPU device executes the method described in any one of claims 7-9.
16. A GPU resource scheduling server, wherein: include: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the GPU resource scheduling server executes the method described in any one of claims 10-12.
17. A computer storage medium, wherein: The computer storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
18. A computer program product, comprising a computer program, wherein the computer program is executed by a processor to implement the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Graphics processing resource allocation method and device, computer equipment and storage medium
CN110597635A
Resource configuration method, data processing method and device, equipment and storage medium
CN114924888A
Method, system and device for using GPU resources and medium
CN115114022A
Coordinated, topology-aware CPU-GPU-memory scheduling for containerized workloads
US20180276044A1
Methods and Devices for Virtualizing a Device Management Client in a Multi-Access Server Separate from a Device
US20210049032A1