Heterogeneous computing power resource management method and device, computer equipment and storage medium

By managing and scheduling heterogeneous computing resources of computing nodes on the control nodes, the problem of difficulty in calling GPU resources across nodes in the prior art is solved, the utilization rate of GPU resources is improved, and the diversified resource pooling needs are met.

CN119960962APending Publication Date: 2025-05-09CHINA TELECOM CLOUD TECH CO LTD
View PDF -1 Cites -1 Cited by

Patent Information

Application Number
CN202411795401.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-09

Smart Images

  • Figure CN119960962A_ABST
    Figure CN119960962A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cloud computing, and discloses a heterogeneous computing power resource management method and device, computer equipment and a storage medium, the method comprises the following steps: obtaining node information of each computing node, the computing nodes comprising a GPU physical machine, a straight-through GPU host machine and a virtualized GPU host machine; under the condition that a heterogeneous computing power resource demand sent by the client node is received, determining a target computing node according to the heterogeneous computing power resource demand and the node information; and creating a target heterogeneous computing power resource according to the heterogeneous computing power resource demand and the heterogeneous computing power resource in the target computing node, and allocating the target heterogeneous computing power resource to the client node. The problem that the utilization rate of all-resource GPUs is not high due to the fact that it is difficult to call GPU resources on any one or more GPU physical machines, the straight-through GPU host machine and the virtualized GPU host machine in a cross-node mode is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cloud computing technology, and in particular to a heterogeneous computing resource management method, device, computer equipment and storage medium. Background Art

[0002] As technology and demand continue to evolve, GPUs (Graphics Processing Units) have become general-purpose computing devices that handle computing tasks in the fields of AI (Artificial Intelligence) deep learning and scientific computing. Because GPUs are expensive, it is necessary to make full use of GPU resources as much as possible. Therefore, mainstream cloud vendors have launched GPU cloud host products to achieve GPU resource sharing, improve GPU resource utilization, and equip different network connection technologies according to customers' actual applications to meet customers' GPU resource needs.

[0003] Cloud vendors provide a variety of IaaS (Infrastructure as a Service) GPU products on public clouds, including GPU physical machines, GPU direct host machines, and GPU virtualized host machines. Sharing GPU resources requires GPU pooling first. However, in the current GPU pooling solution, the server programs of the same resource pool are all deployed on GPU products of the same IaaS product form, for example, all are GPU physical machines or GPU virtualized host machines. GPU resource pooling solutions based on multiple product forms have not been considered. It is difficult to call GPU resources on any one or more GPU physical machines, direct GPU host machines, and virtualized GPU host machines across nodes, resulting in low utilization of all GPU resources, which cannot meet the resource pooling requirements when GPU product forms are diversified in the public cloud resource pool.

[0004] Therefore, the related technology has the problem of difficulty in calling GPU resources on any one or more GPU physical machines, direct GPU host machines, and virtualized GPU host machines across nodes, resulting in low utilization of all GPU resources. Summary of the invention

[0005] In view of this, the present invention provides a heterogeneous computing resource management method, apparatus, computer equipment and storage medium to solve the problem that it is difficult to call GPU resources on any one or more GPU physical machines, direct GPU host machines, and virtualized GPU host machines across nodes, resulting in low utilization of all GPU resources.

[0006] In a first aspect, the present invention provides a heterogeneous computing resource management method, which is applied to a control node and includes:

[0007] Obtain node information of each computing node, where computing nodes include: GPU physical machines, direct GPU host machines, and virtualized GPU host machines;

[0008] When receiving the heterogeneous computing resource requirements sent by the client node, determine the target computing node according to the heterogeneous computing resource requirements and node information;

[0009] According to the heterogeneous computing power resource requirements and the heterogeneous computing power resources in the target computing node, create the target heterogeneous computing power resources and allocate the target heterogeneous computing power resources to the client node.

[0010] The heterogeneous computing power resource management method provided in this embodiment controls the node to obtain the node information of the computing node, and manages the heterogeneous computing power resources in each computing node according to the node information, so as to realize unified management and effective scheduling of the heterogeneous computing power resources in the resource pool. The control node creates the target heterogeneous computing power resources according to the heterogeneous computing power resource requirements sent by the client node, and allocates the target heterogeneous computing power resources to the client node, so that the client node can call the heterogeneous computing power resources in any computing node across nodes, thereby improving the utilization rate of GPU resources. It solves the problem that it is difficult to call the GPU resources on any one or more GPU physical machines, direct GPU host machines, and virtualized GPU host machines across nodes, resulting in low utilization rate of all resource GPUs.

[0011] In some optional implementations, determining a target computing node according to heterogeneous computing resource requirements and node information includes:

[0012] Determine the target graphics processor according to the target model, target software stack and network delay threshold for the graphics processor in the heterogeneous computing resource requirements, wherein the model of the target graphics processor is the target model, the software stack of the target graphics processor is the target software stack, and the network delay of the computing node where the target graphics processor is located is less than or equal to the network delay threshold;

[0013] Taking a computing node including a target graphics processor as a candidate computing node;

[0014] If the number of candidate computing nodes is less than a first preset threshold, taking the candidate computing nodes as target computing nodes;

[0015] If the number of candidate computing nodes is greater than or equal to a first preset threshold, obtaining a utilization rate of a graphics processor in the candidate computing node;

[0016] A target computing node is determined from among the candidate computing nodes according to utilization.

[0017] In this implementation, in the process of determining the target computing node, the software stack is added to better implement the scheduling of GPU resources from multiple manufacturers; the latency requirement is increased to make the scheduling more reasonable. In addition, the computing nodes are screened according to the GPU utilization rate, and the nodes with lower GPU utilization rate are given priority when the prerequisites are met, further ensuring the best user experience for global users in the resource pool.

[0018] In some optional implementations, creating a target heterogeneous computing resource according to the heterogeneous computing resource demand and the heterogeneous computing resources in the target computing node includes:

[0019] Determine the target computing power and target video memory based on the heterogeneous computing resource requirements;

[0020] When the number of target computing nodes is greater than or equal to the second preset threshold, determining intermediate heterogeneous computing power resources at each target computing node according to the target computing power, the target video memory, and the preset ratio;

[0021] Create a target heterogeneous computing resource based on the intermediate heterogeneous computing resource, where the computing power of the target heterogeneous computing resource is the target computing power, and the video memory of the target heterogeneous computing resource is the target video memory;

[0022] When the number of target computing nodes is less than a second preset threshold, target heterogeneous computing resources are determined in the target computing nodes.

[0023] In this embodiment, according to the number of target computing nodes, the target heterogeneous computing power resources are determined in the target computing nodes to ensure that the target heterogeneous computing power resources can meet the target computing power and target video memory required by the client nodes, so that the client nodes can process the target tasks normally.

[0024] In some optional embodiments, the method further comprises:

[0025] When a resource release request is received, the target heterogeneous computing resources are released.

[0026] In this embodiment, the control node determines that the client node no longer needs the target heterogeneous computing power resources based on the resource release request, and releases the target heterogeneous computing power resources in a timely manner to avoid wasting heterogeneous computing power resources and improve the utilization rate of heterogeneous computing power resources.

[0027] In a second aspect, the present invention provides a heterogeneous computing resource management method, which is applied to a client node and includes:

[0028] Obtain the heterogeneous computing resource requirements corresponding to the target task;

[0029] Send heterogeneous computing resource requirements to the control node;

[0030] When receiving the target heterogeneous computing power resources allocated by the control node, the target task is processed based on the target heterogeneous computing power resources, wherein the target heterogeneous computing power resources are obtained by the control node based on the heterogeneous computing power resource requirements and the heterogeneous computing power resources in the target computing node, and the target computing node is obtained based on the heterogeneous computing power resource requirements and the node information of each computing node.

[0031] In the heterogeneous computing power resource management method provided in this embodiment, the client node sends the heterogeneous computing power resource demand to the control node, and the control node creates the target heterogeneous computing power resource according to the heterogeneous computing power resource demand, and allocates the target heterogeneous computing power resource to the client node. The client node processes the target task according to the target heterogeneous computing power resource, realizes the cross-node call of the heterogeneous computing power resources in any computing node, and improves the GPU resource utilization. It solves the problem that it is difficult to call the GPU resources on any one or more GPU physical machines, direct GPU host machines, and virtualized GPU host machines across nodes, resulting in low utilization of all resource GPUs.

[0032] In some optional implementations, after processing the target task based on the target heterogeneous computing power resources, the method further includes:

[0033] When the preset indication information is received, a resource release request is sent to the control node, wherein the preset indication information is used to indicate that the target task has been processed.

[0034] In this embodiment, after the target task is processed, a resource release request is sent to the control node to release the target heterogeneous computing resources to avoid wasting heterogeneous computing resources and causing additional overhead.

[0035] In a third aspect, the present invention provides a heterogeneous computing resource management device, which is deployed on a control node and includes:

[0036] An information acquisition module is used to obtain node information of each computing node, wherein the computing nodes include: a GPU physical machine, a direct GPU host machine, and a virtualized GPU host machine;

[0037] A node determination module is used to determine a target computing node according to the heterogeneous computing resource requirements and node information when receiving the heterogeneous computing resource requirements sent by the client node;

[0038] The resource allocation module is used to create target heterogeneous computing resources according to the heterogeneous computing resource requirements and the heterogeneous computing resources in the target computing nodes, and allocate the target heterogeneous computing resources to the client nodes.

[0039] In a fourth aspect, the present invention provides a heterogeneous computing resource management device, which is deployed on a client node and includes:

[0040] The demand acquisition module is used to obtain the heterogeneous computing resource requirements corresponding to the target task;

[0041] The demand sending module is used to send heterogeneous computing resource requirements to the control node;

[0042] The task processing module is used to process the target task based on the target heterogeneous computing power resources when receiving the target heterogeneous computing power resources allocated by the control node, wherein the target heterogeneous computing power resources are obtained by the control node according to the heterogeneous computing power resource requirements and the heterogeneous computing power resources in the target computing node, and the target computing node is obtained according to the heterogeneous computing power resource requirements and the node information of each computing node.

[0043] In a fifth aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the heterogeneous computing power resource management method of the above-mentioned first aspect or any corresponding embodiment thereof by executing the computer instructions, or executes the heterogeneous computing power resource management method of the above-mentioned second aspect or any corresponding embodiment thereof.

[0044] In a sixth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the heterogeneous computing power resource management method of the above-mentioned first aspect or any corresponding embodiment thereof, or to execute the heterogeneous computing power resource management method of the above-mentioned second aspect or any corresponding embodiment thereof.

[0045] In the seventh aspect, the present invention provides a computer program product, including computer instructions, which are used to enable a computer to execute the heterogeneous computing power resource management method of the above-mentioned first aspect or any corresponding embodiment thereof, or to execute the heterogeneous computing power resource management method of the above-mentioned second aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the related technologies, the drawings required for use in the specific embodiments or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0047] Figure 1 It is a flowchart of a heterogeneous computing resource management method applied to a control node according to an embodiment of the present invention;

[0048] Figure 2 is a remote call logic structure diagram according to an embodiment of the present invention;

[0049] Figure 3 is a flow chart of another heterogeneous computing resource management method applied to a control node according to an embodiment of the present invention;

[0050] Figure 4 It is a flowchart of a heterogeneous computing resource management method applied to a client node according to an embodiment of the present invention;

[0051] Figure 5 is a structural block diagram of a heterogeneous computing power resource management device deployed on a control node according to an embodiment of the present invention;

[0052] Figure 6 is a structural block diagram of a heterogeneous computing power resource management device deployed on a client node according to an embodiment of the present invention;

[0053] Figure 7 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0055] Combined with the application scenarios on which the execution of the heterogeneous computing resource management method depends, the application scenarios are described here. Graphics processing units (GPUs) have more computing units and simpler control units, and are suitable for various scenarios involving a large number of parallel calculations. In order to enable different users to share different graphics card resources on the same GPU physical machine or different computing power / video memory resources on the same graphics card, GPU virtualization is required. Through virtualization, remote calling, elastic scaling and other technologies, the GPUs in the resource pool are grouped into a pool, and the GPU resources in the pool are dynamically managed to achieve on-demand free scheduling. The technology of GPU pooling can effectively improve the GPU utilization of the entire resource. Among them, remote calling is a technology widely used in distributed systems for local clients to call processes on remote servers through the network.

[0056] The above-mentioned virtualization scheme can be divided into three levels, namely the user layer, the kernel layer and the hardware layer. The GPU virtualization technology currently used by the GPU cloud host involves two types: the kernel layer and the hardware layer. For example, GPU virtualization belonging to the hardware layer: direct GPU cloud host technology directly maps the physical GPU to the virtual machine. However, this method can only realize the sharing of different graphics card resources on the same GPU physical machine by different users, and cannot perform more fine-grained segmentation of the graphics card. For example, GPU virtualization belonging to the hardware layer: virtualized GPU cloud host technology requires the installation of specific drivers and management software. Although the same physical GPU can be allocated to multiple users for use, the video memory in this virtualization method is statically partitioned, and the computing power of different GPU cloud hosts can only be time-sharing multiplexed, which is not flexible enough, and the license authorization fee is expensive. In addition, when GPU pooling is realized through remote call technology, it is necessary to install a server program for remotely calling the GPU on each GPU physical node in the resource pool to discover and manage the GPU resource pool. Install a client program for remotely calling the GPU on an ordinary cloud host or GPU cloud host running an AI application to apply for or release GPU resources. However, in the current GPU pooling solution, the server-side programs of the same resource pool are all deployed on GPU products of the same IaaS product form, for example, they are all GPU physical machines or GPU virtualized host machines, and GPU pooling solutions based on multiple IaaS product forms have not been considered. In addition, the reasons for the low current GPU utilization rate also include: the configurations of GPU physical machines and GPU host machines are different, and the number of GPUs, number of CPUs, and memory size on a server are not fixed, resulting in the inability to fully utilize the GPU resources on the server when opening a standard specification GPU physical machine or GPU cloud host. In different application scenarios, customers have different requirements for CPU, memory, and GPU resources. Some enterprise users require the opening of various non-standard specifications to reduce expenses, which further leads to the generation of fragmented GPU resources in the public cloud resource pool. The host deployment ratio of direct-pass and virtualized GPU cloud hosts cannot be adjusted dynamically, but the services of different resource pools have different requirements for direct-pass and virtualized GPU cloud hosts, resulting in low GPU utilization.

[0057] Based on the above content, an embodiment of the present invention provides a heterogeneous computing power resource management method, which combines the GPU pooling technology based on remote calls with the actual scenario of the public cloud, and installs a server-side program on each GPU physical node in the resource pool. The control node obtains the node information of the GPU physical node through the server-side program, and manages the GPU physical node according to the node information. A virtual machine with a client program installed can send heterogeneous computing power resource requirements to the control node, and the control node can allocate GPU resources to the virtual machine, so that the virtual machine can call the GPU resources on any one or more GPU physical machines, GPU direct host machines, and GPU virtualized host machines in the resource pool that have installed the server-side program across nodes, which can effectively improve the GPU resource utilization of the public cloud resource pool. In order to achieve the effect of being able to call the GPU resources on any one or more GPU physical machines, GPU direct host machines, and GPU virtualized host machines across nodes, and improve the GPU resource utilization in the current public cloud resource pool.

[0058] According to an embodiment of the present invention, a heterogeneous computing resource management embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, for example: a computer, a server, etc., and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0059] In this embodiment, a heterogeneous computing resource management method is provided, which is applied to a control node. Figure 1 is a flow chart of a heterogeneous computing resource management method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0060] Step S101, obtaining node information of each computing node, wherein the computing nodes include: a GPU physical machine, a direct GPU host machine, and a virtualized GPU host machine.

[0061] Specifically, the resource pools of mainstream cloud vendors have deployed a variety of IaaS GPU products and equipped with a variety of networks to meet scenarios with high demand for GPU resources and high requirements for storage network latency. GPU cloud hosts meet scenarios with relatively small demand for GPU resources and higher requirements for resource flexibility. However, the current GPU pooling solution only considers that the server-side programs in the same resource pool are deployed on the same GPU physical node, which cannot meet the resource pooling needs when the GPU product forms are diversified in the public cloud resource pool.

[0062] For the resource pool where multiple IaaS product forms of GPU products have been deployed, this embodiment combines the GPU pooling technology based on remote calls with the actual scenarios of public clouds to achieve heterogeneous computing resource management.

[0063] Compute nodes include: physical machine nodes, GPU direct host nodes, GPU virtualized host nodes, such as Figure 2 As shown, the physical machine nodes include x86 physical machine nodes and arm physical machine nodes, the x86 physical machine nodes include GPUs of model 1 or GPUs of model 2, and the arm physical machine nodes include GPUs of model 3; the GPU direct host machine nodes include some x86 host machine nodes and arm host machine nodes, the x86 host machine nodes include GPUs of model 1 or GPUs of model 2, and the arm host machine nodes include GPUs of model 3; the GPU virtualization host machine nodes include the remaining x86 host machine nodes and arm host machine nodes, the x86 host machine nodes include GPUs of model 1 or GPUs of model 2, and the arm host machine nodes include GPUs of model 3. Install a server program for remotely calling the GPU on each computing node in the resource pool.

[0064] Create a common cloud host or GPU cloud host in the host machine to deploy application services. The application service is the target task to be processed by the client node, and install the client program for remotely calling heterogeneous computing resources in the form of Binary. The common cloud host or GPU cloud host becomes the client node. Target tasks include: AI applications such as AI reasoning business and AI data processing business. Among them, installing the client program in Binary form means installing the software through pre-compiled binary files instead of compiling and installing from source code. Heterogeneous computing resources include: GPU resources. Figure 2 As shown, create a cloud host / GPU cloud host, install a client program such as GPU Client program in it, and use it as a client node. The client node is connected to the control node, computing node, etc. through the network.

[0065] Create a common cloud host as a control node, and connect the client programs of all client nodes and the server programs of all computing nodes to this control node through the network. The control node obtains the node information of each computing node through the server program, such as the IP addresses of all nodes in the resource pool, the used and idle GPU information, etc. The control node manages the heterogeneous computing resources in each computing node according to the node information, and realizes unified management and effective scheduling of the heterogeneous computing resources in the resource pool. Networks such as ordinary Ethernet, InfiniBand (infinite bandwidth), RoCE (RDMA over Converged Ethernet, remote direct memory access based on converged Ethernet), etc. RoCE is a network protocol that allows remote direct memory access (RDMA) on Ethernet.

[0066] Step S102, when receiving the heterogeneous computing resource requirements sent by the client node, determine the target computing node according to the heterogeneous computing resource requirements and node information.

[0067] Specifically, heterogeneous computing resource requirements include, for example, GPU requirement information submitted by the client program in the client node. The format of heterogeneous computing resource requirements is, for example, {GPU model, computing power, video memory, software stack, network latency requirement}.

[0068] After the control node receives the heterogeneous computing resource requirements sent by the client node, it determines the client node's GPU model requirements, computing power requirements, video memory requirements, software stack requirements, network latency requirements, etc. based on the heterogeneous computing resource requirements. According to the heterogeneous computing resource requirements and node information, the control node screens each computing node to find the node that can meet the above heterogeneous computing resource requirements and uses it as the target computing node. In addition, the GPU resource costs corresponding to different heterogeneous computing resource requirements submitted by the client node are different.

[0069] Step S103, creating target heterogeneous computing resources according to the heterogeneous computing resource requirements and the heterogeneous computing resources in the target computing node, and allocating the target heterogeneous computing resources to the client node.

[0070] Specifically, according to the computing power requirements and video memory requirements in the heterogeneous computing power resource requirements, determine the heterogeneous computing power resources that need to be occupied in the target computing node. For example, the computing power of the heterogeneous computing power resources that need to be occupied is 16TFLOPS / FP32 and 128TFLOPS / FP16, where 16TFLOPS / FP32 means that the heterogeneous computing power resources can achieve 16 trillion floating-point operations per second in single-precision floating-point operation mode, and FP32 refers to single-precision floating-point numbers; 128TFLOPS / FP16 means that the heterogeneous computing power resources can achieve 128 trillion floating-point operations per second in half-precision floating-point operation mode, and FP16 refers to half-precision floating-point numbers. The video memory of the heterogeneous computing power resources that need to be occupied is 32GB.

[0071] The control node creates target heterogeneous computing resources according to the heterogeneous computing resources required to be occupied as mentioned above, and after the creation is completed, allocates the target heterogeneous computing resources to the client nodes, and the client nodes use the target heterogeneous computing resources to process the target tasks.

[0072] The heterogeneous computing power resource management method provided in this embodiment controls the node to obtain the node information of the computing node, and manages the heterogeneous computing power resources in each computing node according to the node information, so as to realize unified management and effective scheduling of the heterogeneous computing power resources in the resource pool. The control node creates the target heterogeneous computing power resources according to the heterogeneous computing power resource requirements sent by the client node, and allocates the target heterogeneous computing power resources to the client node, so that the client node can call the heterogeneous computing power resources in any computing node across nodes, thereby improving the utilization rate of GPU resources. It solves the problem that it is difficult to call the GPU resources on any one or more GPU physical machines, direct GPU host machines, and virtualized GPU host machines across nodes, resulting in low utilization rate of all resource GPUs.

[0073] In this embodiment, another heterogeneous computing resource management method is provided, which can be used for the above-mentioned control node. Figure 3 is a flow chart of another heterogeneous computing resource management method according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:

[0074] Step S301, obtain node information of each computing node, wherein computing nodes include: GPU physical machine, direct GPU host machine and virtualized GPU host machine. For details, please refer to Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.

[0075] Step S302, when receiving the heterogeneous computing resource requirements sent by the client node, determine the target computing node according to the heterogeneous computing resource requirements and node information.

[0076] Specifically, the above step S302 includes:

[0077] Step S3021, determining a target graphics processor according to a target model, a target software stack, and a network delay threshold for the graphics processor in the heterogeneous computing power resource requirements, wherein the model of the target graphics processor is the target model, the software stack of the target graphics processor is the target software stack, and the network delay of the computing node where the target graphics processor is located is less than or equal to the network delay threshold.

[0078] The format of heterogeneous computing resource requirements is, for example, {GPU model, computing power, video memory, software stack, network latency requirements}. Based on the heterogeneous computing resource requirements, the client node can determine the target model, target software stack, and network latency threshold for the graphics processor. The heterogeneous computing resource requirements received by the control node are, for example, {Model 1, 16TFLOPS / FP32, 128TFLOPS / FP16, 32GB, TensorFlow, lower}. It can be seen that the target model is Model 1 and the target software stack is TensorFlow. If the network delay requirement is low, the corresponding relationship between the network delay requirement and the network delay threshold can be set. For example, the network delay threshold corresponding to the low network delay requirement is 100ms. As long as the network delay is less than or equal to 100ms, the network delay requirement can be met; the network delay threshold corresponding to the moderate network delay is 200ms. As long as the network delay is less than or equal to 200ms, the network delay requirement can be met; the network delay threshold corresponding to the high network delay is 300ms. As long as the network delay is less than or equal to 300ms, the network delay requirement can be met. Therefore, the network delay threshold is determined to be 100ms according to the network delay requirement.

[0079] According to the above target model, target software stack and network delay threshold, the target graphics processor is determined. The model of the target graphics processor is model 1, the software stack is TensorFlow, and the network delay of the computing node where the target graphics processor is located is less than or equal to 100ms.

[0080] Step S3022: taking the computing node including the target GPU as a candidate computing node.

[0081] The GPU in the computing node is matched with the target graphics processor, and the computing node including the target graphics processor is determined as a candidate computing node.

[0082] It should be noted that in the process of matching GPUs, the corresponding computing nodes can be matched according to the network latency requirements in the heterogeneous computing power resource requirements submitted by the client program. If the network latency requirement is extremely high, the computing nodes near the server nodes will be prioritized for scheduling, and the allocated GPU resources will also prioritize the combination with smaller network latency, for example: some bare metal nodes support RDMA network latency. In addition, matching priority rules can also be set. For example, if the network latency requirement is low, the target host where the client node is located is determined first, and the computing nodes on the target host are matched first, because if the client node and the computing node are in the same host, the network latency between the two is low; if the network latency requirement is moderate, the target pod (cluster) where the client node is located is determined first, and the computing nodes on the target pod are matched first; if the network latency requirement is high, the computing nodes on pods other than the target pod are matched first.

[0083] Step S3023: If the number of candidate computing nodes is less than a first preset threshold, the candidate computing nodes are used as target computing nodes.

[0084] The first preset threshold is, for example, 2, 3, etc. The specific value is set according to actual needs. If the number of candidate computing nodes is less than the first preset threshold, it means that the number of candidate computing nodes is small, and there is no need to screen the candidate computing nodes, and the candidate computing nodes are directly used as target computing nodes.

[0085] Step S3024: If the number of candidate computing nodes is greater than or equal to the first preset threshold, the utilization rate of the graphics processor in the candidate computing node is obtained.

[0086] If the number of candidate computing nodes is greater than or equal to the first preset threshold, it means that the number of candidate computing nodes is large and the candidate computing nodes need to be screened.

[0087] The candidate computing nodes may be screened according to the utilization of the graphics processors in the candidate computing nodes. Therefore, the utilization of the graphics processors in the candidate computing nodes is firstly obtained.

[0088] Step S3025: determine a target computing node from the candidate computing nodes according to the utilization rate.

[0089] The priority of the computing nodes is determined according to the utilization of the graphics processors in the computing nodes, and the scheduling is performed according to the computing power / video memory information submitted by the client. For example, if the GPU utilization of computing node 1 is 60%, the GPU utilization of computing node 2 is 50%, and the GPU utilization of computing node 3 is 70%, then the priority is computing node 3, computing node 1, and computing node 2. If computing node 3 can meet the heterogeneous computing power resource requirements, computing node 3 will be used as the target computing node; if computing node 3 cannot meet the requirements, but computing node 3 and computing node 1 can meet the heterogeneous computing power resource requirements together, computing node 3 and computing node 1 will be used as the target computing nodes; if it still cannot be met, but computing node 3, computing node 1 and computing node 2 can meet the heterogeneous computing power resource requirements together, they will all be used as target computing nodes, and so on.

[0090] In this implementation, in the process of determining the target computing node, the software stack is added to better implement the scheduling of GPU resources from multiple manufacturers; the latency requirement is increased to make the scheduling more reasonable. In addition, the computing nodes are screened according to the GPU utilization rate, and the nodes with lower GPU utilization rate are given priority when the prerequisites are met, further ensuring the best user experience for global users in the resource pool.

[0091] Step S303: Create target heterogeneous computing resources according to the heterogeneous computing resource requirements and the heterogeneous computing resources in the target computing node, and allocate the target heterogeneous computing resources to the client node. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.

[0092] In some optional implementations, the above step S303 includes:

[0093] Step a1: Determine the target computing power and target video memory based on the heterogeneous computing resource requirements.

[0094] Step a2, when the number of target computing nodes is greater than or equal to the second preset threshold, determine the intermediate heterogeneous computing power resources at each target computing node according to the target computing power, target video memory and the preset ratio.

[0095] Step a3: Create a target heterogeneous computing resource based on the intermediate heterogeneous computing resource, wherein the computing power of the target heterogeneous computing resource is the target computing power, and the video memory of the target heterogeneous computing resource is the target video memory.

[0096] Step a4: when the number of target computing nodes is less than a second preset threshold, determine the target heterogeneous computing power resources in the target computing nodes.

[0097] Specifically, the target computing power and target video memory are determined according to the heterogeneous computing power resource requirements. For example, if the heterogeneous computing power resource requirements are {model 1, 16TFLOPS / FP32, 128TFLOPS / FP16, 32GB, TensorFlow, lower}, the target computing power is 16TFLOPS / FP32, 128TFLOPS / FP16; the target video memory is 32GB.

[0098] The second preset threshold is, for example, 2, 3... The specific value is set according to actual needs. Take the second preset threshold equal to 2 as an example for illustration. When the number of target computing nodes is greater than or equal to the second preset threshold, it is determined that there are multiple target computing nodes. The preset ratio is, for example, if there are 2 target computing nodes, the preset ratio is 1:1, if there are 3 target computing nodes, the preset ratio is 1:2:1, etc. The specific ratio can be set according to actual needs. According to the target computing power, target video memory and the preset ratio, the intermediate heterogeneous computing power resources are determined for each target computing node. For example, if there are 2 target computing nodes, the preset ratio is 1:1, the target computing power is 16TFLOPS / FP32, 128TFLOPS / FP16, and the target video memory is 32GB, then each computing node occupies an intermediate heterogeneous computing power resource of 8TFLOPS / FP32, 64TFLOPS / FP16 computing power and 16GB video memory.

[0099] Create target heterogeneous computing resources based on intermediate heterogeneous computing resources. The computing power of the target heterogeneous computing resources is the target computing power, and the video memory of the target heterogeneous computing resources is the target video memory.

[0100] When the number of target computing nodes is less than the second preset threshold, it means that there is only one target computing node, and the heterogeneous computing power resources are directly determined in the target computing node. For example, if the target computing power is 16TFLOPS / FP32, 128TFLOPS / FP16, and the target video memory is 32GB, then the computing power of the target heterogeneous computing power resources determined in the target computing node is 16TFLOPS / FP32, 128TFLOPS / FP16, and the video memory is 32GB.

[0101] In this embodiment, according to the number of target computing nodes, the target heterogeneous computing power resources are determined in the target computing nodes to ensure that the target heterogeneous computing power resources can meet the target computing power and target video memory required by the client nodes, so that the client nodes can process the target tasks normally.

[0102] This embodiment provides another heterogeneous computing resource management method, which optimizes the scheduling strategy of the control node for scheduling GPU resources, adds consideration of the software stack in the information submitted by the client that remotely calls the GPU, and better implements the scheduling of GPU resources from multiple manufacturers; at the same time, it increases the network latency requirement, and charges different fees according to the different network latency requirements of different clients, making the scheduling more reasonable. In the remotely called control node, the available GPU resources are sorted according to the GPU utilization rate, and nodes with lower GPU utilization rate are given priority when the prerequisites are met, further ensuring the best user experience for global users in the resource pool.

[0103] In some optional embodiments, the method further comprises:

[0104] When a resource release request is received, the target heterogeneous computing resources are released.

[0105] Specifically, after receiving the resource release request, the control node determines that the client node no longer needs the target heterogeneous computing power resources. In order to avoid wasting heterogeneous computing power resources, the target heterogeneous computing power resources allocated to the client node are released.

[0106] In this embodiment, the control node determines that the client node no longer needs the target heterogeneous computing power resources based on the resource release request, and releases the target heterogeneous computing power resources in a timely manner to avoid wasting heterogeneous computing power resources and improve the utilization rate of heterogeneous computing power resources.

[0107] According to an embodiment of the present invention, an embodiment of a heterogeneous computing resource management method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a client node, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0108] In this embodiment, a heterogeneous computing resource management method is provided, which is applied to client nodes. Figure 4 is a flow chart of a heterogeneous computing resource management method according to an embodiment of the present invention. Figure 4 As shown, the process includes the following steps:

[0109] Step S401, obtaining the heterogeneous computing resource requirements corresponding to the target task.

[0110] Specifically, a common cloud host or GPU cloud host is created in the host machine to deploy application services, which are the target tasks to be processed by the client node, and a client program for remotely calling heterogeneous computing resources is installed in it in binary form. The common cloud host or GPU cloud host becomes the client node. Target tasks include AI applications such as AI reasoning services and AI data processing services. Heterogeneous computing resources include GPU resources. Figure 2 As shown, create a cloud host / GPU cloud host, install a client program such as GPU Client program in it, and use it as a client node. The client node is connected to the control node, computing node, etc. through the network.

[0111] The client node obtains the heterogeneous computing resource requirements of the target task. The heterogeneous computing resource requirements are, for example, the GPU requirement information submitted by the client program in the client node. The format of the heterogeneous computing resource requirements is, for example, {GPU model, computing power, video memory, software stack, network latency requirement}.

[0112] Step S402, sending the heterogeneous computing resource requirements to the control node.

[0113] Step S403, upon receiving the target heterogeneous computing power resources allocated by the control node, processing the target task based on the target heterogeneous computing power resources, wherein the target heterogeneous computing power resources are obtained by the control node based on the heterogeneous computing power resource requirements and the heterogeneous computing power resources in the target computing node, and the target computing node is obtained based on the heterogeneous computing power resource requirements and the node information of each computing node.

[0114] Specifically, after the control node receives the heterogeneous computing power resource requirements sent by the client node, it determines the client node's GPU model requirements, computing power requirements, video memory requirements, software stack requirements, network latency requirements, etc. according to the heterogeneous computing power resource requirements. According to the heterogeneous computing power resource requirements and node information, screen each computing node, find out the node that can meet the above heterogeneous computing power resource requirements, and use it as the target computing node. According to the computing power requirements and video memory requirements in the heterogeneous computing power resource requirements, determine the heterogeneous computing power resources that need to be occupied in the target computing node, for example: the computing power of the heterogeneous computing power resources that need to be occupied is 16TFLOPS / FP32, 128TFLOPS / FP16; the video memory of the heterogeneous computing power resources that need to be occupied is 32GB. The control node creates the target heterogeneous computing power resources according to the above heterogeneous computing power resources that need to be occupied, and allocates the target heterogeneous computing power resources to the client node after the creation is completed.

[0115] After receiving the target heterogeneous computing resources allocated by the control node, the client node uses the target heterogeneous computing resources to process the target task.

[0116] In the heterogeneous computing power resource management method provided in this embodiment, the client node sends the heterogeneous computing power resource demand to the control node, and the control node creates the target heterogeneous computing power resource according to the heterogeneous computing power resource demand, and allocates the target heterogeneous computing power resource to the client node. The client node processes the target task according to the target heterogeneous computing power resource, realizes the cross-node call of the heterogeneous computing power resources in any computing node, and improves the GPU resource utilization. It solves the problem that it is difficult to call the GPU resources on any one or more GPU physical machines, direct GPU host machines, and virtualized GPU host machines across nodes, resulting in low utilization of all resource GPUs.

[0117] In some optional implementations, after processing the target task based on the target heterogeneous computing power resources, the method further includes:

[0118] When the preset indication information is received, a resource release request is sent to the control node, wherein the preset indication information is used to indicate that the target task has been processed.

[0119] Specifically, after the target task is processed, preset indication information can be sent to the client program of the client node. For example, if the target task is an AI reasoning service, after the AI ​​reasoning service is completed, the preset indication information can be entered in the client program to terminate the target task.

[0120] When the client node receives the above-mentioned preset indication information, it determines that the target heterogeneous computing power resources are no longer needed. In order to avoid wasting heterogeneous computing power resources and causing additional overhead, the client node sends a resource release request to the control node. After receiving the resource release request, the control node releases the target heterogeneous computing power resources.

[0121] In this embodiment, after the target task is processed, a resource release request is sent to the control node to release the target heterogeneous computing resources to avoid wasting heterogeneous computing resources and causing additional overhead.

[0122] In this embodiment, a heterogeneous computing power resource management device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated here. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0123] This embodiment provides a heterogeneous computing resource management device, which is deployed on a control node, such as Figure 5 As shown, including:

[0124] The information acquisition module 501 is used to acquire node information of each computing node, wherein the computing nodes include: a GPU physical machine, a direct GPU host machine, and a virtualized GPU host machine;

[0125] The node determination module 502 is used to determine the target computing node according to the heterogeneous computing resource requirements and node information when receiving the heterogeneous computing resource requirements sent by the client node;

[0126] The resource allocation module 503 is used to create target heterogeneous computing resources according to the heterogeneous computing resource requirements and the heterogeneous computing resources in the target computing node, and allocate the target heterogeneous computing resources to the client node.

[0127] In some optional implementations, the node determination module 502 includes:

[0128] A first determining unit is used to determine a target graphics processor according to a target model, a target software stack, and a network delay threshold for the graphics processor in the heterogeneous computing power resource requirements, wherein the model of the target graphics processor is the target model, the software stack of the target graphics processor is the target software stack, and the network delay of the computing node where the target graphics processor is located is less than or equal to the network delay threshold;

[0129] A first setting unit, configured to use a computing node including a target graphics processor as a candidate computing node;

[0130] a second setting unit, configured to use the candidate computing nodes as target computing nodes if the number of the candidate computing nodes is less than a first preset threshold;

[0131] an acquisition unit, configured to acquire a utilization rate of a graphics processor in a candidate computing node if the number of the candidate computing nodes is greater than or equal to a first preset threshold;

[0132] The second determining unit is used to determine a target computing node from among the candidate computing nodes according to the utilization rate.

[0133] In some optional implementations, the resource allocation module 503 includes:

[0134] A third determination unit is used to determine the target computing power and target video memory according to the heterogeneous computing power resource requirements;

[0135] a fourth determining unit, configured to determine, when the number of target computing nodes is greater than or equal to a second preset threshold, intermediate heterogeneous computing power resources at each target computing node according to the target computing power, the target video memory, and the preset ratio;

[0136] A creation unit, used to create a target heterogeneous computing resource according to the intermediate heterogeneous computing resource, wherein the computing power of the target heterogeneous computing resource is the target computing power, and the video memory of the target heterogeneous computing resource is the target video memory;

[0137] The fifth determination unit is used to determine the target heterogeneous computing power resources in the target computing nodes when the number of the target computing nodes is less than the second preset threshold.

[0138] In some optional embodiments, the device further comprises:

[0139] The resource release module is used to release the target heterogeneous computing resources when a resource release request is received.

[0140] This embodiment provides a heterogeneous computing resource management device, which is deployed on a client node, such as Figure 6 As shown, including:

[0141] The demand acquisition module 601 is used to obtain the heterogeneous computing resource requirements corresponding to the target task;

[0142] The demand sending module 602 is used to send the heterogeneous computing resource demand to the control node;

[0143] The task processing module 603 is used to process the target task based on the target heterogeneous computing power resources when receiving the target heterogeneous computing power resources allocated by the control node, wherein the target heterogeneous computing power resources are obtained by the control node based on the heterogeneous computing power resource requirements and the heterogeneous computing power resources in the target computing node, and the target computing node is obtained based on the heterogeneous computing power resource requirements and the node information of each computing node.

[0144] In some optional embodiments, the device further comprises:

[0145] The resource release request module is used to send a resource release request to the control node when receiving preset indication information, wherein the preset indication information is used to indicate that the target task has been processed.

[0146] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0147] The heterogeneous computing resource management device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0148] The embodiment of the present invention also provides a computer device having the above Figure 5 and Figure 6 The heterogeneous computing resource management device shown.

[0149] See also Figure 7 , Figure 7 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 7 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 7 A processor 10 is taken as an example.

[0150] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include an integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable logic gate array, a general purpose array logic or any combination thereof.

[0151] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0152] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0153] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0154] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0155] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0156] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0157] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined in this application.

Claims

1. A heterogeneous computing resource management method, characterized in that: The method is applied to a control node, and the method comprises: Obtain node information of each computing node, wherein the computing nodes include: a GPU physical machine, a direct GPU host machine, and a virtualized GPU host machine; Upon receiving the heterogeneous computing resource requirements sent by the client node, determining the target computing node according to the heterogeneous computing resource requirements and the node information; According to the heterogeneous computing power resource demand and the heterogeneous computing power resources in the target computing node, a target heterogeneous computing power resource is created, and the target heterogeneous computing power resource is allocated to the client node.

2. The method according to claim 1, characterized in that The determining the target computing node according to the heterogeneous computing resource requirements and the node information includes: Determine a target graphics processor according to a target model, a target software stack, and a network delay threshold for a graphics processor in the heterogeneous computing power resource requirements, wherein the model of the target graphics processor is the target model, the software stack of the target graphics processor is the target software stack, and the network delay of a computing node where the target graphics processor is located is less than or equal to the network delay threshold; Taking a computing node including the target graphics processor as a candidate computing node; If the number of the candidate computing nodes is less than a first preset threshold, taking the candidate computing nodes as the target computing nodes; If the number of the candidate computing nodes is greater than or equal to the first preset threshold, obtaining the utilization rate of the graphics processor in the candidate computing node; The target computing node is determined from among the candidate computing nodes according to the utilization rate.

3. The method according to claim 2, characterized in that The creating a target heterogeneous computing power resource according to the heterogeneous computing power resource demand and the heterogeneous computing power resources in the target computing node includes: Determine the target computing power and target video memory according to the heterogeneous computing power resource requirements; When the number of the target computing nodes is greater than or equal to a second preset threshold, determining intermediate heterogeneous computing power resources at each of the target computing nodes according to the target computing power, the target video memory, and the preset ratio; Creating the target heterogeneous computing power resource according to the intermediate heterogeneous computing power resource, wherein the computing power of the target heterogeneous computing power resource is the target computing power, and the video memory of the target heterogeneous computing power resource is the target video memory; When the number of the target computing nodes is less than the second preset threshold, the target heterogeneous computing power resources are determined in the target computing nodes.

4. The method according to claim 1, characterized in that The method further comprises: When a resource release request is received, the target heterogeneous computing resources are released.

5. A heterogeneous computing resource management method, characterized in that: The method is applied to a client node, and the method comprises: Obtain the heterogeneous computing resource requirements corresponding to the target task; Sending the heterogeneous computing power resource requirements to the control node; Upon receiving the target heterogeneous computing power resources allocated by the control node, the target task is processed based on the target heterogeneous computing power resources, wherein the target heterogeneous computing power resources are obtained by the control node based on the heterogeneous computing power resource requirements and the heterogeneous computing power resources in the target computing node, and the target computing node is obtained based on the heterogeneous computing power resource requirements and the node information of each computing node.

6. The method according to claim 5, characterized in that After processing the target task based on the target heterogeneous computing power resources, the method further includes: In case of receiving the preset indication information, a resource release request is sent to the control node, wherein the preset indication information is used to indicate that the processing of the target task is completed.

7. A heterogeneous computing resource management device, characterized in that: The device is deployed on a control node, and includes: An information acquisition module, used to acquire node information of each computing node, wherein the computing nodes include: a GPU physical machine, a direct GPU host machine, and a virtualized GPU host machine; A node determination module is used to determine a target computing node according to the heterogeneous computing resource requirements and the node information when receiving the heterogeneous computing resource requirements sent by the client node; The resource allocation module is used to create target heterogeneous computing resources according to the heterogeneous computing resource requirements and the heterogeneous computing resources in the target computing node, and allocate the target heterogeneous computing resources to the client node.

8. A heterogeneous computing resource management device, characterized in that: The device is deployed on a client node, and includes: The demand acquisition module is used to obtain the heterogeneous computing resource requirements corresponding to the target task; A demand sending module, used to send the heterogeneous computing power resource demand to the control node; A task processing module is used to process the target task based on the target heterogeneous computing power resources allocated by the control node upon receiving the target heterogeneous computing power resources allocated by the control node, wherein the target heterogeneous computing power resources are obtained by the control node according to the heterogeneous computing power resource requirements and the heterogeneous computing power resources in the target computing node, and the target computing node is obtained according to the heterogeneous computing power resource requirements and the node information of each computing node.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the heterogeneous computing resource management method according to any one of claims 1 to 6 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the heterogeneous computing resource management method described in any one of claims 1 to 6.