GPU computing power scheduling method and device, equipment, storage medium and program product

By obtaining and analyzing the topological information and performance baseline data of the GPU partition, dynamically allocating the GPU resources required by the virtual machine, solving the problem that multi-vendor GPU cards cannot be fully interconnected in the same host, improving GPU communication efficiency and task processing performance, and dynamic resource adjustment is achieved under resource fragmentation.

CN120104346AActive Publication Date: 2025-06-06ALIBABA CLOUD FEITIAN (HANGZHOU) CLOUD COMPUTING TECH CO LTD

Patent Information

Application Number
CN202510571609.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-06
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

In a virtualization scenario, GPU cards from multiple different manufacturers cannot be fully interconnected in the same host, resulting in high difficulty in scheduling of GPU computing power and low communication efficiency, which affects the virtual machine task processing efficiency and user task processing performance.

Method used

By obtaining the topological information of each GPU partition in the target host, the performance baseline data of the GPU partition, including communication factors and discrete factors, and allocating corresponding GPU partitions according to the virtual machine's demand data, in order to achieve reasonable and dynamic scheduling of GPU resources.

Benefits of technology

It improves GPU communication efficiency and task processing efficiency, improves user task processing performance, and realizes dynamic resource adjustment under the condition of GPU resource fragmentation, improving resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104346A_ABST
    Figure CN120104346A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a GPU computing power scheduling method and device, equipment, a storage medium and a program product, and is applied to the technical field of computers. The method comprises the steps of obtaining topological information of each GPU partition in a target host, and determining performance baseline data of the GPU partitions according to the topological information; according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, determining a first target GPU partition corresponding to the target virtual machine, so that the target virtual machine performs task processing based on the first target GPU partition; under the condition that the resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition, determining a second target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine in the running state and the performance baseline data of each GPU partition, and enabling the target virtual machine to perform task processing based on the second target GPU partition. Therefore, the inter-card communication efficiency of the GPU can be ensured, and the task processing efficiency and the task processing performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a GPU computing power scheduling method, device, equipment, storage medium and program product. Background Art

[0002] With the continuous development of technologies such as artificial intelligence (AI) and deep learning, the demand for graphics processing units (GPUs) in processes such as model training and reasoning has increased significantly.

[0003] In related technologies, in order to efficiently utilize GPU resources, GPUs from multiple manufacturers are usually deployed in virtualization scenarios. However, due to the lack of high-speed direct connection components in GPU cards from different manufacturers, it is impossible to achieve full interconnection of multiple GPU cards in the same host, resulting in great difficulty in GPU computing power scheduling and inability to ensure communication efficiency, which in turn leads to low virtual machine task processing efficiency and low user task processing performance. Summary of the invention

[0004] Multiple aspects of the present application provide a GPU computing power scheduling method, device, equipment, storage medium and program product, which can realize reasonable and dynamic scheduling of GPU computing power, improve communication efficiency and task processing efficiency, and also improve the user's task processing performance.

[0005] In a first aspect, an embodiment of the present application provides a GPU computing power scheduling method, including:

[0006] Acquire topology information of each GPU partition in the target host, and determine performance baseline data of the GPU partition according to the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition;

[0007] Determine a first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition;

[0008] When the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, the second target GPU partition corresponding to the target virtual machine is determined according to the demand data of the running target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition.

[0009] In a possible implementation, the performance baseline data includes a communication factor and a discrete factor; the communication factor is used to characterize the communication performance of the GPU partition; the discrete factor is used to characterize the discrete degree of the GPU partition; and determining the performance baseline data of the GPU partition according to the topology information includes:

[0010] For each GPU partition, determine the topological structure and interconnection mode of the GPU partition according to the topological information, and determine the communication factor of the GPU partition according to the topological structure and the interconnection mode;

[0011] The physical location code of each GPU in the GPU partition is determined according to the topology information, and the discrete factor of the GPU partition is calculated according to the physical location code.

[0012] In a possible implementation manner, determining the communication factor of the GPU partition according to the topology structure and the interconnection mode includes:

[0013] Determining a topological weight of the GPU partition according to the topological structure and the interconnection mode;

[0014] The communication bandwidth of the GPU partition is measured, and a communication factor of the GPU partition is calculated according to the communication bandwidth and the topological weight.

[0015] In a possible implementation, the method further includes:

[0016] After calculating the communication factor of the first GPU partition, if the topology structure and interconnection mode of the second GPU partition are the same as those of the first GPU partition, assigning the communication factor of the second GPU partition to the communication factor of the first GPU partition;

[0017] If there is an isolated GPU in the GPU partition, the GPU partition is marked as a target state partition, and calculation of performance baseline data of the GPU partition is skipped.

[0018] In a possible implementation, the calculating the discrete factor of the GPU partition according to the physical location encoding includes:

[0019] According to the physical location code, physical distribution data corresponding to the GPU partition is calculated, and the physical distribution data is used as a discrete factor of the GPU partition.

[0020] In a possible implementation, the method further includes:

[0021] Obtaining a real-time discrete factor corresponding to each GPU partition, and calculating a global discrete factor corresponding to the GPU partition according to the real-time discrete factor; determining that resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition when the global discrete factor is greater than a preset discrete threshold; the real-time discrete factor is a discrete factor corresponding to each GPU partition within a preset period; or,

[0022] Obtain task performance data of the target virtual machine; the task performance data includes at least one of actual GPU utilization, actual communication bandwidth, and actual task iteration speed; when the task performance data is less than a preset performance threshold, determine whether resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition.

[0023] In a possible implementation, the method further includes:

[0024] Initializing the second target GPU partition according to the corresponding relationship between the target virtual machine and the second target GPU partition;

[0025] Establishing a binding relationship between each GPU in the second target GPU partition after the initialization process and the target virtual machine, and updating the GPU list of the target virtual machine;

[0026] When the second target GPU partition works normally after the initialization process, the binding relationship between the target virtual machine and the first target GPU partition is released.

[0027] In a possible implementation, the method further includes:

[0028] In each of the target virtual machines, collecting status information of a GPU used by the target virtual machine;

[0029] The status information is received in a target host, and when the status information meets a preset warning condition, a target warning information is output; the preset warning condition includes that the communication bandwidth of the GPU is less than a preset bandwidth threshold, the communication delay of the GPU is greater than a preset delay threshold, or the transmission error count of the GPU is greater than a preset number threshold.

[0030] In a second aspect, an embodiment of the present application provides a GPU computing power scheduling device, including:

[0031] An acquisition module, used to acquire topology information of each GPU partition in the target host, and determine performance baseline data of the GPU partition according to the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition;

[0032] A first determination module, configured to determine a first target GPU partition corresponding to the target virtual machine according to demand data of the target virtual machine and performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition;

[0033] The second determination module is used to determine the second target GPU partition corresponding to the target virtual machine according to the demand data of the running target virtual machine and the performance baseline data of each GPU partition when the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, so that the target virtual machine performs task processing based on the second target GPU partition.

[0034] In a possible implementation, the performance baseline data includes a communication factor and a discrete factor; the communication factor is used to characterize the communication performance of the GPU partition; the discrete factor is used to characterize the discrete degree of the GPU partition; the acquisition module is specifically used to:

[0035] For each GPU partition, determine the topological structure and interconnection mode of the GPU partition according to the topological information, and determine the communication factor of the GPU partition according to the topological structure and the interconnection mode;

[0036] The physical location code of each GPU in the GPU partition is determined according to the topology information, and the discrete factor of the GPU partition is calculated according to the physical location code.

[0037] In a possible implementation manner, the acquisition module is specifically configured to:

[0038] Determining a topological weight of the GPU partition according to the topological structure and the interconnection mode;

[0039] The communication bandwidth of the GPU partition is measured, and a communication factor of the GPU partition is calculated according to the communication bandwidth and the topological weight.

[0040] In a possible implementation manner, the device is further used for:

[0041] After calculating the communication factor of the first GPU partition, if the topology structure and interconnection mode of the second GPU partition are the same as those of the first GPU partition, assigning the communication factor of the second GPU partition to the communication factor of the first GPU partition;

[0042] If there is an isolated GPU in the GPU partition, the GPU partition is marked as a target state partition, and calculation of performance baseline data of the GPU partition is skipped.

[0043] In a possible implementation manner, the acquisition module is specifically configured to:

[0044] According to the physical location code, physical distribution data corresponding to the GPU partition is calculated, and the physical distribution data is used as a discrete factor of the GPU partition.

[0045] In a possible implementation manner, the device is further used for:

[0046] Obtaining a real-time discrete factor corresponding to each GPU partition, and calculating a global discrete factor corresponding to the GPU partition according to the real-time discrete factor; determining that resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition when the global discrete factor is greater than a preset discrete threshold; the real-time discrete factor is a discrete factor corresponding to each GPU partition within a preset period; or,

[0047] Obtain task performance data of the target virtual machine; the task performance data includes at least one of actual GPU utilization, actual communication bandwidth, and actual task iteration speed; when the task performance data is less than a preset performance threshold, determine whether resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition.

[0048] In a possible implementation manner, the device is further used for:

[0049] Initializing the second target GPU partition according to the corresponding relationship between the target virtual machine and the second target GPU partition;

[0050] Establishing a binding relationship between each GPU in the second target GPU partition after the initialization process and the target virtual machine, and updating the GPU list of the target virtual machine;

[0051] When the second target GPU partition works normally after the initialization process, the binding relationship between the target virtual machine and the first target GPU partition is released.

[0052] In a possible implementation manner, the device is further used for:

[0053] In each of the target virtual machines, collecting status information of a GPU used by the target virtual machine;

[0054] The status information is received in a target host, and when the status information meets a preset warning condition, a target warning information is output; the preset warning condition includes that the communication bandwidth of the GPU is less than a preset bandwidth threshold, the communication delay of the GPU is greater than a preset delay threshold, or the transmission error count of the GPU is greater than a preset number threshold.

[0055] In a third aspect, an embodiment of the present application provides a GPU computing power scheduling device, including: a memory and a processor;

[0056] The memory stores computer-executable instructions;

[0057] The processor executes the computer execution instructions stored in the memory, so that the processor executes the GPU computing power scheduling method described in any one of the first aspects.

[0058] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer execution instructions, which, when executed by a processor, are used to implement the GPU computing power scheduling method described in any one of the first aspects.

[0059] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the GPU computing power scheduling method shown in any one of the first aspects.

[0060] In a sixth aspect, an embodiment of the present application provides a GPU computing power scheduling system, including a target host and a target virtual machine;

[0061] The target host is used to obtain topology information of each GPU partition in the target host, and determine performance baseline data of the GPU partition according to the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition;

[0062] The target host is used to determine a first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, and the target virtual machine is used to perform task processing through the first target GPU partition;

[0063] The target host is used to determine a second target GPU partition corresponding to the target virtual machine according to demand data of a running target virtual machine and performance baseline data of each GPU partition when resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition, and the target virtual machine is used to perform task processing through the second target GPU partition.

[0064] In an embodiment of the present application, the topology information of each GPU partition in the target host is obtained, and the performance baseline data of the GPU partition is determined according to the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition; according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, the first target GPU partition corresponding to the target virtual machine is determined, so that the target virtual machine performs task processing based on the first target GPU partition; when the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, according to the demand data of the target virtual machine in operation and the performance baseline data of each GPU partition, the second target GPU partition corresponding to the target virtual machine is determined, so that the target virtual machine performs task processing based on the second target GPU partition. In the present application, based on the topology information of each GPU partition in the target host, the performance baseline data of each GPU partition is determined, and the first target GPU partition corresponding to the target virtual machine is determined according to the demand data of the target virtual machine and the performance baseline data. Later, during the operation of the system, if it is detected that the resource fragmentation degree of multiple GPU partitions in the target host is high, the second target GPU partition corresponding to each target virtual machine can be determined at this time, so that the target virtual machine performs task processing based on the second target GPU partition. In this way, GPU resources are allocated based on the topological information of each GPU. At the same time, dynamic and reasonable configuration of GPU resources is achieved when the degree of resource fragmentation is high, which can ensure the efficiency of inter-GPU communication, thereby improving task processing efficiency and improving users' task processing performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0066] Figure 1 An X-Link topology structure of a host GPU card in the related art;

[0067] Figure 2 A schematic diagram of the steps of a GPU computing power scheduling method provided in an embodiment of the present application;

[0068] Figure 3 A schematic diagram of a topological structure provided in an embodiment of the present application;

[0069] Figure 4 A schematic diagram of target host GPU resource fragmentation provided in an embodiment of the present application;

[0070] Figure 5 A schematic diagram of the steps of another GPU computing power scheduling method provided in an embodiment of the present application;

[0071] Figure 6A schematic diagram of a state after dynamic allocation of GPU resources provided in an embodiment of the present application;

[0072] Figure 7 A schematic diagram of a system architecture for GPU computing power scheduling provided for an exemplary embodiment of the present application;

[0073] Figure 8 A schematic diagram of the structure of a GPU computing power scheduling device provided for an exemplary embodiment of the present application;

[0074] Fig. 9 A schematic diagram of the structure of a GPU computing power scheduling device provided for an exemplary embodiment of the present application.

[0075] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0076] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of them. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0077] The following is an explanation of the professional terms involved in this application:

[0078] Graphics Processing Unit GPU: used to accelerate computing tasks.

[0079] Elastic GPU Service (EGS): provides GPU accelerated computing capabilities and enables ready-to-use and elastic scaling of GPU computing resources.

[0080] NVLink / NVSwitch: A high-speed GPU interconnect technology of NVIDIA, where NVLink is a high-speed GPU interconnect technology that provides high bandwidth; NVSwitch is a switching architecture used to achieve full interconnection of multiple GPUs.

[0081] X-Link: An interconnection technology between GPU cards from domestic manufacturers similar to NVLink.

[0082] Kernel-based Virtual Machine (KVM): A virtualization infrastructure for Linux systems.

[0083] Pass-through: Directly allocates physical hardware devices to virtual machines for use.

[0084] Host: A physical server that runs a virtual machine hypervisor.

[0085] Virtual Machine (VM): A virtual computer system running on a host machine.

[0086] High-speed serial expansion bus (Peripheral Component Interconnect Express, PCle): used to connect peripheral devices.

[0087] With the rapid growth of AI and deep learning tasks, the demand for GPUs has increased significantly. In order to efficiently utilize resources, GPUs from multiple manufacturers can be deployed based on EGS in a virtualized environment. As an AI infrastructure platform, EGS has high flexibility, isolation, and reliability. In a virtualized environment, direct virtualization technology plays an important role in resource management and allocation. For GPU computing resources, the cloud platform pools and encapsulates the GPU resources on the physical machine, and directly connects these GPU resources to different virtual machines according to user needs.

[0088] In related technologies, for hosts that use NVSwitch and NVLink to achieve full interconnection of multiple cards on a single machine, there is a high-speed NVLink direct connection between any two cards, and the virtualization management and control side only needs to allocate the corresponding number of GPU cards according to the computing power requirements. However, for virtualization scenarios where multiple domestic manufacturers' GPUs are deployed, there is a lack of components similar to NVSwitch, and full interconnection of multiple cards on a single machine is not supported. The topology and interconnection method of the GPUs in the same host are relatively complex, resulting in great difficulty in GPU computing power scheduling, inability to ensure communication efficiency, and inefficient virtual machine task processing.

[0089] For example, Figure 1 It is an X-Link topology structure of a host GPU card in the related art. Figure 1As shown, in the virtualization scenario, the same host includes two central processing units (CPUs), namely cpu0 and cpu1. The host uses GPUs from different manufacturers, with a total of 16 GPU cards (GPU0 to GPU15). Since there is no component like NVSwitch, the GPU cards in a single machine can only be fully interconnected based on X-Link within a small partition (usually between 4 specific GPUs). In addition, within the same host, there are three types of interconnection between GPU cards: dual-line interconnection (2X-Link), single-line interconnection (1X-Link), and PCle. For example, there is 2X-Link between GPU0 and GPU1, and 1X-Link between GPU2 and GPU8. Two GPUs that are not directly connected by X-Link can only communicate through PCle, such as GPU0 and GPU8. By Figure 1 It can be seen that in the virtualization scenario, the interconnection method between the GPU cards in the same host is relatively complex. At the same time, due to the diversity of topological states, it is difficult to schedule GPU computing power, which has a great impact on the efficiency of communication between cards, and thus affects the performance of user model training and reasoning tasks.

[0090] In order to solve the above problems, the present application provides a GPU computing power scheduling method, device, equipment, storage medium and program product, which obtains the topology information of each GPU partition in the target host, and determines the performance baseline data of the GPU partition according to the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition; according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, the first target GPU partition corresponding to the target virtual machine is determined, so that the target virtual machine performs task processing based on the first target GPU partition; when the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, according to the demand data of the target virtual machine in operation and the performance baseline data of each GPU partition, the second target GPU partition corresponding to the target virtual machine is determined, so that the target virtual machine performs task processing based on the second target GPU partition. In the present application, based on the topology information of each GPU partition in the target host, the performance baseline data of each GPU partition is determined, and the first target GPU partition corresponding to the target virtual machine is determined according to the demand data of the target virtual machine and the performance baseline data. Later, during the operation of the system, if it is detected that the resource fragmentation degree of multiple GPU partitions in the target host is high, the second target GPU partition corresponding to each target virtual machine can be determined, and the target virtual machine performs task processing based on the second target GPU partition. In this way, GPU resources are allocated based on the topological information of each GPU. At the same time, dynamic and reasonable configuration of GPU resources is achieved when the degree of resource fragmentation is high, which can ensure the efficiency of inter-GPU communication, thereby improving task processing efficiency and improving users' task processing performance.

[0091] The technical solution shown in the present application is described in detail below through specific embodiments. It should be noted that the following embodiments can exist independently or in combination with each other, and the same or similar contents will not be described repeatedly in different embodiments.

[0092] Figure 2 A schematic diagram of the steps of a GPU computing power scheduling method provided in an embodiment of the present application. Figure 2 , the GPU computing power scheduling method may include:

[0093] S201, obtaining topology information of each GPU partition in the target host, and determining performance baseline data of the GPU partition according to the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition.

[0094] The execution subject of the embodiment of the present application can be an electronic device, or a GPU computing power scheduling device set in the electronic device. The GPU computing power scheduling device can be implemented by software, or by a combination of software and hardware. For ease of understanding, the following description is taken as an example of an execution subject being an electronic device. The electronic device can specifically refer to a server or a cloud platform, etc., or it can refer to a target host (such as a GPU centralized control platform in the target host, etc.), and the embodiment of the present application does not limit this.

[0095] In the embodiment of the present application, the target host may refer to a host machine or a physical machine for running a virtual machine. The GPU partition may refer to each partition that may be formed by the GPU in the target host, and each partition may include N GPU cards, where N is a positive integer. Since X-Link is usually used by different manufacturers to fully interconnect 4 cards, the value of N can be 4. Of course, based on the needs of actual tasks, the electronic device can also be partitioned according to 2 GPU cards or 8 GPU cards, which is not limited in the embodiment of the present application.

[0096] The topology information may refer to the topology information corresponding to each GPU partition, and may specifically include the topology structure, interconnection mode, and location information of multiple GPUs in the GPU partition. The topology structure may refer to the topological space structure within the GPU partition, and may specifically include irregular topology structure and regular topology structure. For example, Figure 3 A schematic diagram of a topological structure provided in an embodiment of the present application. Figure 3 As shown in (a), the topological structure of the GPU partition is an irregular topological structure, which is not a regular rectangle in space. Figure 3 As shown in (b), the topological structure of the GPU partition is a regular topological structure, forming a regular rectangle in space.

[0097] The interconnection mode may refer to the connection mode between GPUs in a GPU partition, specifically, dual-line interconnection, single-line interconnection, and PCle interconnection, etc. The location information may refer to the actual physical location of each GPU card in the hardware topology, specifically, including the rack location, motherboard slot, and non-uniform memory access (NUMA) node, etc. Of course, the topology information may also include other information, such as hardware bandwidth information, load information, and heat dissipation partition information, etc., which is not limited in the embodiments of the present application.

[0098] The performance baseline data can be used to characterize the communication performance of each GPU partition, and can be used to guide the allocation strategy of the electronic device for the GPU partition. The performance baseline data may include the communication factor (ComFactor) and the discrete factor (FragFactor) corresponding to the GPU partition. Among them, the communication factor can be used to characterize the communication performance of GPU partitions of different topological forms; the larger the communication factor of the GPU partition, the higher the communication efficiency of the GPU partition. The discrete factor can be used to characterize the degree of discretization of GPU partitions of the same topological form. The larger the discrete factor of the GPU partition, the more serious the degree of resource fragmentation caused by the GPU partition.

[0099] In the embodiment of the present application, the electronic device obtains the topology information of each GPU partition in the target host, and then can calculate the performance baseline data of each GPU partition according to the topology information. When the electronic device needs to allocate GPU resources to the target virtual machine, it can allocate the GPU partition with higher communication efficiency according to the performance baseline data, so as to ensure the task processing performance of the target virtual machine.

[0100] S202: Determine a first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition.

[0101] In the embodiment of the present application, the demand data may refer to the GPU resources actually required by the target virtual machine, which may specifically include the number of GPUs, etc. The first target GPU partition may refer to the GPU partition allocated to the target virtual machine. After determining the performance baseline data of the GPU partition, the electronic device may determine the first target GPU partition according to the demand data of the target virtual machine, and may specifically select an idle GPU partition with the largest communication factor and the smallest discrete factor as the first target GPU partition, so as to ensure that the communication efficiency of the GPU partition selected by the target virtual machine in the current situation is high and the degree of fragmentation is low.

[0102] In one possible implementation, based on the performance baseline data of each GPU partition, the electronic device can sort the GPU partitions with different topological structures in order from large to small according to the communication factor, and sort the GPU partitions with the same topological structure in order from small to large according to the discrete factor. When it is necessary to allocate GPU resources to the target virtual machine, the electronic device can select the idle GPU partition ranked first as the first target GPU partition, which can ensure the communication efficiency and task processing efficiency of the target virtual machine and improve the user's task processing performance.

[0103] S203. When the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, the second target GPU partition corresponding to the target virtual machine is determined according to the demand data of the running target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition.

[0104] In the embodiment of the present application, resource fragmentation information may refer to the overall GPU resource fragmentation degree of each GPU partition in the target host, and may specifically include parameters such as the real-time global discrete factor of each GPU partition in the target host (characterizing the overall discrete degree of the GPU partition) and the real-time task performance data of the target virtual machine (such as actual GPU utilization, actual communication bandwidth, and actual iteration speed, etc.). The preset condition may refer to a pre-set trigger condition for GPU resource reallocation, for example, the preset condition may refer to a global discrete factor greater than a preset discrete threshold, or an actual GPU utilization less than a preset utilization threshold, or an actual iteration speed less than a preset speed threshold, etc. The embodiment of the present application does not limit the specific type of the preset condition.

[0105] During system operation, since users create and release target virtual machines multiple times on the target host, the fragmentation of GPU card resources in the target host gradually increases, and the effective utilization of resources gradually decreases. When the fragmentation of GPU resources is too high, although there are enough idle resources, these idle resources are often scattered in different GPU partitions, resulting in the target virtual machine being unable to be allocated to a GPU partition with higher communication efficiency, and the utilization of GPU resources is low. For example, Figure 4 A schematic diagram of a target host GPU resource fragmentation provided in an embodiment of the present application. Figure 4 As shown, among the 16 GPU cards in the target host, 8 GPU cards are in an occupied state. Although there are 8 idle GPU cards, the communication efficiency of the GPU partition composed of these 8 GPU cards is not high. In this way, the high degree of resource fragmentation leads to the inability to fully utilize the computing power of the GPU resources in the target host, affecting the communication efficiency and task processing performance.

[0106] In an embodiment of the present application, the electronic device can monitor the real-time operating status of each GPU partition in the target host and collect resource fragmentation information of the GPU partition in the target host. When the resource fragmentation information meets the preset conditions, the electronic device can determine that the GPU resource fragmentation degree in the target host is high at this time, and it is necessary to reallocate the GPU resources corresponding to each target virtual machine to improve the utilization rate of GPU resources and ensure the user task processing performance. The electronic device can determine at least one target virtual machine that is currently in operation, and then obtain the demand data corresponding to each target virtual machine. According to the demand data and the performance baseline data of each GPU partition, the GPU resources corresponding to the target virtual machine are reallocated to obtain the second target GPU partition corresponding to each target virtual machine. The target virtual machine can be subsequently bound to the second target GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition, and realizes the dynamic and reasonable allocation of the GPU resources of the target virtual machine in the virtualization scenario, ensuring the communication efficiency and task processing efficiency.

[0107] In an embodiment of the present application, the electronic device obtains the topology information of each GPU partition in the target host, and determines the performance baseline data of the GPU partition according to the topology information; determines the first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition; when the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, determines the second target GPU partition corresponding to each target virtual machine according to the demand data of multiple target virtual machines in operation and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition. In this way, in the present application, the performance baseline data of each GPU partition is determined based on the topology information of each GPU partition in the target host, and the first target GPU partition corresponding to the target virtual machine is determined according to the demand data of the target virtual machine and the performance baseline data. In the subsequent operation process, if it is detected that the resource fragmentation degree of multiple GPU partitions in the target host is high, the second target GPU partition corresponding to each target virtual machine can be determined at this time, so that the target virtual machine performs task processing based on the second target GPU partition. In this way, GPU resources are allocated based on the topological information of each GPU. At the same time, dynamic and reasonable configuration of GPU resources is achieved when the degree of resource fragmentation is high, which can ensure the efficiency of inter-GPU communication, thereby improving task processing efficiency and improving users' task processing performance.

[0108] Based on the above embodiments, Figure 5 A flowchart of another GPU computing power scheduling method provided in an embodiment of the present application. Figure 5, the GPU computing power scheduling method may include:

[0109] S501, obtaining topology information of each GPU partition in the target host; for each GPU partition, determining the topology structure and interconnection mode of the GPU partition according to the topology information, and determining the communication factor of the GPU partition according to the topology structure and the interconnection mode.

[0110] In an embodiment of the present application, the performance baseline data of each possible GPU partition in the target host can be used to characterize the communication performance and discreteness of the GPU partition. The performance baseline data can specifically include the communication factor and discreteness factor of each GPU partition, wherein the communication factor is used to characterize the communication performance of the GPU partition; the discreteness factor is used to characterize the discreteness of the GPU partition. Specifically, the electronic device can parse and read the topological structure and interconnection mode of the GPU partition from the topological information of the GPU partition, and then calculate the communication factor of the GPU partition based on the topological structure and interconnection mode of the GPU partition, so as to realize the calculation and evaluation of the communication performance of GPU partitions with different topological forms.

[0111] In a possible implementation, the communication factor in the performance baseline data may be calculated as follows:

[0112] The topological weight of the GPU partition is determined according to the topological structure and the interconnection mode; the communication bandwidth of the GPU partition is measured, and the communication factor of the GPU partition is calculated according to the communication bandwidth and the topological weight.

[0113] In the embodiment of the present application, the topological weight may refer to the topological structure of the GPU partition and the weight information corresponding to the interconnection mode. Since the communication bandwidths of different interconnection modes are significantly different, for example, the topological weight of the X-link direct connection mode is greater than the topological weight corresponding to the PCle direct connection. The communication bandwidth may refer to the communication bandwidth actually tested by the current GPU partition.

[0114] Specifically, when calculating the communication factor of each GPU partition, the electronic device can first determine the topological structure and interconnection method of the GPU partition. Specifically, the GPU management tool can be used to obtain the topological information of each GPU partition, and the topological structure and interconnection method of the GPU partition can be determined based on the topological information. For example, an adjacency matrix can be used to represent the interconnection method between GPUs. Afterwards, the electronic device can measure the communication bandwidth of the GPU partition through the GPU communication library and topological isolation tools, for example, the point-to-point bandwidth between any two cards in the GPU partition can be measured, and the communication bandwidth of the GPU partition can be obtained by calculating the average value. Afterwards, the electronic device can determine the topological weight corresponding to the GPU partition based on the topological structure and interconnection method of the GPU partition. Afterwards, the electronic device can use the communication bandwidth and topological weight to calculate the communication factor of the current GPU partition. Exemplarily, the electronic device calculates the communication factor of the GPU partition specifically according to the following formulas (1) to (3):

[0115]

[0116]

[0117]

[0118] In the above formula, ComFator 1 It can refer to the first component corresponding to the communication factor, which is used to characterize the ratio of the actual measured communication bandwidth of the GPU partition to the theoretical bandwidth (the theoretical bandwidth value calculated according to the interconnection method). 2 It may refer to the second component corresponding to the communication factor, which is used to characterize the communication bandwidth of the GPU partition after weighted processing; BW i,j It can refer to the actual measured communication bandwidth between GPU card i and GPU card j, W i,j It refers to the topological weight corresponding to the interconnection between GPU card i and GPU card j. α and β refer to the weights corresponding to the first component and the second component respectively, which are used to balance the influence of different components. For example, α can be 0.7 and β can be 0.3. 1 and ComFator 2 It can be a normalized value for ease of calculation and comparison. In the embodiment of the present application, based on the above formulas (1) to (3), the electronic device can calculate the first component and the second component of the GPU partition communication factor, and then calculate the communication factor by weighted summation, which can balance the influencing factors such as communication bandwidth and interconnection mode, and realize reasonable and accurate evaluation of the GPU partition communication factor.

[0119] It should be noted that the above formulas (1) to (3) are merely examples, and the communication factor of the GPU partition may also be calculated in other ways, which is not limited in the embodiments of the present application.

[0120] In the embodiment of the present application, when there are many GPU cards in the target host, there are many GPU partitions that may be formed in the target host. In order to reduce the amount of calculation of performance baseline data, the electronic device may adopt a certain pruning strategy when calculating the performance baseline data of each GPU partition. In a possible implementation, the GPU computing power scheduling method may also include the following steps:

[0121] After calculating the communication factor of the first GPU partition, if the topology and interconnection method of the second GPU partition are the same as those of the first GPU partition, the communication factor of the second GPU partition is assigned to the communication factor of the first GPU partition; if there is an isolated GPU in the GPU partition, the GPU partition is marked as a target state partition, and the calculation of the performance baseline data of the GPU partition is skipped.

[0122] In the embodiment of the present application, if the topological structures and interconnection methods of two GPU partitions are exactly the same, the communication factors of the two GPU partitions can be regarded as the same, and the electronic device only needs to calculate the communication factor of one of the GPU partitions. That is, for the second GPU partition that has the same topological structure and interconnection method as the first GPU partition, the electronic device can directly reuse the communication factor of the first GPU partition as the communication factor of the second GPU partition, so that the electronic device does not need to repeat the test, which can reduce the number of test cases and reduce unnecessary calculation overhead.

[0123] An isolated GPU may refer to a GPU that does not have an X-link direct connection with other GPUs. In this case, the isolated GPU is approximately equivalent to a completely disconnected state, and the communication efficiency is low. The electronic device may directly mark the GPU partition including the isolated GPU as a target state partition, such as a "partition with poor communication efficiency". The electronic device may skip the calculation of the target state partition performance baseline data, or set a minimum communication factor for the target state partition, which is not limited in the embodiments of the present application.

[0124] In addition, the electronic device can also establish a GPU partition performance database corresponding to the target host, and store the performance baseline data of each possible GPU partition, so as to facilitate subsequent reuse, avoid repeated calculation of the performance baseline data by the electronic device, and save computing resources. In an embodiment of the present application, when calculating the performance baseline data of the GPU partition, for GPU partitions with the same topology and the same interconnection method, the electronic device can select one of the GPU partitions to calculate the communication factor, and then assign the communication factor to other GPU partitions, thereby avoiding repeated calculation of the communication factor and reducing computing overhead; for GPU partitions with isolated GPUs, the electronic device directly marks them as target state partitions, and there is no need to calculate the performance baseline data of the target state partition, so as to avoid the calculation of GPU partitions with low communication efficiency, and further save computing resources.

[0125] S502: Determine the physical location code of each GPU in the GPU partition according to the topology information, and calculate the discrete factor of the GPU partition according to the physical location code.

[0126] In the embodiment of the present application, the physical location code may refer to the physical distribution code corresponding to each GPU card in the GPU partition, which may be determined according to the actual physical location of the GPU card, and may be represented by Square (g i ). After determining the topology information of each GPU partition, the electronic device can determine the physical location code of each GPU card in the GPU partition according to the location information in the topology information, and then calculate the discrete factor of the GPU partition according to the physical location code, thereby effectively evaluating the discretization degree of the GPU cards in the GPU partition.

[0127] In a possible implementation, the discrete factor can be calculated as follows:

[0128] According to the physical location encoding, the physical distribution data corresponding to the GPU partition is calculated, and the physical distribution data is used as the discrete factor of the GPU partition.

[0129] In the embodiment of the present application, the physical distribution data may refer to the physical distribution discrete degree data corresponding to each GPU card in the GPU partition, and specifically may refer to the physical distribution variance or physical distribution standard deviation of each GPU card in the GPU partition. After determining the physical position code corresponding to each GPU in the GPU partition, the electronic device may calculate the physical distribution data corresponding to the GPU partition, and use the physical distribution data as the discrete factor of the GPU partition to achieve an effective evaluation of the discretization degree of the GPU partition. For example, the electronic device may calculate the physical distribution data of the GPU partition according to the following formula (4):

[0130]

[0131] In the above formula (4), n is the number of GPU cards in the GPU partition, and μ is the average value of the physical position codes of each GPU card in the GPU partition. After obtaining the physical position codes of each GPU card in the GPU partition, the electronic device obtains the physical distribution data of the GPU partition by calculating the variance, and then obtains the discrete factor of the GPU partition, which can accurately characterize the discrete degree of distribution of each GPU in the GPU partition. For example, if each GPU card in a GPU partition belongs to the same rack, the physical position codes of the GPU cards in the GPU partition are slightly different, the variance is small, and the discrete factor is also small; if each GPU card in a GPU partition belongs to different racks, the physical position codes of the GPU cards in the GPU partition are significantly different, the variance is large, and the discrete factor is also large. Since the communication bandwidth is affected by the physical connection distance, the closer the physical position, the higher the communication efficiency. For example, the communication efficiency of GPU cards in the same rack is greater than the communication efficiency of GPU cards across racks. When the electronic device subsequently allocates GPU resources to the target virtual machine, it can give priority to allocating GPU partitions with smaller discrete factors to ensure the communication efficiency and task performance of users during task processing.

[0132] Of course, the above formula (4) is only an example. The electronic device may also use other methods to calculate the physical distribution data of the GPU partition, and then obtain the discrete factor of the GPU partition. The specific selection can be based on actual needs, and the embodiment of the present application is not limited to this.

[0133] S503: Determine a first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition.

[0134] In the embodiment of the present application, after determining the communication factor and discrete factor of each GPU partition, the electronic device can store the performance baseline data of each GPU partition. When it is necessary to allocate GPU resources to the target virtual machine, the electronic device can determine the first target GPU partition corresponding to the target virtual machine based on the demand data of the target virtual machine and the performance baseline data of each GPU partition, so as to ensure the communication efficiency and task processing performance of the target virtual machine and improve the user experience.

[0135] S504. When the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, the second target GPU partition corresponding to each target virtual machine is determined according to the demand data of multiple target virtual machines in operation and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition.

[0136] In the embodiment of the present application, during the operation of the system, multiple users may create and release target virtual machines for the same target host multiple times, and the degree of GPU resource fragmentation in the target host gradually increases. When the resource fragmentation information corresponding to the GPU partition in the target host meets the preset conditions, the electronic device can dynamically re-allocate the GPU resources of the running target virtual machine to achieve efficient use of GPU resources.

[0137] In a possible implementation, whether the resource fragmentation information corresponding to each GPU partition in the target host meets the preset condition can be determined in the following manner:

[0138] Obtain the real-time discrete factor corresponding to each GPU partition, and calculate the global discrete factor corresponding to the GPU partition according to the real-time discrete factor; when the global discrete factor is greater than a preset discrete threshold, determine that the resource fragmentation information corresponding to each GPU partition in the target host meets the preset condition; the real-time discrete factor is the discrete factor corresponding to each GPU partition within a preset period; or,

[0139] Obtain task performance data of the target virtual machine; the task performance data includes at least one of actual GPU utilization, actual communication bandwidth, and actual task iteration speed; when the task performance data is less than a preset performance threshold, determine whether resource fragmentation information corresponding to each GPU partition in the target host meets preset conditions.

[0140] In an embodiment of the present application, the real-time discrete factor may refer to the real-time discrete factor of the GPU partition in the target host. The global discrete factor may refer to the average real-time discrete factor corresponding to each GPU partition in the target host. The preset discrete threshold may refer to a preset discrete factor critical value of GPU resource reallocation, which may specifically refer to 0.6, 0.5 or 0.4, etc., which is not limited in the embodiment of the present application. Specifically, the electronic device may monitor the discrete factors of each GPU partition in the target host in real time, obtain the real-time discrete factors corresponding to each GPU partition, and then calculate the average value of the real-time discrete factors of each GPU partition to obtain the global discrete factors of each GPU partition in the target host. If the global discrete factor is greater than the preset discrete threshold, the electronic device may determine that the discretization degree of the GPU resources in the target host is high at this time, and the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions.

[0141] The task performance data may refer to the actual performance data of the target virtual machine, which may specifically include the actual GPU utilization, the actual communication bandwidth, and the actual iteration speed of the task. The preset performance threshold may refer to the preset critical value of the task performance data for GPU resource reallocation, which may specifically include the GPU utilization threshold, the communication bandwidth threshold, and the iteration speed threshold. The electronic device may collect the actual performance data of the target virtual machine in real time. When the actual performance data is less than the preset performance threshold, for example, the actual GPU utilization is less than the GPU utilization threshold, or the actual communication bandwidth is less than the communication bandwidth threshold, or the actual iteration speed of the task is less than the iteration speed threshold, the electronic device may determine that the fragmentation degree of the GPU resources is high at this time, and the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions.

[0142] It should be noted that the judgment method of whether the above two types of resource fragmentation information meet the preset conditions is an "or" relationship, that is, as long as any one of them is met, the electronic device can reallocate GPU resources to the target virtual machine. Of course, the preset condition may also include other conditions, for example, the preset condition may also include the number of consecutive times that the scheduler in the target host cannot allocate GPU partitions with communication factors that meet the standard (greater than the preset communication threshold), which is greater than or equal to the preset number threshold (for example, three consecutive times); or, the preset condition may also include that the target virtual machine task iteration speed decrease value is greater than the preset speed change threshold, etc. The embodiment of the present application does not limit the specific type of the preset condition.

[0143] In an embodiment of the present application, during the long-term use of the target host, the GPU resources allocated to the target virtual machine have a high degree of fragmentation. The electronic device can detect and calculate the global discrete factor and task performance data in the target host in real time, and when the global discrete factor is greater than a preset discrete threshold, or the task performance data is less than a preset performance threshold, determine that the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions. The electronic device can reallocate GPU resources to each target virtual machine to ensure the communication efficiency and task processing performance of each target virtual machine.

[0144] S505 , initializing the second target GPU partition according to the correspondence between the target virtual machine and the second target GPU partition; establishing a binding relationship between each GPU in the initialized second target GPU partition and the target virtual machine, and updating the GPU list of the target virtual machine.

[0145] S506: When the second target GPU partition works normally after the initialization process, release the binding relationship between the target virtual machine and the first target GPU partition.

[0146] In an embodiment of the present application, when the degree of GPU resource fragmentation in the target host is high, although there are sufficient idle resources, the target virtual machine cannot be allocated to the GPU partition with higher communication efficiency in theory, and the GPU resource utilization rate is low. The electronic device can re-allocate GPU resources to the target virtual machine when it is detected that the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, improve resource utilization, and ensure communication efficiency. Specifically, the electronic device can first determine the demand data of the target virtual machine in operation, and sort the target virtual machines in order from small to large according to the demand data; then the electronic device can determine the second target GPU partition corresponding to each target virtual machine in turn in the fully idle whole machine multi-card topology according to the performance baseline data of each GPU partition, and realize the marking of the corresponding relationship between the target virtual machine and the second target GPU partition. Afterwards, the electronic device can perform the GPU resource allocation operation based on the corresponding relationship between the target virtual machine and the second target GPU partition. The process can be performed during the period when the target virtual machine task load is low, minimizing the impact on the user as much as possible, and finally realizing the binding of the target virtual machine with the second target CPU partition.

[0147] During the configuration change execution process, the electronic device can first initialize the second target GPU partition corresponding to the target virtual machine, and the initialization process can include resource reservation and hot migration support check, etc., to achieve resource locking and ensure the normal execution of the configuration change. Afterwards, the electronic device can establish a binding relationship between the target virtual machine and each GPU in the second target GPU partition after the initialization process, which can be specifically GPU dynamic addition instructions, etc. to add each GPU in the second target GPU partition in the target virtual machine, and the electronic device can update the GPU list corresponding to the target virtual machine to achieve GPU state synchronization. When each GPU in the second target GPU partition after the initialization process is running normally, the electronic device can release the binding relationship between the target virtual machine and the first target GPU partition, and complete the configuration change process of the target virtual machine. In this way, when the target host GPU resource fragmentation is high, the electronic device re-determines the second target GPU partition corresponding to the target virtual machine, and establishes a binding relationship between the target virtual machine and the GPU in the second target GPU partition, so as to achieve dynamic adjustment of GPU resources in the virtualization scenario, which can ensure communication efficiency and improve resource utilization.

[0148] For example, Figure 6 A schematic diagram of a state after dynamic allocation of GPU resources provided by an embodiment of the present application. Figure 4 and Figure 6It can be seen from the comparison that when the resource fragmentation information of the GPU partition in the target host meets the preset conditions, the GPU resource utilization rate in the target host is not high. The electronic device can re-determine the second target GPU partition corresponding to each target virtual machine according to the demand data of the running target virtual machine and the performance baseline data of each GPU partition, and then gradually adjust the GPU resources bound to the target virtual machine by executing the configuration change, thereby reducing the degree of resource fragmentation in the target host and improving the utilization rate of GPU resources.

[0149] S507. In each target virtual machine, collect status information of a GPU used by the target virtual machine; receive the status information in the target host, and output target warning information when the status information meets preset warning conditions; the preset warning conditions include that the communication bandwidth of the GPU is less than a preset bandwidth threshold, the communication delay of the GPU is greater than a preset delay threshold, or the transmission error count of the GPU is greater than a preset number threshold.

[0150] In the embodiment of the present application, the status information may refer to information such as the usage and connection status of the GPU during actual operation, and may specifically include the actual X-Link bandwidth, communication delay, and transmission error count. The preset warning condition may refer to a preset alarm condition for abnormal GPU status, and may specifically include the communication bandwidth being less than a preset bandwidth threshold, the communication delay being greater than a preset delay threshold, or the transmission error count being greater than a preset number threshold, and may of course include other alarm conditions, which are not limited in the embodiment of the present application. The target warning information may refer to the alarm prompt information when there is an abnormality in the GPU status.

[0151] In the related art, the target host's ability to directly monitor the GPU status in a virtualized scenario is relatively limited, and it is impossible to achieve real-time and effective detection of the GPU.

[0152] In an embodiment of the present application, a proxy tool (GpuBuddy) may be provided in the target virtual machine, and the proxy tool may collect the status information of the GPU in the target virtual machine in real time. Afterwards, the target virtual machine may send the status information to the target host through the proxy tool based on the target interface, for example, a Quick Emulator Guest Agent (QEMU Guest Agent) may be used to communicate with the target host via a socket. The target host receives the status information, and outputs the target warning information when the status information meets the preset warning conditions, so as to facilitate timely handling of the abnormal GPU and ensure the safe and stable operation of the user's target virtual machine tasks.

[0153] It should be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0154] Based on any one of the above embodiments, Figure 7 A schematic diagram of a GPU computing power scheduling system architecture provided for an exemplary embodiment of the present application. Figure 7 As shown, the target host includes a GPU centralized control platform (GPUMaster) and GPU physical resources (including GPU1, GPU2, GPU3 and GPUn), wherein the GPU centralized control platform includes a scheduler (Scheduler), a monitor (Monitor) and an isolation mechanism (Isolation). In addition, multiple target virtual machines are set in the target host, and each target virtual machine can be deployed with an agent tool.

[0155] Specifically, GPU physical resources are interconnected through X-Link and other methods. As the main control module for GPU resource management, the GPU centralized control platform can be used to schedule, manage and allocate GPU resources. The scheduler is used to manage and allocate GPU resources to ensure efficient use of GPU resources. The monitor is used to receive GPU status information collected by the agent tool in the target virtual machine, and to issue an alarm in time when the status information meets the preset warning conditions. The isolation mechanism is used to ensure that the GPU resources between different target virtual machines are independent of each other to prevent resource conflicts and data leakage. The electronic device can specifically integrate the isolation methods of GPUs from different manufacturers to ensure that training or inference data is securely transmitted in the X-Link network inside the target virtual machine, while ensuring that the X-Links inside different target virtual machines are isolated from each other and do not interfere with each other.

[0156] In the embodiment of the present application, the electronic device, in the absence of high-speed interconnection hardware such as NVSwitch, realizes high-speed interconnection and dynamic scheduling between GPU cards in a virtualized environment based on X-Link, realizes dynamic allocation of GPU resources, calculates performance baseline data based on topological information, and allocates GPU partitions with higher communication efficiency to the target virtual machine based on the performance baseline data. Compared with the random allocation of GPU cards in the relevant technology, the GPU computing power scheduling method of the present application improves resource utilization, improves the communication efficiency between GPU cards inside the target virtual machine, and improves the efficiency of user task processing. The isolation mechanism of GPUs of different manufacturers is integrated in the target host for different GPU resources, which improves data security while ensuring task performance. In addition, in the present application, when the degree of fragmentation of GPU resources is high, the electronic device reallocates GPU resources, improves the utilization of GPU resources, and ensures the processing performance of tasks such as user training and reasoning.

[0157] Figure 8 For a schematic diagram of a GPU computing power scheduling device provided by an exemplary embodiment of the present application, see Figure 8 , the GPU computing power scheduling device 80 includes:

[0158] The acquisition module 81 is used to obtain the topology information of each GPU partition in the target host, and determine the performance baseline data of the GPU partition according to the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition;

[0159] A first determination module 82 is used to determine a first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition;

[0160] The second determination module 83 is used to determine the second target GPU partition corresponding to the target virtual machine according to the demand data of the running target virtual machine and the performance baseline data of each GPU partition when the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, so that the target virtual machine performs task processing based on the second target GPU partition.

[0161] In a possible implementation, the performance baseline data includes a communication factor and a discrete factor; the communication factor is used to characterize the communication performance of the GPU partition; the discrete factor is used to characterize the discrete degree of the GPU partition; the acquisition module 81 is specifically used to:

[0162] For each GPU partition, determine the topology structure and interconnection mode of the GPU partition according to the topology information, and determine the communication factor of the GPU partition according to the topology structure and the interconnection mode;

[0163] The physical location code of each GPU in the GPU partition is determined according to the topology information, and the discrete factor of the GPU partition is calculated according to the physical location code.

[0164] In a possible implementation, the acquisition module 81 is specifically configured to:

[0165] Determine the topological weight of GPU partitions based on the topological structure and interconnection mode;

[0166] The communication bandwidth of the GPU partition is measured, and the communication factor of the GPU partition is calculated based on the communication bandwidth and the topological weight.

[0167] In a possible implementation manner, the device 80 is further used for:

[0168] After calculating the communication factor of the first GPU partition, if the topology structure and interconnection mode of the second GPU partition are the same as those of the first GPU partition, assigning the communication factor of the second GPU partition to the communication factor of the first GPU partition;

[0169] If there is an isolated GPU in the GPU partition, the GPU partition is marked as a target state partition, and calculation of performance baseline data for the GPU partition is skipped.

[0170] In a possible implementation, the acquisition module 81 is specifically configured to:

[0171] According to the physical location encoding, the physical distribution data corresponding to the GPU partition is calculated, and the physical distribution data is used as the discrete factor of the GPU partition.

[0172] In a possible implementation manner, the device 80 is further used for:

[0173] Obtain the real-time discrete factor corresponding to each GPU partition, and calculate the global discrete factor corresponding to the GPU partition according to the real-time discrete factor; when the global discrete factor is greater than a preset discrete threshold, determine that the resource fragmentation information corresponding to each GPU partition in the target host meets the preset condition; the real-time discrete factor is the discrete factor corresponding to each GPU partition within a preset period; or,

[0174] Obtain task performance data of the target virtual machine; the task performance data includes at least one of actual GPU utilization, actual communication bandwidth, and actual task iteration speed; when the task performance data is less than a preset performance threshold, determine whether resource fragmentation information corresponding to each GPU partition in the target host meets preset conditions.

[0175] In a possible implementation manner, the device 80 is further used for:

[0176] Initialize the second target GPU partition according to the corresponding relationship between the target virtual machine and the second target GPU partition;

[0177] Establishing a binding relationship between each GPU in the second target GPU partition after the initialization process and the target virtual machine, and updating the GPU list of the target virtual machine;

[0178] When the second target GPU partition works normally after the initialization process, the binding relationship between the target virtual machine and the first target GPU partition is released.

[0179] In a possible implementation manner, the device 80 is further used for:

[0180] In each target virtual machine, collect the status information of the GPU used by the target virtual machine;

[0181] Status information is received in a target host, and target warning information is output when the status information meets preset warning conditions; the preset warning conditions include that the communication bandwidth of the GPU is less than a preset bandwidth threshold, the communication delay of the GPU is greater than a preset delay threshold, or the transmission error count of the GPU is greater than a preset number threshold.

[0182] The GPU computing power scheduling device 80 provided in the embodiment of the present application can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar and will not be repeated here.

[0183] Fig. 9 A schematic diagram of a GPU computing power scheduling device provided for an exemplary embodiment of the present application is shown in FIG. Fig. 9 The GPU computing power scheduling device 90 may include a processor 91 and a memory 92. Exemplarily, the processor 91 and the memory 92 are interconnected via a bus 93.

[0184] Memory 92 stores computer executable instructions;

[0185] The processor 91 executes the computer execution instructions stored in the memory 92, so that the processor 91 executes the GPU computing power scheduling method as shown in the above method embodiment.

[0186] Accordingly, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored. When the computer execution instructions are executed by a processor, they are used to implement the GPU computing power scheduling method of the above method embodiment.

[0187] Accordingly, an embodiment of the present application may also provide a computer program product, including a computer program, which, when executed by a processor, can implement the GPU computing power scheduling method shown in the above method embodiment.

[0188] The embodiment of the present application also provides a GPU computing power scheduling system, including a target host and a target virtual machine;

[0189] The target host is used to obtain the topology information of each GPU partition in the target host, and determine the performance baseline data of the GPU partition according to the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition;

[0190] The target host is used to determine a first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, and the target virtual machine is used to perform task processing through the first target GPU partition;

[0191] The target host is used to determine the second target GPU partition corresponding to the target virtual machine according to the demand data of the running target virtual machine and the performance baseline data of each GPU partition when the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions. The target virtual machine is used to perform task processing through the second target GPU partition.

[0192] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.) containing computer-usable program codes.

[0193] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0194] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0195] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0196] In a typical configuration, a computing device includes one or more processors, input / output interfaces, network interfaces, and memory.

[0197] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0198] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, Parameter Random Access Memory (PRAM), Static RAM (SRAM), Dynamic RAM (DRAM), other types of random access memory (RAM), Read-Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Flash memory or other memory technology, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0199] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0200] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A GPU computing power scheduling method, characterized in that: include: Acquire topology information of each GPU partition in the target host, and determine performance baseline data of the GPU partition according to the topology information; The performance baseline data is used to characterize the communication performance and discreteness of the GPU partition; Determine a first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition; When the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, the second target GPU partition corresponding to the target virtual machine is determined according to the demand data of the running target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition.

2. The method according to claim 1, characterized in that The performance baseline data includes a communication factor and a discrete factor; the communication factor is used to characterize the communication performance of the GPU partition; The discrete factor is used to characterize the discrete degree of the GPU partition; The determining the performance baseline data of the GPU partition according to the topology information includes: For each GPU partition, determine the topological structure and interconnection mode of the GPU partition according to the topological information, and determine the communication factor of the GPU partition according to the topological structure and the interconnection mode; The physical location code of each GPU in the GPU partition is determined according to the topology information, and the discrete factor of the GPU partition is calculated according to the physical location code.

3. The method according to claim 2, characterized in that The determining the communication factor of the GPU partition according to the topological structure and the interconnection mode includes: Determining a topological weight of the GPU partition according to the topological structure and the interconnection mode; The communication bandwidth of the GPU partition is measured, and a communication factor of the GPU partition is calculated according to the communication bandwidth and the topological weight.

4. The method according to claim 2, characterized in that: The method further comprises: After calculating the communication factor of the first GPU partition, if the topology structure and interconnection mode of the second GPU partition are the same as those of the first GPU partition, assigning the communication factor of the second GPU partition to the communication factor of the first GPU partition; If there is an isolated GPU in the GPU partition, the GPU partition is marked as a target state partition, and calculation of performance baseline data of the GPU partition is skipped.

5. The method according to claim 2, characterized in that: The calculating the discrete factor of the GPU partition according to the physical location encoding includes: According to the physical location code, physical distribution data corresponding to the GPU partition is calculated, and the physical distribution data is used as a discrete factor of the GPU partition.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Obtaining a real-time discrete factor corresponding to each GPU partition, and calculating a global discrete factor corresponding to the GPU partition according to the real-time discrete factor; determining that resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition when the global discrete factor is greater than a preset discrete threshold; the real-time discrete factor is a discrete factor corresponding to each GPU partition within a preset period; or, Obtain task performance data of the target virtual machine; the task performance data includes at least one of actual GPU utilization, actual communication bandwidth, and actual task iteration speed; when the task performance data is less than a preset performance threshold, determine whether resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition.

7. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Initializing the second target GPU partition according to the corresponding relationship between the target virtual machine and the second target GPU partition; Establishing a binding relationship between each GPU in the second target GPU partition after the initialization process and the target virtual machine, and updating the GPU list of the target virtual machine; When the second target GPU partition works normally after the initialization process, the binding relationship between the target virtual machine and the first target GPU partition is released.

8. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: In each of the target virtual machines, collecting status information of a GPU used by the target virtual machine; The status information is received in a target host, and when the status information meets a preset warning condition, a target warning information is output; the preset warning condition includes that the communication bandwidth of the GPU is less than a preset bandwidth threshold, the communication delay of the GPU is greater than a preset delay threshold, or the transmission error count of the GPU is greater than a preset number threshold.

9. A GPU computing power scheduling device, characterized in that: include: An acquisition module, used to acquire topology information of each GPU partition in the target host, and determine performance baseline data of the GPU partition according to the topology information; The performance baseline data is used to characterize the communication performance and discreteness of the GPU partition; A first determination module, configured to determine a first target GPU partition corresponding to the target virtual machine according to demand data of the target virtual machine and performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition; The second determination module is used to determine the second target GPU partition corresponding to the target virtual machine according to the demand data of the running target virtual machine and the performance baseline data of each GPU partition when the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, so that the target virtual machine performs task processing based on the second target GPU partition.

10. A GPU computing power scheduling device, characterized in that: include: Memory and processor; The memory stores computer-executable instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the GPU computing power scheduling method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the GPU computing power scheduling method described in any one of claims 1 to 8.

12. A computer program product, characterized in that It includes a computer program, which, when executed by a computer, implements the GPU computing power scheduling method as described in any one of claims 1 to 8.

13. A GPU computing power scheduling system, characterized in that: Including the target host and the target virtual machine; The target host is used to obtain topology information of each GPU partition in the target host, and determine performance baseline data of the GPU partition according to the topology information; The performance baseline data is used to characterize the communication performance and discreteness of the GPU partition; The target host is used to determine a first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, and the target virtual machine is used to perform task processing through the first target GPU partition; The target host is used to determine a second target GPU partition corresponding to the target virtual machine according to demand data of a running target virtual machine and performance baseline data of each GPU partition when resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition, and the target virtual machine is used to perform task processing through the second target GPU partition.

Citation Information

Patent Citations

  • Efficient GPU resource allocation optimization method and system

    CN111930498A

  • GPU resource management method and device

    CN113703961A

  • Distributed training task scheduling method, system and device for intelligent computing

    CN115248728A

  • GPU (Graphics Processing Unit) computing power resource scheduling method and device, equipment and medium

    CN116860391A

  • Hybrid workflow scheduling method and system in cloud computing environment and medium

    CN117519927A

Cited By

  • Processor test method, device, medium and program product

    CN120670239A