GPU computing power scheduling method, device, equipment, storage medium and program product
By obtaining the topology information and performance baseline data of the GPU partitions within the virtual machine, GPU resources are dynamically scheduled, solving the problem of GPU cards being unable to be fully interconnected in virtualization scenarios and improving communication efficiency and task processing performance.
Patent Information
- Application Number
- CN202510571609.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-05-06
AI Technical Summary
In virtualization scenarios, due to the lack of high-speed direct connection components of GPU cards from different manufacturers, multiple GPU cards in the same host cannot be fully interconnected, resulting in difficulty in GPU computing power scheduling and low communication efficiency, which in turn affects the efficiency of virtual machine task processing.
By obtaining the topology information of each GPU partition in the target host, the performance baseline data, including communication factors and discrete factors, is determined. The GPU partitions are dynamically scheduled according to the needs of the virtual machines to achieve reasonable allocation and dynamic configuration of resources and ensure the efficiency of inter-card communication.
It improves GPU resource utilization and task processing efficiency, and enhances users' task processing performance.
Smart Images

Figure CN120104346B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a GPU computing power scheduling method, apparatus, device, storage medium, and program product. Background Art
[0002] With the continuous development of technologies such as artificial intelligence (AI) and deep learning, the demand for graphics processing units (GPUs) in processes such as model training and inference has increased significantly.
[0003] To efficiently utilize GPU resources, virtualization scenarios typically deploy GPUs from multiple vendors. However, due to the lack of high-speed direct connectivity between GPU cards from different vendors, full interconnection between multiple GPUs within the same host is impossible. This makes GPU computing scheduling difficult and communication efficiency uncertain, leading to low virtual machine task processing efficiency and poor user task processing performance. Summary of the Invention
[0004] Multiple aspects of the present application provide a GPU computing power scheduling method, device, equipment, storage medium and program product, which can realize reasonable and dynamic scheduling of GPU computing power, improve communication efficiency and task processing efficiency, and also improve the user's task processing performance.
[0005] In a first aspect, an embodiment of the present application provides a GPU computing power scheduling method, comprising:
[0006] Obtaining topology information of each GPU partition in the target host, and determining performance baseline data of the GPU partition based on the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition;
[0007] Determine, based on the demand data of the target virtual machine and the performance baseline data of each GPU partition, a first target GPU partition corresponding to the target virtual machine, so that the target virtual machine performs task processing based on the first target GPU partition;
[0008] When the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, the second target GPU partition corresponding to the target virtual machine is determined based on the demand data of the target virtual machine in operation and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition.
[0009] In one possible implementation, the performance baseline data includes a communication factor and a discrete factor; the communication factor is used to characterize the communication performance of the GPU partition; the discrete factor is used to characterize the degree of discreteness of the GPU partition; and determining the performance baseline data of the GPU partition based on the topology information includes:
[0010] For each GPU partition, determine a topology structure and an interconnection mode of the GPU partition according to the topology information, and determine a communication factor of the GPU partition according to the topology structure and the interconnection mode;
[0011] The physical location code of each GPU in the GPU partition is determined according to the topology information, and the discrete factor of the GPU partition is calculated according to the physical location code.
[0012] In a possible implementation, determining the communication factor of the GPU partition according to the topology structure and the interconnection mode includes:
[0013] Determining a topological weight of the GPU partition according to the topological structure and the interconnection mode;
[0014] The communication bandwidth of the GPU partition is measured, and a communication factor of the GPU partition is calculated according to the communication bandwidth and the topology weight.
[0015] In one possible implementation, the method further includes:
[0016] After calculating the communication factor of the first GPU partition, if the second GPU partition has the same topology and interconnection method as the first GPU partition, assign the communication factor of the second GPU partition to the communication factor of the first GPU partition;
[0017] If an isolated GPU exists in the GPU partition, the GPU partition is marked as a target state partition, and calculation of performance baseline data of the GPU partition is skipped.
[0018] In a possible implementation, calculating the discrete factor of the GPU partition according to the physical location code includes:
[0019] According to the physical location code, physical distribution data corresponding to the GPU partition is calculated, and the physical distribution data is used as a discrete factor of the GPU partition.
[0020] In one possible implementation, the method further includes:
[0021] Obtaining a real-time discrete factor corresponding to each GPU partition, and calculating a global discrete factor corresponding to the GPU partition based on the real-time discrete factor; determining that resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition when the global discrete factor is greater than a preset discrete threshold; the real-time discrete factor is a discrete factor corresponding to each GPU partition within a preset period; or
[0022] Obtain task performance data of the target virtual machine; the task performance data includes at least one of actual GPU utilization, actual communication bandwidth, and actual task iteration speed; when the task performance data is less than a preset performance threshold, determine that resource fragmentation information corresponding to each GPU partition in the target host meets preset conditions.
[0023] In one possible implementation, the method further includes:
[0024] Initializing the second target GPU partition according to the correspondence between the target virtual machine and the second target GPU partition;
[0025] Establishing a binding relationship between each GPU in the initialized second target GPU partition and the target virtual machine, and updating the GPU list of the target virtual machine;
[0026] When the second target GPU partition works normally after the initialization process, the binding relationship between the target virtual machine and the first target GPU partition is released.
[0027] In one possible implementation, the method further includes:
[0028] In each of the target virtual machines, collecting status information of a GPU used by the target virtual machine;
[0029] The status information is received in a target host, and target warning information is output when the status information meets a preset warning condition; the preset warning condition includes that the communication bandwidth of the GPU is less than a preset bandwidth threshold, the communication delay of the GPU is greater than a preset delay threshold, or the transmission error count of the GPU is greater than a preset number threshold.
[0030] In a second aspect, an embodiment of the present application provides a GPU computing power scheduling device, comprising:
[0031] An acquisition module is used to obtain topology information of each GPU partition in the target host and determine performance baseline data of the GPU partition based on the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition;
[0032] A first determining module is configured to determine a first target GPU partition corresponding to the target virtual machine based on the demand data of the target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition;
[0033] The second determination module is used to determine the second target GPU partition corresponding to the target virtual machine based on the demand data of the running target virtual machine and the performance baseline data of each GPU partition when the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, so that the target virtual machine performs task processing based on the second target GPU partition.
[0034] In one possible implementation, the performance baseline data includes a communication factor and a discrete factor; the communication factor is used to characterize the communication performance of the GPU partition; the discrete factor is used to characterize the degree of discreteness of the GPU partition; the acquisition module is specifically configured to:
[0035] For each GPU partition, determine a topology structure and an interconnection mode of the GPU partition according to the topology information, and determine a communication factor of the GPU partition according to the topology structure and the interconnection mode;
[0036] The physical location code of each GPU in the GPU partition is determined according to the topology information, and the discrete factor of the GPU partition is calculated according to the physical location code.
[0037] In a possible implementation, the acquisition module is specifically configured to:
[0038] Determining a topological weight of the GPU partition according to the topological structure and the interconnection mode;
[0039] The communication bandwidth of the GPU partition is measured, and a communication factor of the GPU partition is calculated according to the communication bandwidth and the topology weight.
[0040] In one possible embodiment, the device is further used for:
[0041] After calculating the communication factor of the first GPU partition, if the second GPU partition has the same topology and interconnection method as the first GPU partition, assign the communication factor of the second GPU partition to the communication factor of the first GPU partition;
[0042] If an isolated GPU exists in the GPU partition, the GPU partition is marked as a target state partition, and calculation of performance baseline data of the GPU partition is skipped.
[0043] In a possible implementation, the acquisition module is specifically configured to:
[0044] According to the physical location code, physical distribution data corresponding to the GPU partition is calculated, and the physical distribution data is used as a discrete factor of the GPU partition.
[0045] In one possible embodiment, the device is further used for:
[0046] Obtaining a real-time discrete factor corresponding to each GPU partition, and calculating a global discrete factor corresponding to the GPU partition based on the real-time discrete factor; determining that resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition when the global discrete factor is greater than a preset discrete threshold; the real-time discrete factor is a discrete factor corresponding to each GPU partition within a preset period; or
[0047] Obtain task performance data of the target virtual machine; the task performance data includes at least one of actual GPU utilization, actual communication bandwidth, and actual task iteration speed; when the task performance data is less than a preset performance threshold, determine that resource fragmentation information corresponding to each GPU partition in the target host meets preset conditions.
[0048] In one possible embodiment, the device is further used for:
[0049] Initializing the second target GPU partition according to the correspondence between the target virtual machine and the second target GPU partition;
[0050] Establishing a binding relationship between each GPU in the initialized second target GPU partition and the target virtual machine, and updating the GPU list of the target virtual machine;
[0051] When the second target GPU partition works normally after the initialization process, the binding relationship between the target virtual machine and the first target GPU partition is released.
[0052] In one possible embodiment, the device is further used for:
[0053] In each of the target virtual machines, collecting status information of a GPU used by the target virtual machine;
[0054] The status information is received in a target host, and target warning information is output when the status information meets a preset warning condition; the preset warning condition includes that the communication bandwidth of the GPU is less than a preset bandwidth threshold, the communication delay of the GPU is greater than a preset delay threshold, or the transmission error count of the GPU is greater than a preset number threshold.
[0055] In a third aspect, an embodiment of the present application provides a GPU computing power scheduling device, including: a memory and a processor;
[0056] The memory stores computer-executable instructions;
[0057] The processor executes the computer execution instructions stored in the memory, so that the processor executes the GPU computing power scheduling method described in any one of the first aspects.
[0058] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer execution instructions, which, when executed by a processor, are used to implement the GPU computing power scheduling method described in any one of the first aspects.
[0059] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the GPU computing power scheduling method shown in any one of the first aspects.
[0060] In a sixth aspect, an embodiment of the present application provides a GPU computing power scheduling system, including a target host and a target virtual machine;
[0061] The target host is used to obtain topology information of each GPU partition in the target host, and determine performance baseline data of the GPU partition based on the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition;
[0062] The target host is used to determine a first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, and the target virtual machine is used to perform task processing through the first target GPU partition;
[0063] The target host is used to determine a second target GPU partition corresponding to the target virtual machine based on demand data of a running target virtual machine and performance baseline data of each GPU partition when resource fragmentation information corresponding to each GPU partition in the target host meets preset conditions. The target virtual machine is used to perform task processing through the second target GPU partition.
[0064] In an embodiment of the present application, topology information of each GPU partition in a target host is obtained, and performance baseline data of the GPU partition is determined based on the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition; based on the demand data of the target virtual machine and the performance baseline data of each GPU partition, a first target GPU partition corresponding to the target virtual machine is determined, so that the target virtual machine performs task processing based on the first target GPU partition; when the resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition, based on the demand data of the target virtual machine in operation and the performance baseline data of each GPU partition, a second target GPU partition corresponding to the target virtual machine is determined, so that the target virtual machine performs task processing based on the second target GPU partition. In the present application, based on the topology information of each GPU partition in the target host, the performance baseline data of each GPU partition is determined, and based on the demand data of the target virtual machine and the performance baseline data, the first target GPU partition corresponding to the target virtual machine is determined. Subsequently, during system operation, if it is detected that the resource fragmentation degree of multiple GPU partitions in the target host is high, the second target GPU partition corresponding to each target virtual machine can be determined, so that the target virtual machine performs task processing based on the second target GPU partition. In this way, GPU resources are allocated based on the topological information of each GPU, and dynamic and reasonable configuration of GPU resources is achieved when the degree of resource fragmentation is high, which can ensure the efficiency of inter-GPU communication, thereby improving task processing efficiency and user task processing performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0066] Figure 1 This is an X-Link topology structure of a host GPU card in the related art;
[0067] Figure 2 A flowchart of the steps of a GPU computing power scheduling method provided in an embodiment of the present application;
[0068] Figure 3 A schematic diagram of a topological structure provided in an embodiment of the present application;
[0069] Figure 4 A schematic diagram of target host GPU resource fragmentation provided in an embodiment of the present application;
[0070] Figure 5 A schematic diagram of the steps of another GPU computing power scheduling method provided in an embodiment of the present application;
[0071] Figure 6A schematic diagram of a state after dynamic allocation of GPU resources provided in an embodiment of the present application;
[0072] Figure 7 A schematic diagram of a system architecture for GPU computing power scheduling provided by an exemplary embodiment of the present application;
[0073] Figure 8 A schematic diagram of the structure of a GPU computing power scheduling device provided by an exemplary embodiment of the present application;
[0074] Figure 9 A schematic diagram of the structure of a GPU computing power scheduling device provided as an exemplary embodiment of the present application.
[0075] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0076] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0077] The following is an explanation of the professional terms involved in this application:
[0078] Graphics Processing Unit (GPU): used to accelerate computing tasks.
[0079] Elastic GPU Service (EGS): Provides GPU-accelerated computing capabilities, enabling ready-to-use and elastic scaling of GPU computing resources.
[0080] NVLink / NVSwitch: NVIDIA's high-speed GPU interconnect technology. NVLink is a high-speed GPU interconnect technology that provides high bandwidth; NVSwitch is a switching architecture used to achieve full interconnection of multiple GPUs.
[0081] X-Link: An interconnection technology between GPU cards from domestic manufacturers similar to NVLink.
[0082] Kernel-based Virtual Machine (KVM): A virtualization infrastructure for Linux systems.
[0083] Pass-through: Allocate physical hardware devices directly to virtual machines.
[0084] Host: A physical server that runs a virtual machine hypervisor.
[0085] Virtual Machine (VM): A virtual computer system running on a host computer.
[0086] High-speed serial expansion bus (Peripheral Component Interconnect Express, PCle): used to connect peripheral devices.
[0087] With the rapid growth of AI and deep learning tasks, the demand for GPUs has increased significantly. To efficiently utilize resources, GPUs from various manufacturers can be deployed in virtualized environments based on EGS. As an AI infrastructure platform, EGS offers high flexibility, isolation, and reliability. In virtualized environments, direct virtualization technology plays a key role in resource management and allocation. For GPU computing resources, the cloud platform pools and encapsulates GPU resources on physical machines and directly distributes these GPU resources to different virtual machines based on user needs.
[0088] In related technologies, for hosts that use NVSwitch and NVLink to achieve full interconnection of multiple GPUs on a single machine, there is a high-speed NVLink direct connection between any two GPUs, and the virtualization management and control side only needs to allocate the corresponding number of GPU cards based on computing power requirements. However, in virtualization scenarios where multiple domestic GPUs are deployed, there is a lack of components similar to NVSwitch, and full interconnection of multiple GPUs on a single machine is not supported. The topology and interconnection methods of GPUs within the same host are relatively complex, making GPU computing power scheduling difficult and unable to ensure communication efficiency, which in turn leads to low virtual machine task processing efficiency.
[0089] For example, Figure 1 This is an X-Link topology structure of a host GPU card in the related art. Figure 1As shown, in the virtualization scenario, the same host includes two central processing units (CPUs), namely cpu0 and cpu1. The host uses GPUs from different manufacturers, with a total of 16 GPU cards (GPU0 to GPU15). Since there is no component similar to NVSwitch, the GPU cards in a single machine can only be fully interconnected based on X-Link within a small partition (usually between specific 4 GPUs). Moreover, within the same host, there are three types of interconnection between GPU cards: dual-line interconnection (2X-Link), single-line interconnection (1X-Link), and PCle. For example, there is 2X-Link between GPU0 and GPU1, and 1X-Link between GPU2 and GPU8. Two GPUs that are not directly connected by X-Link can only communicate through PCle, such as GPU0 and GPU8. By Figure 1 It can be seen that in the virtualization scenario, the interconnection method between various GPU cards in the same host is relatively complex. At the same time, due to the diversity of topological states, it is difficult to schedule GPU computing power, which has a great impact on the efficiency of communication between cards, and thus also affects the performance of user model training and inference tasks.
[0090] In order to solve the above problems, the present application provides a GPU computing power scheduling method, device, equipment, storage medium and program product, which obtains the topology information of each GPU partition in the target host and determines the performance baseline data of the GPU partition based on the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition; based on the demand data of the target virtual machine and the performance baseline data of each GPU partition, the first target GPU partition corresponding to the target virtual machine is determined, so that the target virtual machine performs task processing based on the first target GPU partition; when the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, the second target GPU partition corresponding to the target virtual machine is determined based on the demand data of the target virtual machine in operation and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition. In this application, based on the topology information of each GPU partition in the target host, the performance baseline data of each GPU partition is determined, and the first target GPU partition corresponding to the target virtual machine is determined based on the demand data of the target virtual machine and the performance baseline data. Later, during the operation of the system, if it is detected that the resource fragmentation degree of multiple GPU partitions in the target host is high, the second target GPU partition corresponding to each target virtual machine can be determined, and the target virtual machine performs task processing based on the second target GPU partition. In this way, GPU resources are allocated based on the topological information of each GPU, and dynamic and reasonable configuration of GPU resources is achieved when the degree of resource fragmentation is high, which can ensure the efficiency of inter-GPU communication, thereby improving task processing efficiency and user task processing performance.
[0091] The technical solutions shown in this application are described in detail below through specific embodiments. It should be noted that the following embodiments can exist independently or in combination with each other, and the same or similar contents will not be repeated in different embodiments.
[0092] Figure 2 This is a flowchart of the steps of a GPU computing power scheduling method provided in this application embodiment. Figure 2 , the GPU computing power scheduling method may include:
[0093] S201. Acquire topology information of each GPU partition in the target host, and determine performance baseline data of the GPU partition based on the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition.
[0094] The execution subject of the embodiment of the present application can be an electronic device, or a GPU computing power scheduling device provided in the electronic device. The GPU computing power scheduling device can be implemented by software, or by a combination of software and hardware. For ease of understanding, the following description will be made by taking the execution subject as an electronic device as an example. The electronic device can specifically refer to a server or a cloud platform, etc., or it can refer to a target host (such as a GPU centralized control platform in the target host, etc.), which is not limited in the embodiment of the present application.
[0095] In the embodiments of the present application, a target host may refer to a host machine or physical machine running a virtual machine. A GPU partition may refer to the various partitions that may be formed by the GPU in the target host. Each partition may include N GPU cards, where N is a positive integer. Since X-Link is typically used by different manufacturers to interconnect four cards, the value of N may be 4. Of course, based on actual task requirements, an electronic device may also be partitioned according to two or eight GPU cards, and this embodiment of the present application does not limit this.
[0096] The topology information may refer to the topology information corresponding to each GPU partition, and specifically may include the topology structure, interconnection mode, and location information of multiple GPUs in the GPU partition. The topology structure may refer to the topological space structure within the GPU partition, and specifically may include irregular topology structure and regular topology structure. For example, Figure 3 This is a schematic diagram of a topological structure provided in an embodiment of the present application. Figure 3 As shown in (a), the topological structure of the GPU partition is an irregular topological structure, which is not a regular rectangle in space. Figure 3 As shown in (b), the topological structure of the GPU partition is a regular topological structure, forming a regular rectangle in space.
[0097] The interconnection method may refer to the connection method between GPUs within a GPU partition, and may specifically refer to dual-line interconnection, single-line interconnection, PCIe interconnection, etc. The location information may refer to the actual physical location of each GPU card in the hardware topology, and may specifically include rack location, motherboard slot, and Non-Uniform Memory Access (NUMA) node. Of course, the topology information may also include other information, such as hardware bandwidth information, load information, and heat dissipation partition information, which is not limited in the embodiments of the present application.
[0098] Performance baseline data can be used to characterize the communication performance of each GPU partition and can be used to guide the allocation strategy of electronic devices for GPU partitions. This performance baseline data may include the communication factor (ComFactor) and the discrete factor (FragFactor) corresponding to the GPU partition. The communication factor can be used to characterize the communication performance of GPU partitions with different topologies; a larger communication factor indicates higher communication efficiency for the GPU partition. The discrete factor can be used to characterize the degree of discretization of GPU partitions within the same topology. A larger discrete factor indicates greater resource fragmentation caused by the GPU partition.
[0099] In this embodiment of the present application, the electronic device obtains topology information for each GPU partition within the target host and can then calculate performance baseline data for each GPU partition based on this topology information. When the electronic device needs to allocate GPU resources to the target virtual machine, it can allocate a GPU partition with higher communication efficiency based on this performance baseline data, thereby ensuring the task processing performance of the target virtual machine.
[0100] S202 : Determine a first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition.
[0101] In an embodiment of the present application, the demand data may refer to the GPU resources actually required by the target virtual machine, specifically including the number of GPUs, etc. The first target GPU partition may refer to the GPU partition allocated to the target virtual machine. After determining the performance baseline data of the GPU partition, the electronic device may determine the first target GPU partition based on the demand data of the target virtual machine. Specifically, the electronic device may select an idle GPU partition with the largest communication factor and the smallest discrete factor as the first target GPU partition. This ensures that the GPU partition currently selected by the target virtual machine has high communication efficiency and low fragmentation.
[0102] In one possible implementation, based on the performance baseline data of each GPU partition, the electronic device can sort the GPU partitions of different topologies in descending order according to the communication factor, and sort the GPU partitions of the same topology in descending order according to the discrete factor. When GPU resources need to be allocated to the target virtual machine, the electronic device can select the idle GPU partition ranked first as the first target GPU partition, which can ensure the communication efficiency and task processing efficiency of the target virtual machine and improve the user task processing performance.
[0103] S203. When the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, the second target GPU partition corresponding to the target virtual machine is determined according to the demand data of the target virtual machine in operation and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition.
[0104] In an embodiment of the present application, resource fragmentation information may refer to the overall degree of GPU resource fragmentation for each GPU partition within the target host. Specifically, it may include parameters such as the real-time global discrete factor of each GPU partition within the target host (characterizing the overall degree of discreteness of the GPU partition) and the real-time task performance data of the target virtual machine (e.g., actual GPU utilization, actual communication bandwidth, and actual iteration speed). A preset condition may refer to a pre-set trigger condition for GPU resource reallocation. For example, the preset condition may refer to a global discrete factor greater than a preset discrete threshold, actual GPU utilization less than a preset utilization threshold, or actual iteration speed less than a preset speed threshold. The embodiments of the present application do not limit the specific type of the preset condition.
[0105] During system operation, due to the multiple creation and release of target virtual machines by users on the target host, the fragmentation of GPU card resources in the target host gradually increases, and the effective utilization of resources gradually decreases. When the fragmentation of GPU resources is too high, although there are enough idle resources, these idle resources are often scattered across different GPU partitions, resulting in the target virtual machine being unable to be allocated to a GPU partition with higher communication efficiency, and the utilization of GPU resources is low. For example, Figure 4 A schematic diagram of target host GPU resource fragmentation provided in an embodiment of the present application. Figure 4 As shown, among the 16 GPU cards in the target host, 8 GPU cards are occupied. Although there are 8 idle GPU cards, the communication efficiency of the GPU partition composed of these 8 GPU cards is not high. In this way, the high degree of resource fragmentation leads to the inability to fully utilize the computing power of the GPU resources in the target host, affecting the communication efficiency and task processing performance.
[0106] In an embodiment of the present application, the electronic device can monitor the real-time operating status of each GPU partition in the target host and collect resource fragmentation information of the GPU partition in the target host. When the resource fragmentation information meets the preset conditions, the electronic device can determine that the degree of GPU resource fragmentation in the target host is high at this time, and it is necessary to reallocate the GPU resources corresponding to each target virtual machine to improve the utilization rate of GPU resources and ensure the user's task processing performance. The electronic device can determine at least one target virtual machine that is currently in operation, and then obtain the demand data corresponding to each target virtual machine. According to the demand data and the performance baseline data of each GPU partition, the GPU resources corresponding to the target virtual machine are reallocated to obtain the second target GPU partition corresponding to each target virtual machine. Subsequently, the target virtual machine can be bound to the second target GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition, thereby realizing dynamic and reasonable allocation of GPU resources of the target virtual machine in a virtualization scenario, and ensuring communication efficiency and task processing efficiency.
[0107] In an embodiment of the present application, an electronic device obtains topology information of each GPU partition in a target host and determines performance baseline data of the GPU partition based on the topology information; determines a first target GPU partition corresponding to the target virtual machine based on the demand data of the target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition; and, when resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition, determines a second target GPU partition corresponding to each target virtual machine based on the demand data of multiple target virtual machines in operation and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition. In this way, in the present application, based on the topology information of each GPU partition in the target host, the performance baseline data of each GPU partition is determined, and the first target GPU partition corresponding to the target virtual machine is determined based on the demand data of the target virtual machine and the performance baseline data. Subsequently, during operation, if it is detected that the resource fragmentation level of multiple GPU partitions in the target host is high, the second target GPU partition corresponding to each target virtual machine can be determined, so that the target virtual machine performs task processing based on the second target GPU partition. In this way, GPU resources are allocated based on the topological information of each GPU, and dynamic and reasonable configuration of GPU resources is achieved when the degree of resource fragmentation is high, which can ensure the efficiency of inter-GPU communication, thereby improving task processing efficiency and user task processing performance.
[0108] Based on the above embodiments, Figure 5 This is a flowchart of another GPU computing power scheduling method provided in this application embodiment. Figure 5, the GPU computing power scheduling method may include:
[0109] S501. Obtain topology information of each GPU partition in the target host; for each GPU partition, determine the topology structure and interconnection mode of the GPU partition according to the topology information, and determine the communication factor of the GPU partition according to the topology structure and interconnection mode.
[0110] In an embodiment of the present application, performance baseline data for each possible GPU partition within a target host can be used to characterize the communication performance and degree of discreteness of the GPU partition. This performance baseline data can specifically include a communication factor and a discrete factor for each GPU partition, wherein the communication factor characterizes the communication performance of the GPU partition, and the discrete factor characterizes the degree of discreteness of the GPU partition. Specifically, the electronic device can parse and read the topology structure and interconnection method of the GPU partition based on the topology information of the GPU partition, and then calculate the communication factor of the GPU partition based on the topology structure and interconnection method of the GPU partition, thereby calculating and evaluating the communication performance of GPU partitions with different topological forms.
[0111] In one possible implementation, the communication factor in the performance baseline data may be calculated as follows:
[0112] The topological weight of the GPU partition is determined based on the topological structure and the interconnection mode; the communication bandwidth of the GPU partition is measured, and the communication factor of the GPU partition is calculated based on the communication bandwidth and the topological weight.
[0113] In the embodiments of the present application, the topology weight may refer to the topology structure of the GPU partition and the weight information corresponding to the interconnection method. Due to the significant difference in communication bandwidth between different interconnection methods, for example, the topology weight of the X-Link direct connection method is greater than the topology weight corresponding to the PCIe direct connection method. The communication bandwidth may refer to the communication bandwidth actually tested for the current GPU partition.
[0114] Specifically, when calculating the communication factor of each GPU partition, the electronic device can first determine the topological structure and interconnection method of the GPU partition. Specifically, the GPU management tool can be used to obtain the topological information of each GPU partition, and the topological structure and interconnection method of the GPU partition can be determined based on the topological information. For example, an adjacency matrix can be used to represent the interconnection method between GPUs. Afterwards, the electronic device can measure the communication bandwidth of the GPU partition through the GPU communication library and topological isolation tools, for example, the point-to-point bandwidth between any two cards in the GPU partition can be measured, and the communication bandwidth of the GPU partition can be obtained by calculating the average value. Afterwards, the electronic device can determine the topological weight corresponding to the GPU partition based on the topological structure and interconnection method of the GPU partition. Afterwards, the electronic device can calculate the communication factor of the current GPU partition based on the communication bandwidth and topological weight. For example, the electronic device calculates the communication factor of the GPU partition according to the following formulas (1) to (3):
[0115]
[0116]
[0117]
[0118] In the above formula, ComFator1 can refer to the first component corresponding to the communication factor, which is used to represent the ratio of the actual measured communication bandwidth of the GPU partition to the theoretical bandwidth (the theoretical bandwidth value calculated based on the interconnection method). ComFator2 can refer to the second component corresponding to the communication factor, which is used to represent the communication bandwidth of the GPU partition after weighted processing; BW i,j It can refer to the actual measured communication bandwidth between GPU card i and GPU card j, W i,j It refers to the topological weight corresponding to the interconnection method between GPU card i and GPU card j. α and β refer to the weights corresponding to the first component and the second component, respectively, which are used to balance the influence of different components. For example, α can be 0.7 and β can be 0.3. In addition, ComFator1 and ComFator2 can be normalized values for ease of calculation and comparison. In the embodiment of the present application, based on the above formulas (1) to (3), the electronic device can calculate the first component and the second component of the GPU partition communication factor, and then calculate the communication factor by weighted summation, which can balance the influencing factors such as communication bandwidth and interconnection method, and realize a reasonable and accurate evaluation of the GPU partition communication factor.
[0119] It should be noted that the above formulas (1) to (3) are merely examples, and the communication factor of the GPU partition may also be calculated in other ways, which is not limited in the embodiments of the present application.
[0120] In an embodiment of the present application, when the target host has a large number of GPU cards, a large number of GPU partitions may be formed in the target host. To reduce the amount of performance baseline data calculation, the electronic device may adopt a certain pruning strategy when calculating the performance baseline data of each GPU partition. In one possible implementation, the GPU computing power scheduling method may further include the following steps:
[0121] After calculating the communication factor of the first GPU partition, if the second GPU partition has the same topology and interconnection method as the first GPU partition, the communication factor of the second GPU partition is assigned to the communication factor of the first GPU partition; if there is an isolated GPU in the GPU partition, the GPU partition is marked as a target state partition, and the calculation of the performance baseline data of the GPU partition is skipped.
[0122] In the embodiments of the present application, if the topology and interconnection methods of two GPU partitions are identical, the communication factors of the two GPU partitions can be considered identical, and the electronic device only needs to calculate the communication factor of one of the GPU partitions. That is, for a second GPU partition that has the same topology and interconnection method as the first GPU partition, the electronic device can directly reuse the communication factor of the first GPU partition as the communication factor of the second GPU partition. This eliminates the need for repeated testing, reduces the number of test cases, and reduces unnecessary computational overhead.
[0123] An isolated GPU can be a GPU that lacks a direct X-Link connection to other GPUs. In this case, the isolated GPU is essentially disconnected and has low communication efficiency. The electronic device can directly mark the GPU partition containing the isolated GPU as a target state partition, such as a "partition with poor communication efficiency." The electronic device can skip calculating the performance baseline data for the target state partition or set a minimum communication factor for the target state partition, although this is not limited in this embodiment.
[0124] In addition, the electronic device can also establish a GPU partition performance database corresponding to the target host, storing the performance baseline data of each possible GPU partition for subsequent reuse, avoiding repeated calculation of the performance baseline data by the electronic device, and saving computing resources. In an embodiment of the present application, when calculating the performance baseline data of the GPU partition, for GPU partitions with the same topology and the same interconnection method, the electronic device can select one of the GPU partitions to calculate the communication factor, and then assign the communication factor to the other GPU partitions, avoiding repeated calculation of the communication factor and reducing computing overhead; for GPU partitions with isolated GPUs, the electronic device directly marks them as target state partitions, without calculating the performance baseline data of the target state partitions, thus avoiding the calculation of GPU partitions with lower communication efficiency and further saving computing resources.
[0125] S502: Determine the physical location code of each GPU in the GPU partition according to the topology information, and calculate the discrete factor of the GPU partition according to the physical location code.
[0126] In the embodiment of the present application, the physical location code may refer to the physical distribution code corresponding to each GPU card in the GPU partition, which can be determined based on the actual physical location of the GPU card. i ). After determining the topology information of each GPU partition, the electronic device can determine the physical location code of each GPU card in the GPU partition based on the location information in the topology information. It can then calculate the discrete factor of the GPU partition based on the physical location code, effectively evaluating the degree of discretization of the GPU cards in the GPU partition.
[0127] In one possible implementation, the discrete factor can be calculated as follows:
[0128] According to the physical location code, the physical distribution data corresponding to the GPU partition is calculated, and the physical distribution data is used as the discrete factor of the GPU partition.
[0129] In an embodiment of the present application, the physical distribution data may refer to the physical distribution discreteness data corresponding to each GPU card in the GPU partition, and specifically may refer to the physical distribution variance or physical distribution standard deviation of each GPU card in the GPU partition. After determining the physical position code corresponding to each GPU in the GPU partition, the electronic device may calculate the physical distribution data corresponding to the GPU partition and use the physical distribution data as a discrete factor of the GPU partition to achieve an effective evaluation of the degree of discretization of the GPU partition. For example, the electronic device may calculate the physical distribution data of the GPU partition according to the following formula (4):
[0130]
[0131] In the above formula (4), n is the number of GPU cards in the GPU partition, and μ is the average value of the physical location codes of each GPU card in the GPU partition. After obtaining the physical location codes of each GPU card in the GPU partition, the electronic device obtains the physical distribution data of the GPU partition by calculating the variance, and then obtains the discrete factor of the GPU partition, which can accurately characterize the discrete degree of distribution of each GPU in the GPU partition. For example, if the GPU cards in a GPU partition belong to the same rack, the physical location codes of the GPU cards in the GPU partition are relatively small, the variance is relatively small, and the discrete factor is also relatively small; if the GPU cards in a GPU partition belong to different racks, the physical location codes of the GPU cards in the GPU partition are relatively large, the variance is relatively large, and the discrete factor is also relatively large. Since the communication bandwidth is affected by the physical connection distance, the closer the physical location, the higher the communication efficiency. For example, the communication efficiency of GPU cards in the same rack is greater than the communication efficiency of GPU cards across racks. When the electronic device subsequently allocates GPU resources to the target virtual machine, it can give priority to allocating GPU partitions with smaller discrete factors to ensure the user's communication efficiency and task performance during task processing.
[0132] Of course, the above formula (4) is only an example. The electronic device can also use other methods to calculate the physical distribution data of the GPU partition, and then obtain the discrete factor of the GPU partition. The specific selection can be based on actual needs, and the embodiment of the present application does not limit this.
[0133] S503 : Determine a first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition.
[0134] In an embodiment of the present application, after determining the communication factor and discrete factor for each GPU partition, the electronic device can store performance baseline data for each GPU partition. When GPU resources need to be allocated to a target virtual machine, the electronic device can determine the first target GPU partition corresponding to the target virtual machine based on the target virtual machine's demand data and the performance baseline data for each GPU partition, thereby ensuring the target virtual machine's communication efficiency and task processing performance, thereby improving the user experience.
[0135] S504. When the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, the second target GPU partition corresponding to each target virtual machine is determined based on the demand data of multiple target virtual machines in operation and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition.
[0136] In embodiments of the present application, during system operation, multiple users may create and release target virtual machines multiple times for the same target host, gradually increasing the degree of GPU resource fragmentation within the target host. When the resource fragmentation information corresponding to the GPU partitions within the target host meets preset conditions, the electronic device can dynamically reallocate GPU resources to the running target virtual machine, achieving efficient utilization of GPU resources.
[0137] In a possible implementation, whether the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions can be determined in the following manner:
[0138] Obtaining the real-time discrete factor corresponding to each GPU partition, and calculating the global discrete factor corresponding to the GPU partition based on the real-time discrete factor; when the global discrete factor is greater than a preset discrete threshold, determining that the resource fragmentation information corresponding to each GPU partition in the target host meets the preset condition; the real-time discrete factor is the discrete factor corresponding to each GPU partition within a preset period; or,
[0139] Obtain task performance data of the target virtual machine; the task performance data includes at least one of actual GPU utilization, actual communication bandwidth, and actual task iteration speed; when the task performance data is less than a preset performance threshold, determine that resource fragmentation information corresponding to each GPU partition in the target host meets preset conditions.
[0140] In an embodiment of the present application, the real-time discrete factor may refer to the real-time discrete factor of the GPU partition in the target host. The global discrete factor may refer to the average real-time discrete factor corresponding to each GPU partition in the target host. The preset discrete threshold may refer to a preset discrete factor critical value for GPU resource reallocation, which may specifically refer to 0.6, 0.5 or 0.4, etc., and is not limited to this in an embodiment of the present application. Specifically, the electronic device may monitor the discrete factors of each GPU partition in the target host in real time, obtain the real-time discrete factors corresponding to each GPU partition, and then calculate the average value of the real-time discrete factors of each GPU partition to obtain the global discrete factor of each GPU partition in the target host. If the global discrete factor is greater than the preset discrete threshold, the electronic device may determine that the degree of discretization of the GPU resources in the target host is high at this time, and the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions.
[0141] The task performance data may refer to the actual performance data of the target virtual machine, which may specifically include the actual GPU utilization, the actual communication bandwidth, and the actual iteration speed of the task. The preset performance threshold may refer to the preset critical value of the task performance data for GPU resource reallocation, which may specifically include the GPU utilization threshold, the communication bandwidth threshold, and the iteration speed threshold. The electronic device may collect the actual performance data of the target virtual machine in real time. When the actual performance data is less than the preset performance threshold, for example, the actual GPU utilization is less than the GPU utilization threshold, or the actual communication bandwidth is less than the communication bandwidth threshold, or the actual iteration speed of the task is less than the iteration speed threshold, the electronic device may determine that the degree of fragmentation of the GPU resources is high at this time, and the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions.
[0142] It should be noted that the judgment method for determining whether the above two types of resource fragmentation information meet the preset conditions is an "or" relationship, that is, as long as any one of them is met, the electronic device can reallocate GPU resources to the target virtual machine. Of course, the preset conditions may also include other conditions. For example, the preset conditions may also include the number of consecutive times that the scheduler in the target host fails to allocate GPU partitions with a communication factor that meets the standard (greater than a preset communication threshold), which is greater than or equal to a preset number threshold (for example, three consecutive times); or, the preset conditions may also include the target virtual machine task iteration speed drop value being greater than a preset speed change threshold, etc. The embodiments of the present application do not limit the specific types of preset conditions.
[0143] In an embodiment of the present application, when a target host is used for a long time, the GPU resources allocated to the target virtual machine may be highly fragmented. The electronic device can detect and calculate the global discrete factor and task performance data within the target host in real time. If the global discrete factor is greater than a preset discrete threshold, or the task performance data is less than a preset performance threshold, the electronic device can determine that the resource fragmentation information corresponding to each GPU partition within the target host meets a preset condition. The electronic device can reallocate GPU resources to each target virtual machine to ensure the communication efficiency and task processing performance of each target virtual machine.
[0144] S505 , initializing the second target GPU partition according to the correspondence between the target virtual machine and the second target GPU partition; establishing a binding relationship between each GPU in the initialized second target GPU partition and the target virtual machine, and updating the GPU list of the target virtual machine.
[0145] S506: If the second target GPU partition works normally after the initialization process, release the binding relationship between the target virtual machine and the first target GPU partition.
[0146] In an embodiment of the present application, when the degree of GPU resource fragmentation within the target host is high, despite sufficient idle resources, the target virtual machine cannot be allocated to a GPU partition with theoretically higher communication efficiency, resulting in low GPU resource utilization. The electronic device can, upon detecting that the resource fragmentation information corresponding to each GPU partition within the target host meets a preset condition, reallocate GPU resources to the target virtual machine, thereby improving resource utilization and ensuring communication efficiency. Specifically, the electronic device can first determine the demand data of the target virtual machines in operation and sort the target virtual machines in order of demand data from smallest to largest; then, based on the performance baseline data of each GPU partition, the electronic device can sequentially determine the second target GPU partition corresponding to each target virtual machine in a fully idle multi-card topology, thereby marking the corresponding relationship between the target virtual machine and the second target GPU partition. Subsequently, the electronic device can perform a GPU resource allocation operation based on the corresponding relationship between the target virtual machine and the second target GPU partition. This process can be performed during a period when the target virtual machine has a low task load, minimizing the impact on the user and ultimately achieving the binding of the target virtual machine to the second target CPU partition.
[0147] During the reconfiguration process, the electronic device can first initialize the second target GPU partition corresponding to the target virtual machine. This initialization process can include resource reservation and hot migration support checks, etc., to achieve resource locking and ensure the normal execution of the reconfiguration. Afterwards, the electronic device can establish a binding relationship between the target virtual machine and each GPU in the second target GPU partition after the initialization process. Specifically, the electronic device can add each GPU in the second target GPU partition to the target virtual machine using GPU dynamic addition instructions. At the same time, the electronic device can update the GPU list corresponding to the target virtual machine to achieve GPU state synchronization. If each GPU in the second target GPU partition after the initialization process is operating normally, the electronic device can release the binding relationship between the target virtual machine and the first target GPU partition, completing the reconfiguration process of the target virtual machine. In this way, when the target host GPU resource fragmentation is high, the electronic device re-determines the second target GPU partition corresponding to the target virtual machine and establishes a binding relationship between the target virtual machine and the GPUs in the second target GPU partition, realizing dynamic adjustment of GPU resources in a virtualization scenario, which can not only ensure communication efficiency but also improve resource utilization.
[0148] For example, Figure 6 A schematic diagram of the state of a GPU resource after dynamic allocation provided by an embodiment of the present application. Figure 4 and Figure 6From the comparison, it can be seen that when the resource fragmentation information of the GPU partition in the target host meets the preset conditions, the GPU resource utilization in the target host is not high. The electronic device can re-determine the second target GPU partition corresponding to each target virtual machine based on the demand data of the running target virtual machine and the performance baseline data of each GPU partition, and then gradually adjust the GPU resources bound to the target virtual machine by performing configuration changes, thereby reducing the degree of resource fragmentation in the target host and improving the utilization of GPU resources.
[0149] S507. In each target virtual machine, collect status information of the GPU used by the target virtual machine; receive the status information in the target host, and output target warning information when the status information meets preset warning conditions; the preset warning conditions include that the communication bandwidth of the GPU is less than a preset bandwidth threshold, the communication delay of the GPU is greater than a preset delay threshold, or the transmission error count of the GPU is greater than a preset number threshold.
[0150] In the embodiments of the present application, status information may refer to information such as the actual GPU usage and connection status during operation, and may specifically include the actual X-Link bandwidth, communication delay, and transmission error count. Preset warning conditions may refer to pre-set alarm conditions for abnormal GPU status, and may specifically include the communication bandwidth being less than a preset bandwidth threshold, the communication delay being greater than a preset delay threshold, or the transmission error count being greater than a preset number threshold. Other alarm conditions may also be included, and are not limited in the embodiments of the present application. Target warning information may refer to alarm prompt information when an abnormal GPU status exists.
[0151] In related technologies, the target host's ability to directly monitor the GPU status in a virtualized scenario is relatively limited, and it is impossible to achieve real-time and effective detection of the GPU.
[0152] In an embodiment of the present application, a proxy tool (GpuBuddy) may be installed within the target virtual machine. This proxy tool can collect real-time status information about the GPU within the target virtual machine. The target virtual machine can then use this proxy tool to send this status information to the target host via a target interface. For example, a Quick Emulator Guest Agent (QEMU Guest Agent) can be used to communicate with the target host via sockets. The target host receives this status information and, if the status information meets preset warning conditions, outputs a target warning message. This facilitates timely handling of abnormal GPUs and ensures the safe and stable operation of the user's target virtual machine tasks.
[0153] It should be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0154] Based on any of the above embodiments, Figure 7 A schematic diagram of a GPU computing power scheduling system architecture provided by an exemplary embodiment of this application. Figure 7 As shown, the target host includes a centralized GPU management platform (GPUMaster) and GPU physical resources (including GPU1, GPU2, GPU3, and GPUn). The GPU centralized management platform includes a scheduler, monitor, and isolation mechanism. Furthermore, the target host is configured with multiple target virtual machines, each of which can be deployed with an agent tool.
[0155] Specifically, GPU physical resources are interconnected through X-Link and other means. As the main control module for GPU resource management, the GPU centralized control platform can be used to schedule, manage and allocate GPU resources. The scheduler is used to manage and allocate GPU resources to ensure efficient use of GPU resources. The monitor is used to receive GPU status information collected by the agent tool in the target virtual machine, and to issue an alarm in time when the status information meets the preset warning conditions. The isolation mechanism is used to ensure that the GPU resources between different target virtual machines are independent of each other to prevent resource conflicts and data leakage. Electronic devices can specifically integrate isolation methods of GPUs from different manufacturers to ensure that training or inference data is securely transmitted in the X-Link network inside the target virtual machine, while ensuring that the X-Links inside different target virtual machines are isolated from each other and do not interfere with each other.
[0156] In the embodiment of the present application, the electronic device, in the absence of high-speed interconnection hardware such as NVSwitch, realizes high-speed interconnection and dynamic scheduling between GPU cards in a virtualized environment based on X-Link, realizes dynamic allocation of GPU resources, calculates performance baseline data based on topology information, and allocates GPU partitions with higher communication efficiency to the target virtual machine based on the performance baseline data. Compared with the random allocation of GPU cards in related technologies, the GPU computing power scheduling method of the present application improves resource utilization, improves the communication efficiency between GPU cards in the target virtual machine, and improves the efficiency of user task processing. The isolation mechanism of GPUs from different manufacturers is integrated in the target host for different GPU resources, which improves data security while ensuring task performance. In addition, when the degree of GPU resource fragmentation is high, the electronic device in the present application reallocates GPU resources, improves the utilization of GPU resources, and ensures the processing performance of tasks such as user training and reasoning.
[0157] Figure 8 For a schematic diagram of a GPU computing power scheduling device provided by an exemplary embodiment of this application, see Figure 8 , the GPU computing power scheduling device 80 includes:
[0158] The acquisition module 81 is used to obtain the topology information of each GPU partition in the target host and determine the performance baseline data of the GPU partition based on the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition;
[0159] A first determining module 82 is configured to determine a first target GPU partition corresponding to the target virtual machine based on the demand data of the target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition;
[0160] The second determination module 83 is used to determine the second target GPU partition corresponding to the target virtual machine based on the demand data of the running target virtual machine and the performance baseline data of each GPU partition when the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions, so that the target virtual machine performs task processing based on the second target GPU partition.
[0161] In one possible implementation, the performance baseline data includes a communication factor and a discrete factor; the communication factor is used to characterize the communication performance of the GPU partition; the discrete factor is used to characterize the discreteness of the GPU partition; the acquisition module 81 is specifically used to:
[0162] For each GPU partition, determine the topology structure and interconnection mode of the GPU partition based on the topology information, and determine the communication factor of the GPU partition based on the topology structure and interconnection mode;
[0163] The physical location code of each GPU in the GPU partition is determined according to the topology information, and the discrete factor of the GPU partition is calculated according to the physical location code.
[0164] In a possible implementation, the acquisition module 81 is specifically configured to:
[0165] Determine the topological weight of the GPU partition based on the topological structure and interconnection mode;
[0166] Measure the communication bandwidth of the GPU partition and calculate the communication factor of the GPU partition based on the communication bandwidth and topology weight.
[0167] In one possible implementation, the device 80 is further configured to:
[0168] After calculating the communication factor of the first GPU partition, if the second GPU partition has the same topology and interconnection method as the first GPU partition, assign the communication factor of the second GPU partition to the communication factor of the first GPU partition;
[0169] If there is an isolated GPU in the GPU partition, the GPU partition is marked as a target state partition, and calculation of performance baseline data for the GPU partition is skipped.
[0170] In a possible implementation, the acquisition module 81 is specifically configured to:
[0171] According to the physical location code, the physical distribution data corresponding to the GPU partition is calculated, and the physical distribution data is used as the discrete factor of the GPU partition.
[0172] In one possible implementation, the device 80 is further configured to:
[0173] Obtaining the real-time discrete factor corresponding to each GPU partition, and calculating the global discrete factor corresponding to the GPU partition based on the real-time discrete factor; when the global discrete factor is greater than a preset discrete threshold, determining that the resource fragmentation information corresponding to each GPU partition in the target host meets the preset condition; the real-time discrete factor is the discrete factor corresponding to each GPU partition within a preset period; or,
[0174] Obtain task performance data of the target virtual machine; the task performance data includes at least one of actual GPU utilization, actual communication bandwidth, and actual task iteration speed; when the task performance data is less than a preset performance threshold, determine that resource fragmentation information corresponding to each GPU partition in the target host meets preset conditions.
[0175] In one possible implementation, the device 80 is further configured to:
[0176] Initialize the second target GPU partition according to the corresponding relationship between the target virtual machine and the second target GPU partition;
[0177] Establishing a binding relationship between each GPU in the second target GPU partition after the initialization process and the target virtual machine, and updating the GPU list of the target virtual machine;
[0178] When the second target GPU partition works normally after the initialization process, the binding relationship between the target virtual machine and the first target GPU partition is released.
[0179] In one possible implementation, the device 80 is further configured to:
[0180] In each target virtual machine, collect the status information of the GPU used by the target virtual machine;
[0181] Status information is received in the target host, and target warning information is output when the status information meets preset warning conditions; the preset warning conditions include that the GPU communication bandwidth is less than a preset bandwidth threshold, the GPU communication delay is greater than a preset delay threshold, or the GPU transmission error count is greater than a preset number threshold.
[0182] The GPU computing power scheduling device 80 provided in the embodiment of the present application can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar and will not be repeated here.
[0183] Figure 9 For a schematic diagram of the structure of a GPU computing power scheduling device provided by an exemplary embodiment of this application, see Figure 9 The GPU computing power scheduling device 90 may include a processor 91 and a memory 92. Exemplarily, the processor 91 and the memory 92 are interconnected via a bus 93.
[0184] Memory 92 stores computer-executable instructions;
[0185] The processor 91 executes the computer execution instructions stored in the memory 92, so that the processor 91 executes the GPU computing power scheduling method shown in the above method embodiment.
[0186] Accordingly, an embodiment of the present application provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, they are used to implement the GPU computing power scheduling method of the above-mentioned method embodiment.
[0187] Accordingly, an embodiment of the present application may also provide a computer program product, including a computer program, which, when executed by a processor, can implement the GPU computing power scheduling method shown in the above method embodiment.
[0188] The embodiment of the present application also provides a GPU computing power scheduling system, including a target host and a target virtual machine;
[0189] The target host is used to obtain the topology information of each GPU partition in the target host and determine the performance baseline data of the GPU partition based on the topology information; the performance baseline data is used to characterize the communication performance and discreteness of the GPU partition;
[0190] The target host is used to determine a first target GPU partition corresponding to the target virtual machine based on the demand data of the target virtual machine and the performance baseline data of each GPU partition, and the target virtual machine is used to perform task processing through the first target GPU partition;
[0191] The target host is used to determine the second target GPU partition corresponding to the target virtual machine based on the demand data of the running target virtual machine and the performance baseline data of each GPU partition when the resource fragmentation information corresponding to each GPU partition in the target host meets the preset conditions. The target virtual machine is used to perform task processing through the second target GPU partition.
[0192] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.) containing computer-usable program code.
[0193] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0194] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0195] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0196] In a typical configuration, a computing device includes one or more processors, input / output interfaces, network interfaces, and memory.
[0197] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0198] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, parameter random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0199] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0200] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A GPU computing power scheduling method, characterized in that: include: Obtaining topology information of each GPU partition in the target host, and determining performance baseline data of the GPU partition based on the topology information; The performance baseline data is used to characterize the communication performance and dispersion degree of the GPU partition; the performance baseline data includes a communication factor and a dispersion factor; Determine, based on the demand data of the target virtual machine and the performance baseline data of each GPU partition, a first target GPU partition corresponding to the target virtual machine, so that the target virtual machine performs task processing based on the first target GPU partition; Obtaining a real-time discrete factor corresponding to each GPU partition, and calculating a global discrete factor corresponding to the GPU partition based on the real-time discrete factor; When the global discrete factor is greater than a preset discrete threshold, it is determined that the resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition; the real-time discrete factor is the discrete factor corresponding to each GPU partition within a preset period; the real-time discrete factor is the real-time discrete factor of each GPU partition; The global discrete factor is the average real-time discrete factor corresponding to each GPU partition; or, Obtaining task performance data of the target virtual machine; the task performance data including at least one of actual GPU utilization, actual communication bandwidth, and actual task iteration speed; When the task performance data is less than a preset performance threshold, determining that resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition; When resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition, determining a second target GPU partition corresponding to the target virtual machine based on demand data of a running target virtual machine and performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition; Determining the performance baseline data of the GPU partition according to the topology information includes: For each GPU partition, determine a topology structure and an interconnection mode of the GPU partition according to the topology information, and determine a communication factor of the GPU partition according to the topology structure and the interconnection mode; The physical location code of each GPU in the GPU partition is determined according to the topology information, and the physical distribution data corresponding to the GPU partition is calculated according to the physical location code, and the physical distribution data is used as the discrete factor of the GPU partition; the physical distribution data refers to the physical distribution discretization degree data corresponding to each GPU card in the GPU partition.
2. The method according to claim 1, characterized in that Determining the communication factor of the GPU partition according to the topology structure and the interconnection mode includes: Determining a topological weight of the GPU partition according to the topological structure and the interconnection mode; The communication bandwidth of the GPU partition is measured, and a communication factor of the GPU partition is calculated according to the communication bandwidth and the topology weight.
3. The method according to claim 1, characterized in that The method further comprises: After calculating the communication factor of the first GPU partition, if the second GPU partition has the same topology and interconnection method as the first GPU partition, assign the communication factor of the second GPU partition to the communication factor of the first GPU partition; If an isolated GPU exists in the GPU partition, the GPU partition is marked as a target state partition, and calculation of performance baseline data of the GPU partition is skipped.
4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: Initializing the second target GPU partition according to the correspondence between the target virtual machine and the second target GPU partition; Establishing a binding relationship between each GPU in the initialized second target GPU partition and the target virtual machine, and updating the GPU list of the target virtual machine; If the second target GPU partition works normally after the initialization process, the binding relationship between the target virtual machine and the first target GPU partition is released.
5. The method according to any one of claims 1 to 3, characterized in that The method further comprises: In each of the target virtual machines, collecting status information of a GPU used by the target virtual machine; The status information is received in a target host, and target warning information is output when the status information meets a preset warning condition; the preset warning condition includes that the communication bandwidth of the GPU is less than a preset bandwidth threshold, the communication delay of the GPU is greater than a preset delay threshold, or the transmission error count of the GPU is greater than a preset number threshold.
6. A GPU computing power scheduling device, characterized in that: include: An acquisition module is used to obtain topology information of each GPU partition in the target host and determine performance baseline data of the GPU partition based on the topology information; The performance baseline data is used to characterize the communication performance and dispersion degree of the GPU partition; the performance baseline data includes a communication factor and a dispersion factor; A first determination module is configured to determine a first target GPU partition corresponding to the target virtual machine based on the demand data of the target virtual machine and the performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the first target GPU partition; The acquisition module is further configured to acquire a real-time discrete factor corresponding to each GPU partition, and calculate a global discrete factor corresponding to the GPU partition based on the real-time discrete factor; when the global discrete factor is greater than a preset discrete threshold, determine that the resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition; the real-time discrete factor is the discrete factor corresponding to each GPU partition within a preset period; the real-time discrete factor is the real-time discrete factor of each GPU partition; the global discrete factor is the average real-time discrete factor corresponding to each GPU partition; or, The acquisition module is further configured to acquire task performance data of the target virtual machine; the task performance data includes at least one of actual GPU utilization, actual communication bandwidth, and actual task iteration speed; When the task performance data is less than a preset performance threshold, determining that resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition; a second determining module configured to determine, when resource fragmentation information corresponding to each GPU partition in the target host satisfies a preset condition, a second target GPU partition corresponding to the target virtual machine based on demand data of a running target virtual machine and performance baseline data of each GPU partition, so that the target virtual machine performs task processing based on the second target GPU partition; The acquisition module is specifically configured to determine, for each GPU partition, a topology structure and an interconnection mode of the GPU partition according to the topology information, and determine a communication factor of the GPU partition according to the topology structure and the interconnection mode; The physical location code of each GPU in the GPU partition is determined according to the topology information, and the physical distribution data corresponding to the GPU partition is calculated according to the physical location code, and the physical distribution data is used as the discrete factor of the GPU partition; the physical distribution data refers to the physical distribution discretization degree data corresponding to each GPU card in the GPU partition.
7. A GPU computing power scheduling device, characterized in that: include: memory and processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the GPU computing power scheduling method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the GPU computing power scheduling method according to any one of claims 1 to 5.
9. A computer program product, characterized in that The method comprises a computer program, which, when executed by a computer, implements the GPU computing power scheduling method according to any one of claims 1 to 5.
10. A GPU computing power scheduling system, characterized in that: Including target host and target virtual machine; The target host is used to obtain topology information of each GPU partition in the target host, and determine performance baseline data of the GPU partition according to the topology information; The performance baseline data is used to characterize the communication performance and dispersion degree of the GPU partition; the performance baseline data includes a communication factor and a dispersion factor; The target host is used to determine a first target GPU partition corresponding to the target virtual machine according to the demand data of the target virtual machine and the performance baseline data of each GPU partition, and the target virtual machine is used to perform task processing through the first target GPU partition; The target host is used to obtain the real-time discrete factors corresponding to each GPU partition, and calculate the global discrete factors corresponding to the GPU partition according to the real-time discrete factors; When the global discrete factor is greater than a preset discrete threshold, it is determined that the resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition; the real-time discrete factor is the real-time discrete factor of each GPU partition; The global discrete factor is the average real-time discrete factor corresponding to each GPU partition; or The target host is used to obtain task performance data of the target virtual machine; the task performance data includes at least one of actual GPU utilization, actual communication bandwidth, and actual task iteration speed; When the task performance data is less than a preset performance threshold, determining that resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition; The target host is configured to determine, when resource fragmentation information corresponding to each GPU partition in the target host meets a preset condition, a second target GPU partition corresponding to the target virtual machine based on demand data of a running target virtual machine and performance baseline data of each GPU partition, and the target virtual machine is configured to perform task processing through the second target GPU partition; Determining the performance baseline data of the GPU partition according to the topology information includes: For each GPU partition, determine a topology structure and an interconnection mode of the GPU partition according to the topology information, and determine a communication factor of the GPU partition according to the topology structure and the interconnection mode; The physical location code of each GPU in the GPU partition is determined according to the topology information, and the physical distribution data corresponding to the GPU partition is calculated according to the physical location code, and the physical distribution data is used as the discrete factor of the GPU partition; the physical distribution data refers to the physical distribution discretization degree data corresponding to each GPU card in the GPU partition.
Citation Information
Patent Citations
Efficient GPU resource allocation optimization method and system
CN111930498A