AI Computing Power Scheduling Method, Device and Equipment for Cloud Platform

By extracting VGPU information in virtual machine migration instructions from the control node of the cloud platform, determining alternative computing nodes, and detecting computing nodes that meet the conditions, the problem of difficult to determine the optimal target computing node for thermal migration in the SR-IOV scenario is solved, and efficient utilization and balance of GPU computing power resources are achieved.

CN119917224BActive Publication Date: 2025-06-27WUHAN YANGTZE COMPUTING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510404977.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-27
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The prior art cannot determine the optimal target computing node for thermal migration in SR-IOV scenarios, resulting in waste and imbalance of GPU computing resources.

Method used

After receiving the virtual machine migration instructions in the control node of the cloud platform, the target model, target slot and target number are extracted from the VGPU information of the source virtual machine, the alternative computing node is determined, and whether there is a first computing node or a second computing node that meets the target conditions is detected, thereby determining the target computing node.

Benefits of technology

It realizes the optimal target computing node for thermal migration in the SR-IOV scenario, improves the utilization rate and balance of GPU computing power resources, and reduces business interruption and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917224B_ABST
    Figure CN119917224B_ABST
Patent Text Reader

Abstract

The present application provides a cloud platform AI computing power scheduling method, apparatus and device. The method includes: after receiving a virtual machine migration instruction, extracting a target model, a target slot and a target quantity from the VGPU information of the source virtual machine, where the target model is the model of the GPU that provides VGPUs to the source virtual machine, the target slot is the slot where the GPU that provides VGPUs to the source virtual machine is located, and the target quantity is the quantity of VGPUs provided by the GPUs in the target slot to the source virtual machine; determining alternative computing nodes from other computing nodes, where the GPU model in each target slot of the alternative computing nodes is the target model; detecting whether there is a first computing node among the alternative computing nodes, where the number of idle VGPUs of the GPUs in each target slot of the first computing node is greater than or equal to the corresponding target quantity; if there is a first computing node, determining a target computing node from the first computing nodes. Through the present application, the optimal target computing node for hot migration in the SR-IOV scenario can be determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of virtualization technology, and particularly to a cloud platform AI computing power scheduling method, device, and equipment. Background Art

[0002] With the wide application and popularization of artificial intelligence and large models, the demand for AI computing power resources is increasing day by day, showing an exponential growth trend. Currently, the main demand for servers is AI servers in intelligent computing centers. There are multiple GPU cards on AI servers to provide AI computing power, and the AI services in the data center are deployed on the GPU cards in the resource pool of these AI servers. Currently, the AI servers of the servers mainly use GPUs to execute AI computing tasks, and the CPUs are responsible for data input and the execution of AI service programs. In the intelligent computing center, GPUs are the main AI computing power and are scheduled and managed through the cloud platform. In order to improve the management efficiency and utilization rate of computing power resources, current servers need to use virtualization technology to schedule and manage computing power resources through the cloud platform.

[0003] In related technologies, the GPU card uses the SR-IOV virtualization technology to virtualize the PF (Physical Function) into multiple VF (Virtual Function) devices for virtual machines to use, thereby improving the utilization rate of GPU computing power. However, currently, the cold migration technology is mostly used in the SR-IOV scenario. Cold migration will interrupt for a long time, which has a great impact on scenarios that require continuous operation such as training and inference, and will also cause waste and imbalance of GPU computing power resources. Moreover, the computing power scheduling algorithms used in the hot migration technology in other scenarios cannot determine the optimal target computing node for hot migration in the SR-IOV scenario because they do not fully consider the characteristics of the SR-IOV scenario. Summary of the Invention

[0004] This application provides a cloud platform AI computing power scheduling method, device, and equipment, which can solve the technical problem in the prior art that the optimal target computing node for hot migration in the SR-IOV scenario cannot be determined.

[0005] In a first aspect, an embodiment of this application provides a cloud platform AI computing power scheduling method, which is applied to a control node of the cloud platform. The cloud platform AI computing power scheduling method includes:

[0006] After receiving a virtual machine migration instruction, extract the target model, target slot, and target quantity from the VGPU information of the source virtual machine, where the target model is the model of the GPU that provides VGPU to the source virtual machine, the target slot is the slot where the GPU that provides VGPU to the source virtual machine is located, the target quantity is the quantity of VGPUs provided by the GPU in the target slot to the source virtual machine, and the target slot corresponds to the target quantity one by one;

[0007] Determine alternative computing nodes from other computing nodes, where the GPU models in each target slot of the alternative computing nodes are all target models;

[0008] Detect whether there is a first computing node in the alternative computing nodes, where the number of idle VGPUs of the GPUs in each target slot of the first computing node is greater than or equal to the corresponding target number;

[0009] If there is a first computing node, determine the target computing node from the first computing nodes.

[0010] Further, in one embodiment, after detecting whether there is a first computing node in the alternative computing nodes, it further includes:

[0011] If there is no first computing node, detect whether there is a second computing node in the alternative computing nodes, where the sum of the number of idle VGPUs of all the GPUs of the second computing node with the target models is greater than or equal to the sum of all the target numbers;

[0012] If there is a second computing node, determine the target computing node from the second computing nodes, and send a first release instruction to the target computing node, where the first release instruction is used to control the target computing node to release the VGPUs required for a new virtual machine through the GPUs of the target model, so as to convert the target computing node into a first computing node.

[0013] Further, in one embodiment, after detecting whether there is a second computing node in the alternative computing nodes, it further includes:

[0014] If there is no second computing node, determine the target computing node from the alternative computing nodes, and send a second release instruction to the target computing node, where the second release instruction is used to control the target computing node to release the VGPUs required for a new virtual machine through the GPUs and CPUs of the target model, so as to convert the target computing node into a first computing node.

[0015] Further, in one embodiment, the determining the target computing node from the alternative computing nodes includes:

[0016] Determine the target computing node from the alternative computing nodes whose CPUs support HBM.

[0017] Further, in one embodiment, when there are multiple alternative computing nodes for the target computing node, determine the computing node with the lowest GPU occupancy rate as the target computing node.

[0018] Further, in one embodiment, when there are multiple computing nodes with the lowest GPU occupancy rate, determine the computing node with the lowest CPU occupancy rate as the target computing node.

[0019] Further, in one embodiment, the cloud platform AI computing power scheduling method further includes:

[0020] After receiving a virtual machine migration instruction, send a first synchronization instruction to the source computing node, where the first synchronization instruction is used to control the source computing node to synchronize the VGPU video memory of the source virtual machine to the memory of the source computing node;

[0021] After selecting a target computing node, send a second synchronization instruction to the target computing node, where the second synchronization instruction is used to control the target computing node to synchronize the VGPU video memory of the source virtual machine from the memory of the source computing node to the VGPU video memory of the new virtual machine.

[0022] Further, in one embodiment, the source computing node synchronizes the VGPU video memory of the source virtual machine to the memory of the source computing node through DMA technology;

[0023] The target computing node synchronizes the VGPU video memory of the source virtual machine from the memory of the source computing node to the VGPU video memory of the new virtual machine through GPU DRMA technology.

[0024] In a second aspect, an embodiment of the present application further provides a cloud platform AI computing power scheduling device, which is applied to a control node of a cloud platform. The cloud platform AI computing power scheduling device includes:

[0025] An information extraction module, configured to extract a target model, a target slot, and a target quantity from the VGPU information of the source virtual machine after receiving a virtual machine migration instruction, where the target model is the model of the GPU that provides the VGPU to the source virtual machine, the target slot is the slot where the GPU that provides the VGPU to the source virtual machine is located, the target quantity is the quantity of VGPUs provided by the GPUs in the target slot to the source virtual machine, and the target slot and the target quantity are in one-to-one correspondence;

[0026] An alternative determination module, configured to determine alternative computing nodes from other computing nodes, where the GPU models in each target slot of the alternative computing nodes are all the target models;

[0027] A first detection module, configured to detect whether there is a first computing node among the alternative computing nodes, where the number of idle VGPUs of the GPUs in each target slot of the first computing node is greater than or equal to the corresponding target quantity;

[0028] A target determination module, configured to, if there is a first computing node, determine a target computing node from the first computing nodes.

[0029] In a third aspect, an embodiment of the present application further provides a cloud platform AI computing power scheduling device, which includes a processor, a memory, and a cloud platform AI computing power scheduling program stored on the memory and executable by the processor. When the cloud platform AI computing power scheduling program is executed by the processor, the steps of the above cloud platform AI computing power scheduling method are implemented.

[0030] In the present application, after receiving a virtual machine migration instruction, the control node extracts the target model, target slot, and target quantity from the VGPU information of the source virtual machine. According to the extracted information, it first finds alternative computing nodes among other computing nodes whose GPU models and slots meet the requirements of the target model and target slot. When there is a first computing node among the alternative computing nodes whose available VGPU quantity in the target slot meets the target quantity requirement, the target computing node is determined from the first computing node. Through the present application, fully considering the characteristics of the SR-IOV scenario, the optimal target computing node for hot migration in the SR-IOV scenario can be determined. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a flowchart of the cloud platform AI computing power scheduling method in an embodiment of the present application;

[0032] Figure 2 is a schematic diagram of a computing node in an embodiment of the present application;

[0033] Figure 3 is a schematic diagram of VGPU video memory synchronization in an embodiment of the present application;

[0034] Figure 4 is a schematic diagram of the functional modules of the cloud platform AI computing power scheduling device in an embodiment of the present application;

[0035] Figure 5 is a schematic diagram of the hardware structure of the cloud platform AI computing power scheduling device involved in the solution of the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0037] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the drawings.

[0038] In a first aspect, an embodiment of the present application provides a cloud platform AI computing power scheduling method, which is applied to a control node of the cloud platform.

[0039] Figure 1 The flowchart of the cloud platform AI computing power scheduling method in an embodiment of the present application is shown.

[0040] Referring to Figure 1 , in one embodiment, the cloud platform AI computing power scheduling method includes the following steps:

[0041] S1. After receiving a virtual machine migration instruction, extract the target model, target slot, and target quantity from the VGPU information of the source virtual machine. Here, the target model is the model of the GPU that provides the VGPU to the source virtual machine, the target slot is the slot where the GPU that provides the VGPU to the source virtual machine is located, the target quantity is the quantity of VGPUs provided by the GPU in the target slot to the source virtual machine, and the target slot corresponds one-to-one with the target quantity.

[0042] Specifically, the cloud platform includes a control node and multiple computing nodes. Multiple virtual machines run in the computing nodes. The physical resources of the computing nodes include CPUs and multiple GPUs. The VF obtained by virtualizing the GPU as a PF through the SR-IOV virtualization technology is called a VGPU, and the virtual machine configures the VGPU as a peripheral. Different GPUs in the same computing node are set in different slots, and the slot lists owned by different computing nodes in the cloud platform are the same. The control node is the entry and management center for logging in to the cloud platform, responsible for operations such as the creation, migration, and monitoring of virtual machines. The control node can obtain the GPU information, CPU information, etc. of all computing nodes as needed, as well as the VGPU information, VCPU information, XML configuration files, etc. of the virtual machines in the computing nodes.

[0043] Exemplarily, the slot information can be the BDF (Bus, Device, Function) number.

[0044] The virtual machine migration instruction is manually triggered when the user needs to perform maintenance, such as when it is necessary to replace the hardware on the computing node (such as a server), maintain the computing node, add new computing node resources, etc. After the control node receives the virtual machine migration instruction, it needs to decide which computing node will be used as the target computing node. After the hot migration is completed, the target computing node will continue to execute the AI service of the source virtual machine through the new virtual machine, and the VGPU configuration of the new virtual machine needs to be consistent with that of the source virtual machine.

[0045] A virtual machine can be configured with one or more VGPUs. When multiple VGPUs are configured, these VGPUs can come from one or more GPUs. The number of VGPUs provided by different GPUs can be different, but the models of these GPUs need to be the same. Therefore, in this embodiment, the target model is unique, and the target slot and target quantity may not be unique. Each target slot corresponds to a target quantity.

[0046] Figure 2 FIG. shows a schematic diagram of a computing node in an embodiment of the present application.

[0047] Refer to Figure 2 , for example, the computing node runs virtual machine 1 and virtual machine 2. CPU1 is provided in slot 1, and CPU2 is provided in slot 2. Both GPU1 and GPU2 are of model A. Virtual machine 1 is configured with VGPU1 of GPU1, VGPU1 and VGPU2 of GPU2, and virtual machine 2 is configured with VGPU2 of GPU1. When virtual machine 1 is the source virtual machine, the target model is model A, the target slots are slot 1 and slot 2, the target quantity corresponding to slot 1 is 1, and the target quantity corresponding to slot 2 is 2. Correspondingly, the GPUs in slot 1 and slot 2 on the target computing node also need to be of model A, and the GPU in slot 1 provides 1 VGPU to the new virtual machine, and the GPU in slot 2 provides 2 VGPUs to the new virtual machine.

[0048] S2. Determine alternative computing nodes from other computing nodes, where the GPU models in each target slot of the alternative computing nodes are all the target model.

[0049] In this embodiment, the GPU models and slots of the alternative computing nodes meet the requirements of the target model and target slots, and there is no need to change the hardware configuration. When there are no alternative computing nodes among other computing nodes, the virtual machine migration instruction cannot be successfully responded to.

[0050] For example, when Figure 2 the virtual machine 1 in is the source virtual machine, the GPUs in slot 1 and slot 2 on the alternative computing node are of model A.

[0051] S3. Detect whether there is a first computing node among the alternative computing nodes, where the number of idle VGPUs of the GPUs in each target slot of the first computing node is greater than or equal to the corresponding target quantity.

[0052] In this embodiment, the number of idle VGPUs in the target slots of the first computing node meets the target quantity requirements, and there is no need to perform a resource release operation before hot migration, and the VGPUs required to create a new virtual machine can be provided.

[0053] For example, when Figure 2When virtual machine 1 in is the source virtual machine, the GPUs in slots 1 and 2 on the first computing node are model A, the number of idle VGPUs in slot 1 is greater than or equal to 1, and the number of idle VGPUs in slot 2 is greater than or equal to 2.

[0054] S4. If the first computing node exists, determine the target computing node from the first computing node.

[0055] In this embodiment, the candidate computing nodes are the complete optional range of the target computing nodes, the first computing node is the primary optional range of the target computing node, and the target computing node is determined preferentially from the first computing node, which can avoid unnecessary resource release operations before hot migration as much as possible.

[0056] It can be understood that when the first computing node is unique, it is determined as the target computing node. When the first computing node is not unique, the target computing node can be determined by referring to other factors, such as GPU occupancy, CPU occupancy, memory occupancy, disk occupancy, etc.

[0057] Therefore, in this embodiment, after receiving the virtual machine migration instruction, the control node extracts the target model, target slot and target quantity from the VGPU information of the source virtual machine, and first finds the candidate computing node from other computing nodes based on the extracted information, and when there is a first computing node among the candidate computing nodes, determines the target computing node from the first computing node. Through this embodiment, the characteristics of the SR-IOV scenario are fully considered, and the optimal target computing node for hot migration in the SR-IOV scenario can be determined.

[0058] It should be noted that the native virtualization layer software on the computing nodes of the current cloud platform does not support hot migration in SR-IOV scenarios. The virtualization layer software needs to be optimized to support hot migration of virtual machines configured with PCIE pass-through devices.

[0059] For example, the operation of modifying the libvirt virtualization software on the computing node includes:

[0060] 1. Modify the libvirt source code to remove the check of the PCIE passthrough device configured for the source virtual machine when executing the hot migration instruction on the source computing node. Remove the restriction in the original code that stops executing the hot migration instruction if the source virtual machine is detected to have a PCIE passthrough device configuration, and continue to send the virtual machine hot migration instruction to the libvirt of the target computing node.

[0061] 2. Modify the libvirt source code. On the target computing node, first obtain the virtual machine configuration of the source computing node. If the virtual machine configuration contains a PCIe passthrough device, call the OS interface to obtain the BDF numbers of the PCIe devices on the current computing node. Compare the BDF numbers of the PCIe passthrough devices in the source computing node configuration with the obtained BDF numbers of the target computing node. If the target computing node contains all the BDF numbers of the PCIe passthrough devices in the virtual machine configuration files, it proves that the hardware configuration of the target computing node supports PCIe device passthrough. Then send a live migration instruction to qemu to complete the live migration of the virtual machine. Otherwise, stop sending the live migration instruction to qemu, return that the computing node hardware does not support PCIe device passthrough live migration, and the program exits.

[0062] Further, in one embodiment, after detecting whether there is a first computing node in the alternative computing nodes, it further includes:

[0063] If there is no first computing node, detect whether there is a second computing node in the alternative computing nodes, where the sum of the idle VGPUs of all the target models of GPUs in the second computing node is greater than or equal to the sum of all the target quantities;

[0064] If there is a second computing node, determine the target computing node from the second computing node, and send a first release instruction to the target computing node, where the first release instruction is used to control the target computing node to release the VGPUs required for the new virtual machine through the GPUs of the target model, so as to convert the target computing node into a first computing node.

[0065] In this embodiment, considering that when performing computing power scheduling inside the computing node, the VGPUs of the same model of GPUs can play an equivalent role and are not restricted by the slots. Since the sum of the idle VGPUs of all the target models of GPUs in the second computing node is greater than or equal to the sum of all the target quantities, the second computing node can, under the control of the first release instruction, only rely on the replacement of VGPUs between the GPUs of the target model to migrate a part of the AI services of the VGPUs to another part of the VGPUs, thereby releasing the VGPUs required for the new virtual machine and converting itself into a first computing node.

[0066] After the resource release operation, the corresponding AI services are still executed by the VGPUs of the same model of GPUs, and the execution effect remains basically unchanged. Therefore, in this embodiment, the second computing node is used as the secondary optional range of the target computing node. When there is no first computing node and there is a second computing node, determining the target computing node from the second computing node can minimize the negative impact brought by the resource release operation.

[0067] Exemplarily, when Figure 2When the virtual machine 1 in it is the source virtual machine, there is a computing node. The GPUs in slot 1, slot 2, and slot 3 are of model A. The GPU in slot 1 has no idle VGPUs, the number of idle VGPUs in slot 2 is equal to 1, and the number of idle VGPUs in slot 3 is equal to 2. The total number of idle VGPUs of all target model GPUs is equal to 3, and the total number of target quantities is equal to 3. Since the two are equal, this computing node is the second computing node. When this computing node is the target computing node, 2 idle VGPUs in slot 3 are used to replace 1 occupied VGPU in slot 1 and slot 2 respectively.

[0068] Specifically, the resource release operation corresponding to the first release instruction includes: determining the VGPUs to be released from the target slots where the number of idle VGPUs is less than the target quantity, and the quantity is equal to the difference between the target quantity and the number of idle VGPUs. Check the virtual machines corresponding to the VGPUs to be released, and modify the VGPUs to be released to the target VGPUs that meet the requirements in the XML configuration file of the virtual machine, and restart the virtual machine so that the VGPUs to be released are updated from the occupied state to the idle state. The requirements that the target VGPUs need to meet include: the target VGPUs are in the idle state, the GPU models they belong to are the target models, and the GPUs they belong to do not include the VGPUs to be released.

[0069] Optionally, the VGPUs with a lower occupancy rate are used as the VGPUs to be released, and the target VGPUs are selected from the GPUs with a lower occupancy rate, so as to further reduce the negative impact brought by the resource release operation.

[0070] Furthermore, in one embodiment, after detecting whether there is a second computing node in the alternative computing nodes, it further includes:

[0071] If there is no second computing node, determine the target computing node from the alternative computing nodes, and send a second release instruction to the target computing node, where the second release instruction is used to control the target computing node to release the VGPUs required for the new virtual machine through the GPUs and CPUs of the target model, so as to convert the target computing node into the first computing node.

[0072] It can be understood that since there is no second computing node in the alternative computing nodes, only by replacing the VGPUs between the GPUs of the target model, it is impossible to release enough VGPUs for the new virtual machine to use.

[0073] In this embodiment, considering that the CPU can also execute part of the AI services, when there is no second computing node, the target computing node is controlled by the second release instruction. Not only is part of the AI services of the VGPUs migrated to another part of the VGPUs by means of VGPU replacement between GPUs of the target model, but also part of the AI services of the VGPUs are migrated to the CPU, thereby releasing the VGPUs required for the new virtual machine and converting the target computing node into a first computing node. Through this embodiment, the computing power resources of the computing node are fully utilized, thereby improving the response success rate of the virtual machine migration instruction.

[0074] Exemplarily, when Figure 2 the virtual machine 1 in is the source virtual machine, there is a computing node. The GPUs in slot 1, slot 2, and slot 3 are of model A. The GPU in slot 1 has no free VGPUs. The number of free VGPUs in slot 2 is equal to 1. The number of free VGPUs in slot 3 is equal to 1. The sum of the free VGPUs of all the GPUs of the target model is equal to 2. The sum of the target numbers is equal to 3. The former is less than the latter. This computing node is not a second computing node. When this computing node is the target computing node, 1 free VGPU in slot 3 is used to replace 1 occupied VGPU in slot 1 or slot 2, and the AI services of the other occupied VGPU are migrated to the CPU.

[0075] Specifically, compared with the resource release operation corresponding to the first release instruction, the resource release operation corresponding to the second release instruction preferentially modifies the VGPU to be released into a target VGPU that meets the requirements in the XML configuration file of the virtual machine. When the target VGPU cannot be found, the VGPU to be released can also be directly deleted.

[0076] Further, in one embodiment, determining the target computing node from the alternative computing nodes includes:

[0077] Determining the target computing node from the alternative computing nodes whose CPUs support HBM.

[0078] In this embodiment, considering that a CPU supporting HBM (High Bandwidth Memory) can execute AI services more efficiently than an ordinary CPU, when the CPU of the target computing node supports HBM, the negative impact brought by the resource release operation can be further reduced.

[0079] Further, in one embodiment, when there are multiple alternative computing nodes for the target computing node, the computing node with the lowest GPU occupancy rate is determined as the target computing node.

[0080] Through this embodiment, it helps to achieve the GPU load balancing of different computing nodes. In addition, when there is no first computing node, the negative impact brought by the resource release operation can be further reduced.

[0081] Further, in one embodiment, when there are multiple computing nodes with the lowest GPU occupancy rate, the computing node with the lowest CPU occupancy rate is determined as the target computing node.

[0082] Through this embodiment, it helps to achieve the overall load balancing of different computing nodes. In addition, when there is no second computing node, the negative impact brought by resource release operations can be further reduced.

[0083] Further, in one embodiment, the AI computing power scheduling method of the cloud platform further includes:

[0084] After receiving the virtual machine migration instruction, send a first synchronization instruction to the source computing node, where the first synchronization instruction is used to control the source computing node to synchronize the VGPU video memory of the source virtual machine to the memory of the source computing node;

[0085] After selecting the target computing node, send a second synchronization instruction to the target computing node, where the second synchronization instruction is used to control the target computing node to synchronize the VGPU video memory of the source virtual machine from the memory of the source computing node to the VGPU video memory of the new virtual machine.

[0086] In the related art, the virtual machine live migration solution only calculates the dirty pages of the memory of the source virtual machine and performs dirty page migration to ensure the consistency of the programs and data in the memory of the source virtual machine and the new virtual machine. However, it does not process and monitor the dirty pages of the VGPU video memory of the source virtual machine, which will cause the VGPU video memory of the source virtual machine not to be migrated, resulting in the VGPU of the new virtual machine needing to reload and execute the original services and data, wasting GPU computing power resources and reducing the execution efficiency of AI services. For example, for AI services that require multiple rounds of training, the training framework and the completed training data can be retained, but the training data in the middle of training will be lost.

[0087] Through the embodiment, the VGPU video memory of the source virtual machine can be synchronized to the VGPU video memory of the new virtual machine, thereby reducing the waste of GPU computing power resources and improving the execution efficiency of AI services.

[0088] Figure 3 Shows a schematic diagram of VGPU video memory synchronization in an embodiment of the present application.

[0089] Refer to Figure 3 , Further, in one embodiment, the source computing node synchronizes the VGPU video memory of the source virtual machine to the memory of the source computing node through the DMA technology;

[0090] The target computing node synchronizes the VGPU video memory of the source virtual machine from the memory of the source computing node to the VGPU video memory of the new virtual machine through the GPU DRMA technology.

[0091] In this embodiment, the DMA (Direct Memory Access) technology can ensure the integrity and security of the VGPU video memory data synchronization operation in the source computing node, and the GPU DRMA (Remote Direct Memory Access) technology can improve the VGPU video memory synchronization performance from the source computing node to the target computing node.

[0092] In a second aspect, an embodiment of the present application further provides a cloud platform AI computing power scheduling device, which is applied to the control node of the cloud platform.

[0093] Figure 4 The figure shows a schematic diagram of the functional modules of the cloud platform AI computing power scheduling device in an embodiment of the present application.

[0094] Referring to Figure 4 , in an embodiment, the cloud platform AI computing power scheduling device includes:

[0095] An information extraction module 10, configured to extract a target model, a target slot, and a target quantity from the VGPU information of the source virtual machine after receiving a virtual machine migration instruction, where the target model is the model of the GPU that provides the VGPU for the source virtual machine, the target slot is the slot where the GPU that provides the VGPU for the source virtual machine is located, the target quantity is the quantity of VGPUs provided by the GPUs in the target slot for the source virtual machine, and the target slot corresponds to the target quantity one by one;

[0096] An alternative determination module 20, configured to determine alternative computing nodes from other computing nodes, where the GPU models in each target slot of the alternative computing nodes are all the target models;

[0097] A first detection module 30, configured to detect whether there is a first computing node in the alternative computing nodes, where the quantity of idle VGPUs of the GPUs in each target slot of the first computing node is greater than or equal to the corresponding target quantity;

[0098] A target determination module 40, configured to determine a target computing node from the first computing nodes if there is a first computing node.

[0099] Further, in an embodiment, the cloud platform AI computing power scheduling device further includes a second detection module, configured to detect whether there is a second computing node in the alternative computing nodes if there is no first computing node, where the sum of the quantities of idle VGPUs of all the GPUs of the second computing node with the target models is greater than or equal to the sum of all the target quantities;

[0100] The target determination module 40 is further configured to, if there is a second computing node, determine a target computing node from the second computing nodes, and send a first release instruction to the target computing node, where the first release instruction is used to control the target computing node to release the VGPUs required for a new virtual machine through the GPUs of the target model, so as to convert the target computing node into a first computing node.

[0101] Further, in an embodiment, the target determination module 40 is further configured to, if there is no second computing node, determine a target computing node from the alternative computing nodes, and send a second release instruction to the target computing node, where the second release instruction is used to control the target computing node to release the VGPUs required for a new virtual machine through the GPUs and CPUs of the target model, so as to convert the target computing node into a first computing node.

[0102] Further, in an embodiment, the target determination module 40 is configured to determine a target computing node from the alternative computing nodes whose CPUs support HBM.

[0103] Further, in an embodiment, the target determination module 40 is configured to, when there are multiple alternative computing nodes for the target computing node, determine the computing node with the lowest GPU occupancy rate as the target computing node.

[0104] Further, in an embodiment, the target determination module 40 is configured to, when there are multiple computing nodes with the lowest GPU occupancy rate, determine the computing node with the lowest CPU occupancy rate as the target computing node.

[0105] Further, in an embodiment, the cloud platform AI computing power scheduling device further includes a video memory synchronization module, which is configured to:

[0106] After receiving a virtual machine migration instruction, send a first synchronization instruction to the source computing node, where the first synchronization instruction is used to control the source computing node to synchronize the VGPU video memory of the source virtual machine to the memory of the source computing node;

[0107] After selecting the target computing node, send a second synchronization instruction to the target computing node, where the second synchronization instruction is used to control the target computing node to synchronize the VGPU video memory of the source virtual machine from the memory of the source computing node to the VGPU video memory of the new virtual machine.

[0108] Further, in an embodiment, the source computing node synchronizes the VGPU video memory of the source virtual machine to the memory of the source computing node through the DMA technology;

[0109] The target computing node synchronizes the VGPU video memory of the source virtual machine from the memory of the source computing node to the VGPU video memory of the new virtual machine through the GPU DRMA technology.

[0110] Among them, the function implementation of each module in the above cloud platform AI computing power scheduling device corresponds to each step in the above cloud platform AI computing power scheduling method embodiment, and its function and implementation process will not be elaborated here one by one.

[0111] In a third aspect, an embodiment of the present application provides a cloud platform AI computing power scheduling device, and the cloud platform AI computing power scheduling device may be a device with data processing functions such as a personal computer (PC), a laptop computer, a server, etc.

[0112] Figure 5 The hardware structure diagram of the cloud platform AI computing power scheduling device involved in the embodiment of the present application is shown.

[0113] Referring to Figure 5 In the embodiment of the present application, the cloud platform AI computing power scheduling device may include a processor, a memory, a communication interface, and a communication bus.

[0114] Among them, the communication bus can be of any type and is used to interconnect the processor, the memory, and the communication interface.

[0115] The communication interface includes input / output (I / O) interfaces, physical interfaces, and logical interfaces, etc., which are used to implement the interconnection of internal components of the cloud platform AI computing power scheduling device, and interfaces for implementing the interconnection of the cloud platform AI computing power scheduling device with other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.; the user device can be a display screen (Display), a keyboard (Keyboard), etc.

[0116] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical memory, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0117] The processor may be a general-purpose processor, which can call the cloud platform AI computing power scheduling program stored in the memory and execute the cloud platform AI computing power scheduling method provided in the embodiments of the present application. For example, the general-purpose processor may be a central processing unit (CPU). Among them, the method executed when the cloud platform AI computing power scheduling program is called may refer to the various embodiments of the cloud platform AI computing power scheduling method of the present application, which will not be elaborated here.

[0118] Those skilled in the art can understand that Figure 5 the hardware structure shown in does not constitute a limitation on the present application, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0119] It should be noted that the serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0120] The terms "including" and "having" and any variations thereof in the specification, claims and drawings of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include unlisted steps or units, or may optionally further include other steps or units inherent to these processes, methods, products or devices. The descriptions of "first", "second", "third", etc. are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit that "first", "second", and "third" are different types.

[0121] In the description of the embodiments of the present application, "exemplary", "for example" or "for instance" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary", "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example" or "for instance" is intended to present related concepts in a specific manner.

[0122] In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0123] In some processes described in the embodiments of the present application, there are multiple operations or steps that appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of the present application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.

[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes several instructions for causing a terminal device to execute the methods described in the various embodiments of the present application.

[0125] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A cloud platform AI computing power scheduling method, characterized in that: Applied to the control node of the cloud platform, the cloud platform AI computing power scheduling method includes: After receiving the virtual machine migration instruction, extract the target model, target slot and target quantity from the VGPU information of the source virtual machine, where the target model is the model of the GPU that provides the VGPU to the source virtual machine, the target slot is the slot where the GPU that provides the VGPU to the source virtual machine is located, and the target quantity is the number of VGPUs provided by the GPU in the target slot to the source virtual machine, and the target slot corresponds to the target quantity one by one; Determine a candidate computing node from other computing nodes, wherein the GPU model of each target slot of the candidate computing node is the target model; Detect whether there is a first computing node among the candidate computing nodes, wherein the number of idle VGPUs of the GPU of each target slot of the first computing node is greater than or equal to the corresponding target number; If the first computing node exists, the target computing node is determined from the first computing node.

2. The cloud platform AI computing power scheduling method according to claim 1, characterized in that: After detecting whether the first computing node exists in the candidate computing nodes, the method further includes: If the first computing node does not exist, then detecting whether there is a second computing node among the candidate computing nodes, wherein the sum of the number of idle VGPUs of all target models of GPUs of the second computing node is greater than or equal to the sum of all target numbers; If there is a second computing node, determine the target computing node from the second computing node, and send a first release instruction to the target computing node, wherein the first release instruction is used to control the target computing node to release the VGPU required for the new virtual machine through the GPU of the target model to convert the target computing node into the first computing node.

3. The cloud platform AI computing power scheduling method according to claim 2, characterized in that: After detecting whether there is a second computing node among the candidate computing nodes, the method further includes: If the second computing node does not exist, the target computing node is determined from the candidate computing nodes, and a second release instruction is sent to the target computing node, wherein the second release instruction is used to control the target computing node to release the VGPU required for the new virtual machine through the GPU and CPU of the target model to convert the target computing node into the first computing node.

4. The cloud platform AI computing power scheduling method according to claim 3, characterized in that: The step of determining the target computing node from the candidate computing nodes includes: A target computing node is determined from candidate computing nodes whose CPUs support HBM.

5. The cloud platform AI computing power scheduling method according to any one of claims 1 to 4, characterized in that: When there are multiple optional computing nodes for the target computing node, the computing node with the lowest GPU occupancy rate is determined as the target computing node.

6. The cloud platform AI computing power scheduling method according to claim 5, characterized in that: When there are multiple computing nodes with the lowest GPU occupancy, the computing node with the lowest CPU occupancy is determined as the target computing node.

7. The cloud platform AI computing power scheduling method according to any one of claims 1 to 4, characterized in that: The cloud platform AI computing power scheduling method also includes: After receiving the virtual machine migration instruction, sending a first synchronization instruction to the source computing node, wherein the first synchronization instruction is used to control the source computing node to synchronize the VGPU video memory of the source virtual machine to the memory of the source computing node; After selecting the target computing node, a second synchronization instruction is sent to the target computing node, wherein the second synchronization instruction is used to control the target computing node to synchronize the VGPU video memory of the source virtual machine to the VGPU video memory of the new virtual machine from the memory of the source computing node.

8. The cloud platform AI computing power scheduling method according to claim 7, characterized in that: The source computing node synchronizes the VGPU memory of the source virtual machine to the memory of the source computing node through DMA technology; The target compute node uses GPU DRMA technology to synchronize the VGPU video memory of the source virtual machine from the memory of the source compute node to the VGPU video memory of the new virtual machine.

9. A cloud platform AI computing power scheduling device, characterized in that: Applied to the control node of the cloud platform, the cloud platform AI computing power scheduling device includes: An information extraction module, configured to extract a target model, a target slot, and a target quantity from the VGPU information of the source virtual machine after receiving a virtual machine migration instruction, wherein the target model is the model of a GPU that provides a VGPU to the source virtual machine, the target slot is the slot where the GPU that provides a VGPU to the source virtual machine is located, and the target quantity is the number of VGPUs provided by the GPU in the target slot to the source virtual machine, and the target slot corresponds to the target quantity one by one; An alternative determination module, used to determine an alternative computing node from other computing nodes, wherein the GPU model of the alternative computing node in each target slot is the target model; A first detection module is used to detect whether there is a first computing node among the candidate computing nodes, wherein the number of idle VGPUs of the GPU of each target slot of the first computing node is greater than or equal to the corresponding target number; The target determination module is used to determine the target computing node from the first computing node if the first computing node exists.

10. A cloud platform AI computing power scheduling device, characterized in that: The cloud platform AI computing power scheduling device includes a processor, a memory, and a cloud platform AI computing power scheduling program stored in the memory and executable by the processor. When the cloud platform AI computing power scheduling program is executed by the processor, the steps of the cloud platform AI computing power scheduling method as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • VGPU management method and device, electronic equipment and storage medium

    CN112463392A

  • Virtual machine scheduling method, device and system, storage medium and program product

    CN119484534A