Graphics processing unit (GPU) scheduling method and apparatus, and storage medium

In the multi-core GPU scenario, the migration command is used to establish the association relationship between the virtual machine and the target GPU core, and migrate the workload to the idle GPU core, the problem of load imbalance between cores is solved, and load balancing and performance improvement is achieved.

WO2025113561A1PCT designated stage expired Publication Date: 2025-06-05MOORE THREADS TECH CO LTD

Patent Information

Application Number
PCT/CN2024/135253
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-11-28
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

In the multi-core GPU scenario, the workload of the virtual machine can only be processed on the initially selected GPU core, resulting in unbalanced load between cores, wasted hardware resources and increased pressure on the high-load GPU core.

Method used

By obtaining migration commands, establish an association between the source virtual machine VM and the corresponding hardware identifier on the target GPU core, and migrate the VM's workload to the target GPU core for processing.

Benefits of technology

Load balancing between multi-core GPU cores is realized, which reduces the pressure of high-load GPU cores, shortens response time, and improves throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135253_05062025_PF_FP_ABST
    Figure CN2024135253_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of Graphics Processing Units (GPUs), and in particular to a GPU scheduling method and apparatus, and a storage medium. The method comprises: acquiring a migration command, the migration command being used for migrating a workload of a source Virtual Machine (VM) from a source GPU core to a target GPU core; on the basis of the migration command, establishing an association relationship between the source VM and a corresponding hardware identifier on the target GPU core; and on the basis of the association relationship after the migration, using the target GPU core to process the workload from the source VM. According to embodiments of the present application, while the VM is not shut down, the target GPU core is used to process the workload from the source VM, and thus the workload of the source VM is migrated from a high-load source GPU core to an idle target GPU core, so that the thermal migration between multi-core GPU cores can be achieved, the load between a plurality of GPU cores can be balanced, the pressure of a high-load GPU core can be reduced, the response time can be shortened, and the throughput rate can be increased.
Need to check novelty before this filing date? Find Prior Art

Description

Graphics Processing Unit (GPU) scheduling method, device, and storage medium

[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on November 30, 2023, with application number 202311628327.9 and invention name “Graphics Processing Unit GPU Scheduling Method, Device and Storage Medium”, the entire contents of which are incorporated by reference into this disclosure. Technical Field

[0002] The present disclosure relates to the field of graphics processor technology, and in particular to a graphics processor (GPU) scheduling method, device, and storage medium. Background Art

[0003] Graphics processing units (GPUs) play a crucial role in graphics rendering, parallel computing, artificial intelligence, and other fields. In GPU virtualization, to support multiple virtual machines (VMs) using a single GPU simultaneously, the GPU's hardware resources are partitioned into multiple parts, providing independent hardware resources for each VM.

[0004] In the case of multi-core GPUs, current technical solutions typically process VM workloads only on the GPU core initially selected at startup. Since VMs can be powered on and off at any time, there's a chance that the workloads of multiple VMs will be concentrated on a single GPU core, leaving other GPU cores idle. This not only wastes hardware resources but also increases pressure on the loaded GPU core, leading to load imbalance between cores. Therefore, a new GPU scheduling method is urgently needed to balance load across multiple GPU cores, shorten response times, and improve throughput. Summary of the Invention

[0005] In view of this, the present disclosure proposes a graphics processing unit (GPU) scheduling method, device, and storage medium.

[0006] According to one aspect of the present disclosure, a method for scheduling a graphics processing unit (GPU) is provided. The method comprises:

[0007] Get the migration command, which is used to migrate the workload of the source virtual machine VM from the source GPU core to the target GPU core;

[0008] According to the migration command, an association relationship is established between the source VM and the corresponding hardware identifier on the target GPU core;

[0009] Based on the migrated associations, the target GPU core is used to process the workload from the source VM.

[0010] In one possible implementation, establishing an association between the source VM and the corresponding hardware identifiers on the target GPU core according to the migration command includes:

[0011] According to the migration command, a mapping is established between the register groups corresponding to the hardware identifiers on the source VM and the target GPU core;

[0012] According to the migration command, a mapping is established between the corresponding hardware identifier on the target GPU core and the command queue of the source VM, and a mapping is established between the corresponding hardware identifier on the target GPU core and the general video memory of the source VM.

[0013] In one possible implementation, a mapping is established between register groups corresponding to hardware identifiers on the source VM and the target GPU core according to the migration command, including:

[0014] According to the migration command, obtain the first address of the register group corresponding to the hardware identifier on the source GPU core and the first address of the register group corresponding to the hardware identifier on the target GPU core, where the first address is used to indicate the physical video memory address of the host;

[0015] Based on the first address of the register group corresponding to the hardware identifier on the source GPU core, the secondary page table of the source VM is updated so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core, so as to establish a mapping between the register groups corresponding to the hardware identifier on the source VM and the target GPU core, where the second address is used to indicate the virtual video memory address of the VM.

[0016] In one possible implementation, updating the secondary page table of the source VM based on the first address of the register group corresponding to the hardware identifier on the source GPU core, so that the updated secondary page table indicates a mapping between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core, includes:

[0017] Based on the first address of the register group corresponding to the hardware identifier on the source GPU core, the corresponding segment in the secondary page table of the source VM is changed so that the access to the register group corresponding to the hardware identifier on the source GPU core is trapped;

[0018] After the trap, the secondary page table of the source VM is updated so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core.

[0019] In one possible implementation, updating the secondary page table of the source VM after the trap so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core includes:

[0020] After the trap, a predetermined error handling function in the host driver is called to perform error handling registered by the GPU core. Based on the first address of the register group corresponding to the hardware identifier on the target GPU core, the hypervisor mapping interface is called to update the secondary page table, so that the updated secondary page table indicates the mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core.

[0021] In one possible implementation, establishing a mapping between the hardware identifier corresponding to the target GPU core and the command queue of the source VM, and a mapping between the hardware identifier corresponding to the target GPU core and the general video memory of the source VM according to the migration command includes:

[0022] Obtaining, according to the migration command, a first page table of a target command queue, the first page table indicating a mapping relationship between a third address of the command queue and a first address of the command queue, and between a third address of the general video memory and a first address of the general video memory, the target command queue being a command queue indicated by a corresponding hardware identifier on a target GPU core, the first address being used to indicate a physical video memory address of the host, and the third address being used to indicate a virtual video memory address of the host;

[0023] When the size of the third address used by the command queues between VMs is the same, the first page table of the target command queue is replaced with the first page table of the source VM command queue to establish a mapping between the corresponding hardware identifier on the target GPU core and the command queue of the source VM, and a mapping between the corresponding hardware identifier on the target GPU core and the general video memory of the source VM.

[0024] In a possible implementation, when the sizes of the third addresses used by the command queues between the VMs are different, establishing, according to the migration command, a mapping between the hardware identifier corresponding to the target GPU core and the command queue of the source VM, and a mapping between the hardware identifier corresponding to the target GPU core and the general graphics memory of the source VM, further includes:

[0025] Obtaining, according to the migration command, a second page table of the target command queue, where the second page table indicates a mapping relationship between the second address of the command queue and the third address of the command queue, and between the second address of the general graphics memory and the third address of the general graphics memory;

[0026] The second page table of the target command queue is replaced with the second page table of the source VM command queue.

[0027] In a possible implementation, the method further includes:

[0028] When the host driver is initialized, a first page table and a second page table are created for the command queue and general video memory corresponding to the hardware identifier of the GPU core.

[0029] In one possible implementation, based on the migrated association, the target GPU core is used to process the workload from the source VM, including:

[0030] In response to a write operation to a register group corresponding to a hardware identifier on a target GPU core, based on the association relationship after migration, using a microcontroller MCU of the target GPU core to obtain workload information from the source VM, the workload information including a second address corresponding to the workload and a third address of a page table root directory associated with the workload;

[0031] Using the MCU of the target GPU core, the second address corresponding to the workload and the third address of the page table root directory associated with the current workload are configured to the engine of the target GPU core, so that the engine of the target GPU core can address the host's video memory and process the current workload.

[0032] In one possible implementation, establishing an association between the source VM and the corresponding hardware identifiers on the target GPU core according to the migration command includes:

[0033] In the case where there are idle resources on the target GPU core, an association relationship is established between the source VM and the corresponding hardware identifiers on the target GPU core according to the migration command.

[0034] In a possible implementation, the migration command includes an identifier of a target GPU core and a hardware identifier corresponding to the target GPU core.

[0035] According to another aspect of the present disclosure, a graphics processing unit (GPU) scheduling device is provided. The device includes:

[0036] An acquisition module is used to obtain a migration command, which is used to migrate the workload of the source virtual machine VM from the source GPU core to the target GPU core;

[0037] A first establishing module is used to establish an association relationship between the source VM and the corresponding hardware identifier on the target GPU core according to the migration command;

[0038] The processing module is configured to process the workload from the source VM using the target GPU core based on the migrated association relationship.

[0039] In a possible implementation, the first establishing module is configured to:

[0040] According to the migration command, a mapping is established between the register groups corresponding to the hardware identifiers on the source VM and the target GPU core;

[0041] According to the migration command, a mapping is established between the corresponding hardware identifier on the target GPU core and the command queue of the source VM, and a mapping is established between the corresponding hardware identifier on the target GPU core and the general video memory of the source VM.

[0042] In one possible implementation, a mapping is established between register groups corresponding to hardware identifiers on the source VM and the target GPU core according to the migration command, including:

[0043] According to the migration command, obtain the first address of the register group corresponding to the hardware identifier on the source GPU core and the first address of the register group corresponding to the hardware identifier on the target GPU core, where the first address is used to indicate the physical video memory address of the host;

[0044] Based on the first address of the register group corresponding to the hardware identifier on the source GPU core, the secondary page table of the source VM is updated so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core, so as to establish a mapping between the register groups corresponding to the hardware identifier on the source VM and the target GPU core, where the second address is used to indicate the virtual video memory address of the VM.

[0045] In one possible implementation, updating the secondary page table of the source VM based on the first address of the register group corresponding to the hardware identifier on the source GPU core, so that the updated secondary page table indicates a mapping between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core, includes:

[0046] Based on the first address of the register group corresponding to the hardware identifier on the source GPU core, the corresponding segment in the secondary page table of the source VM is changed so that the access to the register group corresponding to the hardware identifier on the source GPU core is trapped;

[0047] After the trap, the secondary page table of the source VM is updated so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core.

[0048] In one possible implementation, updating the secondary page table of the source VM after the trap so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core includes:

[0049] After the trap, a predetermined error handling function in the host driver is called to perform error handling registered by the GPU core. Based on the first address of the register group corresponding to the hardware identifier on the target GPU core, the hypervisor mapping interface is called to update the secondary page table, so that the updated secondary page table indicates the mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core.

[0050] In one possible implementation, establishing a mapping between the hardware identifier corresponding to the target GPU core and the command queue of the source VM, and a mapping between the hardware identifier corresponding to the target GPU core and the general video memory of the source VM according to the migration command includes:

[0051] Obtaining, according to the migration command, a first page table of a target command queue, the first page table indicating a mapping relationship between a third address of the command queue and a first address of the command queue, and between a third address of the general video memory and a first address of the general video memory, the target command queue being a command queue indicated by a corresponding hardware identifier on a target GPU core, the first address being used to indicate a physical video memory address of the host, and the third address being used to indicate a virtual video memory address of the host;

[0052] When the size of the third address used by the command queues between VMs is the same, the first page table of the target command queue is replaced with the first page table of the source VM command queue to establish a mapping between the corresponding hardware identifier on the target GPU core and the command queue of the source VM, and a mapping between the corresponding hardware identifier on the target GPU core and the general video memory of the source VM.

[0053] In a possible implementation, when the sizes of the third addresses used by the command queues between the VMs are different, establishing, according to the migration command, a mapping between the hardware identifier corresponding to the target GPU core and the command queue of the source VM, and a mapping between the hardware identifier corresponding to the target GPU core and the general graphics memory of the source VM, further includes:

[0054] Obtaining, according to the migration command, a second page table of the target command queue, where the second page table indicates a mapping relationship between the second address of the command queue and the third address of the command queue, and between the second address of the general graphics memory and the third address of the general graphics memory;

[0055] The second page table of the target command queue is replaced with the second page table of the source VM command queue.

[0056] In a possible implementation, the device further includes:

[0057] The second establishing module is used to establish a first page table and a second page table for the command queue and general video memory corresponding to the hardware identification of the GPU core when the host driver is initialized.

[0058] In a possible implementation, the processing module is configured to:

[0059] In response to a write operation to a register group corresponding to a hardware identifier on a target GPU core, based on the association relationship after migration, using a microcontroller MCU of the target GPU core to obtain workload information from the source VM, the workload information including a second address corresponding to the workload and a third address of a page table root directory associated with the workload;

[0060] Using the MCU of the target GPU core, the second address corresponding to the workload and the third address of the page table root directory associated with the current workload are configured to the engine of the target GPU core, so that the engine of the target GPU core can address the host's video memory and process the current workload.

[0061] In a possible implementation, the first establishing module is configured to:

[0062] In the case where there are idle resources on the target GPU core, an association relationship is established between the source VM and the corresponding hardware identifiers on the target GPU core according to the migration command.

[0063] In a possible implementation, the migration command includes an identifier of a target GPU core and a hardware identifier corresponding to the target GPU core.

[0064] According to another aspect of the present disclosure, a graphics processing unit (GPU) scheduling device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0065] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.

[0066] According to another aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0067] According to an embodiment of the present application, by obtaining a migration command and establishing an association relationship between the source VM and the corresponding hardware identifiers on the target GPU core according to the migration command, based on the association relationship after migration, the VM can use the target GPU core to process the workload from the source VM without shutting down, thereby migrating the workload of the source VM from the high-loaded source GPU core to the idle target GPU core, realizing hot migration between multi-core GPU cores, balancing the load between multiple GPU cores, reducing the pressure on the high-load GPU core, shortening the response time, and improving the throughput.

[0068] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.

[0070] FIG1 is a schematic diagram showing an application scenario according to an embodiment of the present application.

[0071] FIG2 is a schematic diagram showing an application scenario according to an embodiment of the present application.

[0072] FIG3 shows a flowchart of a GPU scheduling method according to an embodiment of the present application.

[0073] FIG4 shows a flowchart of a GPU scheduling method according to an embodiment of the present application.

[0074] FIG5 is a schematic diagram showing a register group resource partition according to an embodiment of the present application.

[0075] FIG6 shows a flowchart of a GPU scheduling method according to an embodiment of the present application.

[0076] FIG7 shows a flowchart of a GPU scheduling method according to an embodiment of the present application.

[0077] FIG8 is a schematic diagram showing the division of graphics memory resources according to an embodiment of the present application.

[0078] FIG9 shows a flowchart of a GPU scheduling method according to an embodiment of the present application.

[0079] FIG10 is a schematic diagram showing an application scenario according to an embodiment of the present application.

[0080] FIG11 shows a flowchart of a GPU scheduling method according to an embodiment of the present application.

[0081] FIG12 shows a structural diagram of a GPU scheduling device according to an embodiment of the present application.

[0082] FIG13 is a block diagram showing a device 1900 for scheduling a GPU according to an exemplary embodiment. DETAILED DESCRIPTION

[0083] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0084] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0085] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0086] GPUs are very important in areas such as graphics rendering, parallel computing, and artificial intelligence. In GPU virtualization technology, in order to support multiple VMs using a single GPU simultaneously, the GPU's hardware resources are divided into multiple parts to provide each VM with independent hardware resources. Therefore, in the context of a multi-core GPU, the current technical solution typically processes workloads from VMs only on the GPU core initially selected at startup. Since VMs can be powered on and off at any time, it's possible that the workloads of multiple VMs are concentrated on a single GPU core, leaving other GPU cores idle. This not only wastes hardware resources, but also increases the pressure on the GPU core, leading to load imbalance between cores. Therefore, there is an urgent need for a new GPU scheduling method to balance loads on multiple GPU cores, shorten response times, and increase throughput.

[0087] In view of this, the present application proposes a graphics processor GPU scheduling method. The method of the embodiment of the present application obtains a migration command and establishes an association relationship between the source VM and the corresponding hardware identifiers on the target GPU core according to the migration command. Based on the association relationship after migration, the VM can use the target GPU core to process the workload from the source VM without shutting down, thereby migrating the workload of the source VM from the high-loaded source GPU core to the idle target GPU core, realizing hot migration between multi-core GPU cores. In this way, the load between multiple GPU cores can be balanced, the pressure on the high-loaded GPU core can be reduced, the response time can be shortened, and the throughput can be improved.

[0088] FIG1 shows a schematic diagram of an application scenario according to an embodiment of the present application. The GPU scheduling method of the embodiment of the present application can be applied to the scenario of virtualization on a multi-core GPU. GPU virtualization (graphics card virtualization) is about to divide hardware resources and allocate them to different virtual machines for use. In the scenario of a multi-core GPU, each GPU core can be divided into multiple resources, and different resources can be indicated by different hardware identifiers (hardware IDs), and each resource can be allocated to a VM. As shown in FIG1 , each GPU core (such as GPU core 1 and GPU core 2 in the figure) may include one or more engines (engines), microcontrollers (MCUs) and memory management units (MMUs). In the scenario of the embodiment of the present application, each GPU core also includes a D_MMU, and the D_MMU can be an input-output memory management unit (IOMMU) or an address translation unit. The GPU core can be connected to the video memory via a bus, and the video memory can include a command queue random access memory (RAM) and a general-purpose RAM.

[0089] After the VM's guest driver creates a workload, the MCU can dispatch the VM's workload to the engines for processing. During this process, when the guest driver creates the workload, the address in the commands to be processed by the engines is filled with the VM's virtual video memory address, called the GVA (guest virtual address). The workload can also include the host's virtual video memory address of the associated page table root directory, called the DVA (device virtual address), which is written to the video memory command queue. If a GPU core register is written, the MCU can respond to the write operation to retrieve the VM's workload and assign it to the corresponding engine for processing. When the MCU schedules a task, in addition to sending the workload content to the engine, it also sets the workload's associated page table root directory to the engine. The MMU can convert the GVA to the DVA, and the D_MMU can convert the DVA to the host's physical video memory address, called the DPA (device physical address), which can be sent to the bus to address the video memory. The engine can then access video memory based on the workload's GVA and the DVA of the workload's associated page table root directory to process the corresponding workload. The MCU needs to be able to access the command queues of all VMs, so space must be reserved for all supported VM command queues on the current GPU core and mapped to the MCU's GVA (this is accomplished by the subsequent GPU core management module).

[0090] In the above scenario, the VM workload is typically processed on the GPU core selected when the device is created. This inevitably leads to load imbalance between GPU cores. For example, if the GPU has four cores, each core is divided into eight resources, corresponding to eight hardware IDs. This means that each core can support up to eight VMs. Since VMs can be powered on and off at any time, the workloads of all eight VMs may be concentrated on a single core, while the other cores are idle. This creates load imbalance between cores, wastes hardware resources, and impacts the user experience.

[0091] Based on this, Figure 2 shows a schematic diagram of an application scenario according to an embodiment of the present application. Based on the GPU scheduling system of the embodiment of the present application shown in Figure 2, workloads from VMs can be migrated between cores, scheduling the GPU and balancing the load on multiple cores to address the aforementioned imbalanced load between cores. As shown in Figure 2, the GPU scheduling system may include a driver control module, a GPU core management module, a D_MMU control module, and a virtual machine manager (VMM) memory management module.

[0092] Among them, the driver control module is a module in the host driver, which can be used to provide an interface for migration commands to the application. It registers the misc device with the kernel during the initialization phase and provides multiple control APIs to the control program through the driver's ioctl callback. Alternatively, it can be implemented through other APIs provided by the operating system. The control API includes but is not limited to querying the GPU coreID, hardware ID, and VM workload migration control. When controlling the VM workload migration, the parameters passed to the module are the GPU core ID and hardware ID of the source VM, and the return value indicates success or failure.

[0093] The GPU core management module can be used to initialize the GPU core (which may include command queue initialization, MCU initialization (creating a page table for the MCU firmware and setting corresponding registers so that the GVA can be used after the firmware is started), D_MMU initialization, etc.), initialize and map the part of the video memory reserved for management by the GPU core, maintain the correspondence between the GPU core and the hardware ID, and between the hardware ID and the VM, apply for interrupt numbers, register interrupt handling functions, load MCU firmware, etc.

[0094] The D_MMU control module can be used to provide an external interface for initializing and configuring the D_MMU on the GPU core. The input parameters are the GPU core ID, hardware ID, the mapping destination address DPA, DVA, and the video memory address that may be used to store the page table. The VMM memory management module can be provided by the hypervisor to establish and maintain the second-stage page table for memory virtualization.

[0095] In addition, the GPU core can also provide functions such as creation, destruction, and control of virtual devices. These functions can be implemented through the IO virtualization framework provided by the hypervisor.

[0096] The following, based on Figures 1 and 2, describes in detail the GPU scheduling method of the embodiment of the present application through Figures 3 to 9.

[0097] 3, which shows a flow chart of a GPU scheduling method according to an embodiment of the present application. The GPU scheduling method of the embodiment of the present application can be applied to the above-mentioned GPU scheduling system. As shown in FIG3, the method may include:

[0098] Step S301: Obtain a migration command.

[0099] The migration command is used to migrate the source VM's workload from the source GPU core to the target GPU core. The source GPU core may have a heavier load than the target GPU core. In other words, the source GPU core may have a larger number of VM workloads, while the target GPU core may have a smaller number of VM workloads. Therefore, the migration command can be used to migrate the source VM's workload from the source GPU core to the target GPU core.

[0100] The migration command can be input by the user or automatically generated according to the current load conditions between GPU cores. The migration command may include the identifier of the target GPU core and the corresponding hardware identifier on the target GPU core. For example, the identifier of the target GPU core and the corresponding hardware identifier (i.e., hardware ID) on the target GPU core can be obtained by the above-mentioned driver control module. The identifier of the target GPU core can be used to indicate a unique target GPU core, and the corresponding hardware ID on the target GPU core can be used to indicate one of the resources on the target GPU core. Since each resource on the GPU can be associated with a VM, the association relationship between the source VM and the source hardware ID can be changed to an association relationship with the target hardware ID through the migration command, and an association relationship between the source VM and the target hardware ID is established. See below. The source hardware ID can be used to indicate one of the resources on the source GPU core.

[0101] Step S302: establishing an association relationship between the source VM and the corresponding hardware identifiers on the target GPU core according to the migration command.

[0102] Since the target GPU core can be divided into multiple resources, the hardware identifier (hardware ID) corresponding to the target GPU core can be used to identify one of the resources on the target GPU core. That is, the hardware ID corresponds to a specific resource on the target GPU core. By establishing an association between the source VM and the hardware ID, the workload of the source VM can be migrated to a resource on the target GPU core corresponding to the hardware ID for processing.

[0103] Optionally, step S302 may include:

[0104] In the case where there are idle resources on the target GPU core, an association relationship is established between the source VM and the corresponding hardware identifiers on the target GPU core according to the migration command.

[0105] For example, the GPU core management module can determine whether there are idle resources on the target GPU core. Based on the correspondence between the hardware ID of the current target GPU core and the VM, it can determine whether there is a hardware ID on the target GPU core that can be assigned to the source VM. If there are no idle resources on the target GPU core, the module can return a corresponding message to the driver control module, which can then return a message indicating a migration failure to the application.

[0106] This allows for flexible migration of VM workloads between GPU cores based on the current load between GPU cores, shortening response times, increasing GPU throughput, and enhancing user experience. The VM doesn't need to be aware of this process.

[0107] The implementation of step S302 will be described in detail below.

[0108] FIG4 shows a flow chart of a GPU scheduling method according to an embodiment of the present application. As shown in FIG4 , step S302 may include:

[0109] Step S401 : establishing a mapping between a source VM and a register group corresponding to a hardware identifier on a target GPU core according to a migration command.

[0110] In GPU virtualization scenarios, one of the goals is to enable multiple VMs to use a single GPU device simultaneously. This means supporting multiple VMs submitting workloads to the GPU concurrently; enabling the GPU to identify workloads from different VMs; and notifying the VM initiating the workload of completion events. To achieve this, the GPU core can partition hardware resources, including interrupt lines, register banks, and video memory, into multiple partitions, providing each VM with independent hardware resources.

[0111] The register group is fixed in length and provides the necessary registers for workload and interrupt handling. It is bound to the GPU core and cannot be accessed across GPU cores. A register group can include registers necessary for workload and interrupt handling, such as doorbell registers, interrupt status registers, and interrupt control registers.

[0112] Hardware identifiers can be used to represent different parts of hardware resources, so that the hardware identifiers and register groups can be mapped one to one. In a virtualized scenario, by associating the hardware identifier with the VM, the hardware identifier can also be used to index the VM. Referring to Figure 5, a schematic diagram of the division of register group resources according to an embodiment of the present application is shown. As shown in Figure 5, the hardware resources related to the register group can be divided into n parts, each of which can be assigned to a VM, then the n register groups can correspond to n VMs respectively (such as VM1, VM2...VMn in the figure). The register resources may be the video memory on the GPU card or the system memory, and the GPU provides a certain degree of isolation through the on-chip MMU. All register resources are visible to the MCU to support the scheduling of different VMs. The access isolation of register resources is guaranteed by the GPU itself and the hypervisor using the mechanism of the hardware virtual machine, and the allocation of register resources can be the responsibility of the host driver.

[0113] Because hardware identifiers correspond one-to-one with register groups, and hardware identifiers are associated with VMs, in this application, to establish an association between a source VM and the corresponding hardware identifier on a target GPU core, the mapping between the source VM and the source register group can be changed first, and a new mapping between the source VM and a new register group can be established. This new register group can be the register group corresponding to the hardware identifier on the target GPU core. The method for establishing a new mapping between the source VM and the new register group is described below.

[0114] FIG6 shows a flow chart of a GPU scheduling method according to an embodiment of the present application. As shown in FIG6 , step S401 may include:

[0115] Step S601 : acquiring, according to a migration command, a first address of a register group corresponding to a hardware identifier on a source GPU core and a first address of a register group corresponding to a hardware identifier on a target GPU core.

[0116] The length of the register group allocated to each VM may be fixed, and the first address may be used to indicate the physical video memory address of the host, ie, the above-mentioned DPA.

[0117] Step S602: Based on the first address of the register group corresponding to the hardware identifier on the source GPU core, the secondary page table of the source VM is updated, so that the updated secondary page table indicates the mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core, so as to establish a mapping between the source VM and the register group corresponding to the hardware identifier on the target GPU core.

[0118] The first address is the DPA, and the second address can be used to indicate the VM's virtual memory address, namely the GVA. The second-stage page table can be used to indicate the mapping relationship between the VM's GVA and the register group's DPA. The second-stage page table can be established and maintained by the VMM memory management module.

[0119] Optionally, the process of updating the secondary page table of the source VM based on the first address of the register group corresponding to the hardware identifier on the source GPU core to establish a mapping between the source VM and the register group corresponding to the hardware identifier on the target GPU core can be implemented based on a trap mechanism. As described below, step S602 may include:

[0120] Based on the first address of the register group corresponding to the hardware identifier on the source GPU core, the corresponding segment in the secondary page table of the source VM is changed so that the access to the register group corresponding to the hardware identifier on the source GPU core is trapped. After the trap, the secondary page table of the source VM is updated so that the updated secondary page table indicates the mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core.

[0121] The trap generated when accessing the register group corresponding to the hardware identifier may be a trap generated when accessing registers such as doorbell in the register group.

[0122] After the trap, a predetermined error handling function in the host driver can be called to perform error handling registered by the GPU core, and based on the first address of the register group corresponding to the hardware identifier on the target GPU core, the hypervisor mapping interface is called to update the secondary page table, so that the updated secondary page table indicates the mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core.

[0123] For example, changing the corresponding fragment in the secondary page table of the source VM may include: clearing the content in the 2nd-stage page table allocated to the source VM through the VMM management module, and caching information such as the identifier of the target GPU core and the corresponding hardware ID on the target GPU core; adding an illegal value in the corresponding command to cause access to registers such as doorbell to be trapped; after the trap, the corresponding error handling function in the host driver can be called to perform error handling of GPU core registration. The error handling process may include using the cached target GPU core identifier, the corresponding hardware ID on the target GPU core, and other information to determine the DPA (i.e., the first address) of the register group corresponding to the hardware identifier on the target GPU core, and calling the hypervisor mapping interface based on the DPA to re-establish the 2nd-stage page table for the source VM, so that the new 2nd-stage page table indicates the mapping relationship between the GVA of the source VM and the DPA of the register group corresponding to the hardware identifier on the target GPU core, thereby updating the secondary page table of the source VM and obtaining an updated secondary page table.

[0124] During this process, the status of the current host can also be updated to the migration status.

[0125] According to an embodiment of the present application, it is possible to update the secondary page table using a trap mechanism, so that the new secondary page table can indicate the mapping relationship between the GVA of the source VM and the DPA of the register group corresponding to the hardware identifier on the target GPU core, so as to realize the migration of the mapping relationship related to the register group, creating conditions for balancing the load between GPU cores.

[0126] Because the only channel for the guest driver to send messages to the MCU is the command queue, when the MCU receives a doorbell register write event, the only sideband signal is the hardware ID. In order to allow the MCU to easily obtain the location of the command queue, the hardware ID and the command queue can be made to correspond one to one.

[0127] Returning to Figure 4, in order to establish an association relationship between the source VM and the corresponding hardware identifiers on the target GPU core in this application, the mapping between the source hardware ID and the command queue of the source VM, and between the source hardware ID and the general video memory of the source VM can also be changed to establish a new mapping, as shown below.

[0128] Step S402 : establishing a mapping between the corresponding hardware identifier on the target GPU core and the command queue of the source VM, and a mapping between the corresponding hardware identifier on the target GPU core and the general graphics memory of the source VM according to the migration command.

[0129] The hardware identifier corresponding to the target GPU core can be the hardware ID of the resource to which the source VM's workload is to be migrated. The location of the command queue in the hardware can be seen in the command queue RAM in Figure 1, and the location of the general-purpose video memory in the hardware can be seen in the general-purpose RAM in Figure 1. General-purpose video memory can be used to store the GPU core page table (i.e., the second page table) created by the VM's guest driver, workload command data, and so on.

[0130] The mappings between the hardware ID and the command queue of the source VM, and between the hardware ID and the general video memory of the source VM that existed before the new mapping was established can be statically set by the GPU core management module during the host driver initialization phase.

[0131] Before S301 , the GPU core management module may be used for initialization. When the host driver is initialized, a first page table and a second page table may be established for the command queue and general video memory corresponding to the hardware identifier of the GPU core.

[0132] According to an embodiment of the present application, by establishing a mapping between the register groups of the source VM and the corresponding hardware identifiers on the target GPU core, a mapping between the corresponding hardware identifiers on the target GPU core and the command queue of the source VM, and a mapping between the corresponding hardware identifiers on the target GPU core and the general video memory of the source VM according to the migration command, the workload of the source VM can be migrated from the high-loaded source GPU core to the idle target GPU core, so as to balance the load among multiple GPU cores, reduce the pressure on the high-load GPU core, shorten the response time, and improve the throughput.

[0133] The following is a detailed description of the process of step S402. In this application, by introducing a third address (i.e., the aforementioned DVA) during this process, it is not necessary to change the DPA during the process of establishing the command queue and remapping the general-purpose video memory. This will not trigger a trap during this process, eliminating the need for the central processing unit (CPU) to perform additional operations. Referring to FIG. 7 , a flowchart of a GPU scheduling method according to an embodiment of the present application is shown. As shown in FIG. 7 , step S402 may include:

[0134] Step S701 : Obtain the first page table of the target command queue according to the migration command.

[0135] Based on the target GPU core identifier and hardware ID in the migration command, the target command queue associated with the hardware ID can be determined to obtain its first page table. This process can be implemented using the memory virtualization technology provided by the hypervisor and will not be detailed here.

[0136] Among them, the first page table can indicate the mapping relationship between the third address of the command queue and the first address of the command queue, and between the third address of the general video memory and the first address of the general video memory. The target command queue can be the command queue indicated by the corresponding hardware identifier on the target GPU core. The first address can be used to indicate the physical video memory address of the host (that is, the above-mentioned DPA), and the third address can be used to indicate the virtual video memory address of the host (that is, the above-mentioned DVA).

[0137] Step S702, when the size of the third address used by the command queues between VMs is the same, replace the first page table of the target command queue with the first page table of the source VM command queue to establish a mapping between the corresponding hardware identifier on the target GPU core and the command queue of the source VM, and a mapping between the corresponding hardware identifier on the target GPU core and the general graphics memory of the source VM.

[0138] Refer to Figure 8, which shows a schematic diagram of video memory resource partitioning according to an embodiment of the present application. As shown in Figure 8, video memory-related hardware resources can be divided into n shares, each of which can include command queue RAM and general RAM. Each share can be allocated to a VM, so the n shares of resources can correspond to n VMs (such as VM1, VM2, ..., VMn in the figure).

[0139] Based on the memory resource division diagram shown in Figure 8, since each resource indicated by the hardware ID includes the command queue and the general video memory, when the first page table of the command queue is replaced, the first page table of the general video memory will also be replaced, thereby establishing a new mapping, namely, the mapping between the corresponding hardware ID on the target GPU core and the command queue of the source VM, and the mapping between the corresponding hardware ID on the target GPU core and the general video memory of the source VM.

[0140] In this application, by introducing D_MMU (see Figure 1), the first page table of the source VM command queue can represent the mapping relationship between DVA and DPA. At this time, all VMs can use the same DVA space, and the only difference is the DPA. In this way, since the size of the third address used by the command queues between VMs is the same, only the first page table needs to be replaced, so that the remapping of the VM general video memory does not require additional operations. Since the DPA of the VM command queue and general video memory has not changed at this time, and the guest's access to the command queue and video memory is implemented through the memory virtualization technology provided by the hypervisor, the access will not trigger a trap. Therefore, from the CPU's perspective, no additional operations are required at this time.

[0141] Optionally, when the sizes of the third addresses used by the command queues of different VMs are different, that is, when the DVA spaces used by the VMs are different, step S402 may further include:

[0142] According to the migration command, a second page table of the target command queue is obtained; and the second page table of the target command queue is replaced with the second page table of the source VM command queue.

[0143] The second page table can indicate the mapping relationship between the second address of the command queue and the third address of the command queue, and between the second address of the general graphics memory and the third address of the general graphics memory. The second address can be used to indicate the virtual graphics memory address of the VM (i.e., the aforementioned GVA). That is, in this case, when the first page table is replaced, the second page table must also be replaced.

[0144] In this way, the workload of the source virtual machine VM can be migrated from the source GPU core to the target GPU core. When the MCU on the target GPU core subsequently uses GVA to access the command queue associated with the corresponding hardware ID, it can access the command queue of the source VM.

[0145] Referring back to FIG. 3 , after the migration is complete, the workload from the source VM can be processed by the target GPU core, as described below.

[0146] Step S303: Based on the migrated association relationship, the target GPU core is used to process the workload from the source VM.

[0147] That is, the workload from the source VM can be processed using the target GPU core based on the association between the migrated source VM and the target hardware ID. The detailed process is described below.

[0148] According to an embodiment of the present application, by obtaining a migration command and establishing an association relationship between the source VM and the corresponding hardware identifiers on the target GPU core according to the migration command, based on the association relationship after migration, the VM can use the target GPU core to process the workload from the source VM without shutting down, thereby migrating the workload of the source VM from the high-loaded source GPU core to the idle target GPU core, realizing hot migration between multi-core GPU cores, balancing the load between multiple GPU cores, reducing the pressure on the high-load GPU core, shortening the response time, and improving the throughput.

[0149] FIG9 is a flowchart of a GPU scheduling method according to an embodiment of the present application. As shown in FIG9 , step S303 may include:

[0150] Step S901 : in response to a write operation on a register group corresponding to a hardware identifier on a target GPU core, based on the association relationship after migration, the microcontroller MCU of the target GPU core is used to obtain workload information from a source VM.

[0151] Among them, the write operation on the register group corresponding to the hardware identifier on the target GPU core can be a write operation on the doorbell register in the register group, and the workload information can include the second address corresponding to this workload (i.e., GVA) and the third address of the page table root directory associated with this workload (i.e., DVA).

[0152] Step S902: Using the MCU of the target GPU core, the second address corresponding to the current workload and the third address of the page table root directory associated with the current workload are configured to the engine of the target GPU core, so as to use the engine of the target GPU core to address the host's video memory and process the current workload.

[0153] Based on the secondary page table, the GVA can be converted to the DPA to obtain the DPA of the current workload. The D_MMU can also be used to convert the DVA to the DPA based on the primary page table to obtain the DPA of the page table root directory associated with the current workload. Based on the DPA of the page table root directory and the DPA of the workload, the target GPU core engine can address and access the host's video memory to process the current workload, and notify the source VM after processing.

[0154] This allows for load balancing between GPU cores. After migration, VMs associated with other hardware IDs on the target GPU core remain unaffected, effectively utilizing existing hardware and improving user experience.

[0155] Through the use of DVA in the embodiments of the present application, there is no need to copy the command queue content during migration, and there is no need to include DPA in the workload. During migration, there is no need to parse or modify the DPA in the command to the DPA of the new GPU core during the initialization phase, making the migration process more efficient.

[0156] This application also provides a GPU scheduling method.

[0157] Figure 10 shows a schematic diagram of an application scenario according to an embodiment of the present application. As shown in Figure 10, the GPU scheduling method of the embodiment of the present application can be applied to the scenario of virtualization on a multi-core GPU. In the scenario of a multi-core GPU, each GPU core of the multi-core GPU (such as GPU core 1 and GPU core 2 in the figure) runs independently without affecting each other, and each GPU core can be divided into multiple resources. Different resources can be indicated by different hardware identifiers (hardware IDs), and each resource can be allocated to a VM. As shown in Figure 1, each GPU core (such as GPU core1 and GPU core2 in the figure) may include one or more engines (engine, responsible for executing workload), microcontroller (micro controller unit, MCU) and memory management unit (memory management unit, MMU). GPU core can be connected to video memory via a bus, and video memory may include command queue random access memory (random access memory, RAM) and general RAM.

[0158] After the VM's guest driver creates a workload, the firmware running on the MCU is responsible for scheduling and dispatching the VM's workload to the engines for processing. During this process, when the guest driver creates the workload, the address in the commands to be processed by the engines is filled with the VM's virtual video memory address, called the GVA (guest virtual address). If a GPU core register is written, the MCU can respond to the write operation to retrieve the VM's workload and assign it to the corresponding engine for processing. The MMU can convert the GVA into the host's physical video memory address, called the DPA (device physical address). The DPA is sent to the bus to address the video memory. (On a terminal device or server, the graphics card is connected to the system via the PCIe bus, so the DPA in this case is the DPA of the graphics card's internal bus. When addressing video memory, DPA addressing can be used. If system memory is used as video memory, the DPA is converted to a PCIe bus domain address by the PCIe module and then used by the RC to address system memory.) The engine can then access the video memory based on the GVA in the workload to process the corresponding workload.

[0159] FIG11 is a flowchart of a GPU scheduling method according to an embodiment of the present application. As shown in FIG11 , the method may include:

[0160] Step S1101: Host driver initialization.

[0161] During initialization, the host driver allocates a hardware ID, register set, and a predetermined amount of video memory to the virtual machine (VM). The host driver also maintains a record of the correspondence between the hardware ID and the VM. The hardware ID represents the hardware resources (register set, video memory, etc.) allocated to the VM by the host driver.

[0162] The host driver can also construct initialization structures in the video memory allocated to the VM.

[0163] In step S1102 , the VM is started through a hypervisor (a type of system software), and the register group and video memory allocated to the VM are mapped to the address space of the VM to construct a secondary page table.

[0164] The second-stage page table (i.e., the 2nd-stage page table) can be used to indicate the mapping relationship between the VM's GVA and the DPA of the register group.

[0165] This enables the guest driver in the VM to start executing the corresponding workload.

[0166] In step S1103 , the guest driver in the VM allocates space for the workload from the allocated video memory and constructs a corresponding index to write into a command queue.

[0167] In step S1104 , the guest driver in the VM performs a write operation on the register corresponding to the hardware identifier on the GPU core based on the secondary page table.

[0168] This register may be a doorbell register corresponding to the hardware ID, and the GPU may be informed by this write operation that a new command has been added to the command queue.

[0169] Step S1105 : Based on a write operation to a register corresponding to a hardware identifier on the GPU core, the microcontroller MCU of the GPU core is used to obtain workload information from the VM.

[0170] The write operation on the register group corresponding to the hardware identifier on the GPU core may be a write operation on the doorbell register in the register group, and the workload information may include the GVA corresponding to the current workload.

[0171] Step S1106: Based on the workload information from the VM, the workload is executed by the MCU scheduling engine.

[0172] The GVA may be converted into the DPA based on the secondary page table to obtain the DPA of the current workload, so as to address the video memory of the host through the DPA and access the video memory to process the current workload.

[0173] Step S1107 : After the engine on the GPU core completes executing the workload, it sends an interrupt signal through the interrupt line corresponding to the hardware ID.

[0174] In step S1108 , the hypervisor (a type of system software) executes an interrupt handling process, determines the corresponding hardware ID through registers or interrupt signals, and injects the interrupt into the VM corresponding to the hardware ID for subsequent processing.

[0175] The register may be an interrupt status register.

[0176] In this way, the GPU can efficiently execute and schedule workloads.

[0177] FIG12 shows a structural diagram of a GPU scheduling device according to an embodiment of the present application. As shown in FIG12 , the device may include:

[0178] An acquisition module 1201 is configured to acquire a migration command, where the migration command is configured to migrate the workload of a source virtual machine VM from a source GPU core to a target GPU core.

[0179] A first establishing module 1202 is configured to establish an association relationship between the source VM and the corresponding hardware identifiers on the target GPU core according to the migration command;

[0180] The processing module 1203 is configured to process the workload from the source VM using the target GPU core based on the migrated association relationship.

[0181] In a possible implementation, the first establishing module 1202 is configured to:

[0182] According to the migration command, a mapping is established between the register groups corresponding to the hardware identifiers on the source VM and the target GPU core;

[0183] According to the migration command, a mapping is established between the corresponding hardware identifier on the target GPU core and the command queue of the source VM, and a mapping is established between the corresponding hardware identifier on the target GPU core and the general video memory of the source VM.

[0184] In one possible implementation, a mapping is established between register groups corresponding to hardware identifiers on the source VM and the target GPU core according to the migration command, including:

[0185] According to the migration command, obtain the first address of the register group corresponding to the hardware identifier on the source GPU core and the first address of the register group corresponding to the hardware identifier on the target GPU core, where the first address is used to indicate the physical video memory address of the host;

[0186] Based on the first address of the register group corresponding to the hardware identifier on the source GPU core, the secondary page table of the source VM is updated so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core, so as to establish a mapping between the register groups corresponding to the hardware identifier on the source VM and the target GPU core, where the second address is used to indicate the virtual video memory address of the VM.

[0187] In one possible implementation, updating the secondary page table of the source VM based on the first address of the register group corresponding to the hardware identifier on the source GPU core, so that the updated secondary page table indicates a mapping between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core, includes:

[0188] Based on the first address of the register group corresponding to the hardware identifier on the source GPU core, the corresponding segment in the secondary page table of the source VM is changed so that the access to the register group corresponding to the hardware identifier on the source GPU core is trapped;

[0189] After the trap, the secondary page table of the source VM is updated so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core.

[0190] In one possible implementation, updating the secondary page table of the source VM after the trap so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core includes:

[0191] After the trap, a predetermined error handling function in the host driver is called to perform error handling registered by the GPU core. Based on the first address of the register group corresponding to the hardware identifier on the target GPU core, the hypervisor mapping interface is called to update the secondary page table, so that the updated secondary page table indicates the mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core.

[0192] In one possible implementation, establishing a mapping between the hardware identifier corresponding to the target GPU core and the command queue of the source VM, and a mapping between the hardware identifier corresponding to the target GPU core and the general video memory of the source VM according to the migration command includes:

[0193] Obtaining, according to the migration command, a first page table of a target command queue, the first page table indicating a mapping relationship between a third address of the command queue and a first address of the command queue, and between a third address of the general video memory and a first address of the general video memory, the target command queue being a command queue indicated by a corresponding hardware identifier on a target GPU core, the first address being used to indicate a physical video memory address of the host, and the third address being used to indicate a virtual video memory address of the host;

[0194] When the size of the third address used by the command queues between VMs is the same, the first page table of the target command queue is replaced with the first page table of the source VM command queue to establish a mapping between the corresponding hardware identifier on the target GPU core and the command queue of the source VM, and a mapping between the corresponding hardware identifier on the target GPU core and the general video memory of the source VM.

[0195] In a possible implementation, when the sizes of the third addresses used by the command queues between the VMs are different, establishing, according to the migration command, a mapping between the hardware identifier corresponding to the target GPU core and the command queue of the source VM, and a mapping between the hardware identifier corresponding to the target GPU core and the general graphics memory of the source VM, further includes:

[0196] Obtaining, according to the migration command, a second page table of the target command queue, where the second page table indicates a mapping relationship between the second address of the command queue and the third address of the command queue, and between the second address of the general graphics memory and the third address of the general graphics memory;

[0197] The second page table of the target command queue is replaced with the second page table of the source VM command queue.

[0198] In a possible implementation, the device further includes:

[0199] The second establishing module is used to establish a first page table and a second page table for the command queue and general video memory corresponding to the hardware identification of the GPU core when the host driver is initialized.

[0200] In a possible implementation, the processing module 1203 is configured to:

[0201] In response to a write operation to a register group corresponding to a hardware identifier on a target GPU core, based on the association relationship after migration, using a microcontroller MCU of the target GPU core to obtain workload information from the source VM, the workload information including a second address corresponding to the workload and a third address of a page table root directory associated with the workload;

[0202] Using the MCU of the target GPU core, the second address corresponding to the workload and the third address of the page table root directory associated with the current workload are configured to the engine of the target GPU core, so that the engine of the target GPU core can address the host's video memory and process the current workload.

[0203] In a possible implementation, the first establishing module 1202 is configured to:

[0204] In the case where there are idle resources on the target GPU core, an association relationship is established between the source VM and the corresponding hardware identifiers on the target GPU core according to the migration command.

[0205] In a possible implementation, the migration command includes an identifier of a target GPU core and a hardware identifier corresponding to the target GPU core.

[0206] According to an embodiment of the present application, by obtaining a migration command and establishing an association relationship between the source VM and the corresponding hardware identifiers on the target GPU core according to the migration command, based on the association relationship after migration, the VM can use the target GPU core to process the workload from the source VM without shutting down, thereby migrating the workload of the source VM from the high-loaded source GPU core to the idle target GPU core, realizing hot migration between multi-core GPU cores, balancing the load between multiple GPU cores, reducing the pressure on the high-load GPU core, shortening the response time, and improving the throughput.

[0207] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0208] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0209] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0210] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0211] FIG13 is a block diagram of an apparatus 1900 for scheduling a GPU, according to an exemplary embodiment. For example, apparatus 1900 may be provided as a server or a terminal device. Referring to FIG13 , apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions executable by processing component 1922, such as applications. The applications stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, processing component 1922 is configured to execute instructions to perform the above-described method.

[0212] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958 (I / O interface). The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2003. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM or similar.

[0213] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the apparatus 1900 to perform the above-described method.

[0214] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0215] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0216] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0217] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0218] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0219] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0220] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0221] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0222] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A graphics processor GPU scheduling method, characterized in that: The method comprises: Obtain a migration command, where the migration command is used to migrate the workload of the source virtual machine VM from the source GPU core to the target GPU core; According to the migration command, an association relationship is established between the source VM and the corresponding hardware identifier on the target GPU core; Based on the association relationship after migration, the target GPU core is used to process the workload from the source VM.

2. The method according to claim 1, characterized in that The step of establishing an association relationship between the source VM and a hardware identifier corresponding to the target GPU core according to the migration command includes: According to the migration command, a mapping is established between the source VM and the register group corresponding to the hardware identification on the target GPU core; According to the migration command, a mapping between the corresponding hardware identifier on the target GPU core and the command queue of the source VM, and a mapping between the corresponding hardware identifier on the target GPU core and the general video memory of the source VM are established.

3. The method according to claim 2, characterized in that The step of establishing a mapping between the source VM and a register group corresponding to the hardware identifier on the target GPU core according to the migration command includes: According to the migration command, obtaining a first address of a register group corresponding to a hardware identifier on a source GPU core and a first address of a register group corresponding to a hardware identifier on a target GPU core, wherein the first address is used to indicate a physical video memory address of a host; Based on the first address of the register group corresponding to the hardware identifier on the source GPU core, the secondary page table of the source VM is updated so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core, so as to establish a mapping between the source VM and the register group corresponding to the hardware identifier on the target GPU core, wherein the second address is used to indicate the virtual video memory address of the VM.

4. The method according to claim 3, characterized in that: The method of updating the secondary page table of the source VM based on the first address of the register group corresponding to the hardware identifier on the source GPU core, so that the updated secondary page table indicates a mapping between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core, includes: Based on the first address of the register group corresponding to the hardware identifier on the source GPU core, the corresponding segment in the secondary page table of the source VM is changed so that the access to the register group corresponding to the hardware identifier on the source GPU core is trapped; After the trap, the secondary page table of the source VM is updated so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core.

5. The method according to claim 4, characterized in that The updating of the secondary page table of the source VM after the trap so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core includes: After the trap, a predetermined error handling function in the host driver is called to perform error handling registered by the GPU core, and based on the first address of the register group corresponding to the hardware identifier on the target GPU core, a hypervisor mapping interface is called to update the secondary page table, so that the updated secondary page table indicates a mapping relationship between the second address of the source VM and the first address of the register group corresponding to the hardware identifier on the target GPU core.

6. The method according to claim 2, characterized in that The step of establishing, according to the migration command, a mapping between the hardware identifier corresponding to the target GPU core and the command queue of the source VM, and a mapping between the hardware identifier corresponding to the target GPU core and the general video memory of the source VM, comprises: According to the migration command, a first page table of the target command queue is obtained, wherein the first page table indicates a mapping relationship between a third address of the command queue and a first address of the command queue, and between a third address of the general video memory and a first address of the general video memory, wherein the target command queue is a command queue indicated by a corresponding hardware identifier on a target GPU core, wherein the first address is used to indicate a physical video memory address of a host, and the third address is used to indicate a virtual video memory address of the host; When the size of the third address used by the command queues between VMs is the same, the first page table of the target command queue is replaced by the first page table of the source VM command queue to establish a mapping between the corresponding hardware identifier on the target GPU core and the command queue of the source VM, and a mapping between the corresponding hardware identifier on the target GPU core and the general video memory of the source VM.

7. The method according to claim 2, characterized in that: The step of establishing, according to the migration command, a mapping between the hardware identifier corresponding to the target GPU core and the command queue of the source VM, and a mapping between the hardware identifier corresponding to the target GPU core and the general video memory of the source VM, comprises: When the sizes of the third addresses used by the command queues between the VMs are different, obtaining, according to the migration command, a second page table of the target command queue, wherein the second page table indicates a mapping relationship between the second address of the command queue and the third address of the command queue, and between the second address of the general video memory and the third address of the general video memory; Replacing the second page table of the target command queue with the second page table of the source VM command queue; According to the migration command, a first page table of the target command queue is obtained, wherein the first page table indicates a mapping relationship between a third address of the command queue and a first address of the command queue, and between a third address of the general video memory and a first address of the general video memory, wherein the target command queue is a command queue indicated by a corresponding hardware identifier on a target GPU core, and the first address is used to indicate a physical video memory address of the host; The first page table of the target command queue is replaced with the first page table of the source VM command queue to establish a mapping between the corresponding hardware identifier on the target GPU core and the command queue of the source VM, and a mapping between the corresponding hardware identifier on the target GPU core and the general video memory of the source VM.

8. The method according to claim 7, characterized in that The method further comprises: When the host driver is initialized, a first page table and a second page table are established for the command queue and general video memory corresponding to the hardware identification of the GPU core.

9. The method according to any one of claims 1 to 8, characterized in that: The using the target GPU core to process the workload from the source VM based on the association relationship after the migration includes: In response to a write operation to a register group corresponding to a hardware identifier on a target GPU core, based on the association relationship after migration, using a microcontroller MCU of the target GPU core to obtain workload information from a source VM, wherein the workload information includes a second address corresponding to the current workload and a third address of a page table root directory associated with the current workload; Using the MCU of the target GPU core, the second address corresponding to the current workload and the third address of the page table root directory associated with the current workload are configured to the engine of the target GPU core, so that the engine of the target GPU core can address the video memory of the host and process the current workload.

10. The method according to claim 1, characterized in that The step of establishing an association relationship between the source VM and a hardware identifier corresponding to the target GPU core according to the migration command includes: In the case where there are idle resources on the target GPU core, an association relationship between the source VM and the corresponding hardware identifier on the target GPU core is established according to the migration command.

11. The method according to claim 1, characterized in that: The migration command includes an identifier of a target GPU core and a corresponding hardware identifier on the target GPU core.

12. A graphics processor GPU scheduling device, characterized in that: The device comprises: An acquisition module is used to acquire a migration command, where the migration command is used to migrate the workload of the source virtual machine VM from the source GPU core to the target GPU core; A first establishing module, configured to establish an association relationship between a source VM and a corresponding hardware identifier on the target GPU core according to the migration command; The processing module is used to process the workload from the source VM using the target GPU core based on the migrated association relationship.

13. A graphics processor GPU scheduling device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method described in any one of claims 1 to 11 when executing the instructions stored in the memory.

14. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Techniques for load balancing GPU enabled virtual machines

    CN102402462A

  • Method for non-uniform I / O access of virtual machine resource migration in virtual multi-nuclear environment

    CN106095576A

  • Virtual machine migration method and device, upgrading method and server

    CN115599494A

  • Graphic processing unit GPU scheduling method and device and storage medium

    CN117331704A

  • Prepopulating page tables for memory of workloads during live migrations

    US20230195533A1

Cited By

  • Hardware resource scheduling method, device and system, storage medium and program product

    CN120610827A

  • GPU virtualization system based on CUDA cross-level translation and multi-pooling scheduling

    CN121143952A

  • Process scheduling method and device, chip, electronic equipment and storage medium

    CN121255405A

  • Process management system, method and equipment and storage medium

    CN121900979A