Process migration method and device

By obtaining video memory snapshots and mapping relationships in the kernel state, the migration of process data and commands is solved, and the problem of reduced efficiency during the process migration is achieved and seamless process migration efficiency is achieved.

CN120256046APending Publication Date: 2025-07-04LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510279397.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

During the process migration process, the termination of the original process and recreating its state affects the process's running efficiency, resulting in a decrease in efficiency.

Method used

The memory snapshot and the second mapping relationship of the target process are obtained in the kernel state, and the process data is migrated to the second memory based on the storage address of the process data, and the write command is mapped to the second memory using the second mapping relationship to realize seamless migration of the process.

Benefits of technology

Process migration can be achieved without restarting the target process, which improves the efficiency of process migration and reduces the impact on process operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256046A_ABST
    Figure CN120256046A_ABST
Patent Text Reader

Abstract

The invention discloses a process migration method and device.The method comprises the steps that a video memory snapshot of a target process and a second mapping relation are obtained in a kernel mode, the video memory snapshot comprises process data of the target process and a storage address of the process data, and the second mapping relation is stored in the target process; the second mapping relationship is created based on a first mapping relationship and a second video memory, and the first mapping relationship is a mapping relationship between the first video memory and an address space of the target process; migrating the process data to the second video memory based on the storage address of the process data; and mapping the command written from the address space to the second video memory by using the second mapping relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of electronic technology, and relate to, but are not limited to, a process migration method and apparatus. Background Art

[0002] Process migration refers to moving a running process from the current location (or processor) to another specified processor, so that it can continue to access all resources and run on the other processor. Before migration, it is necessary to terminate (or exit) the original process, and recreate and restore the state of the process at the target location. The operation of terminating the original process affects the efficiency of the process running. How to reduce the impact on the process efficiency during the process migration has become a technical problem to be solved urgently. Summary of the Invention

[0003] In view of this, the embodiments of the present application provide a process migration method, apparatus, device and storage medium.

[0004] The technical solution of the embodiments of the present application is implemented as follows:

[0005] In a first aspect, the embodiments of the present application provide a process migration method, including:

[0006] Obtaining a video memory snapshot and a second mapping relationship of a target process in the kernel state, where the video memory snapshot includes the process data of the target process and the storage address of the process data, and the second mapping relationship is created based on a first mapping relationship and a second video memory, and the first mapping relationship is the mapping relationship between the first video memory and the address space of the target process;

[0007] Migrating the process data to the second video memory based on the storage address of the process data;

[0008] Mapping the command written to the address space to the second video memory by using the second mapping relationship.

[0009] In a second aspect, the embodiments of the present application provide a process migration apparatus, including:

[0010] A first acquisition module, configured to obtain a video memory snapshot and a second mapping relationship of a target process in the kernel state, where the video memory snapshot includes the process data of the target process and the storage address of the process data, and the second mapping relationship is created based on a first mapping relationship and a second video memory, and the first mapping relationship is the mapping relationship between the first video memory and the address space of the target process;

[0011] A migration module, configured to migrate the process data to the second video memory based on the storage address of the process data;

[0012] A writing module, configured to map the command written to the address space to the second video memory by using the second mapping relationship.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it realizes obtaining a video memory snapshot and a second mapping relationship of a target process in the kernel mode. The video memory snapshot includes the process data of the target process and the storage address of the process data. The second mapping relationship is created based on the first mapping relationship and the second video memory. The first mapping relationship is the mapping relationship between the first video memory and the address space of the target process; migrating the process data to the second video memory based on the storage address of the process data; and mapping the command written from the address space to the second video memory by using the second mapping relationship.

[0014] In a fourth aspect, an embodiment of the present application provides a storage medium storing executable instructions, which are used to realize, when executed by a processor, obtaining a video memory snapshot and a second mapping relationship of a target process in the kernel mode. The video memory snapshot includes the process data of the target process and the storage address of the process data. The second mapping relationship is created based on the first mapping relationship and the second video memory. The first mapping relationship is the mapping relationship between the first video memory and the address space of the target process; migrating the process data to the second video memory based on the storage address of the process data; and mapping the command written from the address space to the second video memory by using the second mapping relationship.

[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program or instructions. When the computer program or instructions are executed by a processor, it realizes obtaining a video memory snapshot and a second mapping relationship of a target process in the kernel mode. The video memory snapshot includes the process data of the target process and the storage address of the process data. The second mapping relationship is created based on the first mapping relationship and the second video memory. The first mapping relationship is the mapping relationship between the first video memory and the address space of the target process; migrating the process data to the second video memory based on the storage address of the process data; and mapping the command written from the address space to the second video memory by using the second mapping relationship. Description of the Drawings

[0016] Figure 1 It is a schematic flowchart of an implementation process of a process migration method provided by an embodiment of the present application;

[0017] Figure 2A It is a schematic diagram of a process migration scenario provided by an embodiment of the present application;

[0018] Figure 2B It is a schematic diagram of a process migration process provided by an embodiment of the present application;

[0019] Figure 2C It is a schematic flowchart of an implementation process of a process migration provided by an embodiment of the present application;

[0020] Figure 3AThis application provides a time-space division hybrid virtualization system in an embodiment;

[0021] Figure 3B The schematic diagram of the implementation process for determining the second video memory through time-space multiplexing provided in an embodiment of this application;

[0022] Figure 3C The time-division schematic diagram that meets the target throughput provided in an embodiment of this application;

[0023] Figure 3D The schematic diagram of peak throughput and instances provided in an embodiment of this application;

[0024] Figure 4A The schematic diagram of time-space hybrid orchestration provided in an embodiment of this application;

[0025] Figure 4B The schematic diagram of the merged result provided in an embodiment of this application;

[0026] Figure 4C The schematic diagram of resource utilization for different usage scenarios provided in an embodiment of this application;

[0027] Figure 5 The schematic diagram of the composition structure of a process migration device provided in an embodiment of this application;

[0028] Figure 6 The schematic diagram of a hardware entity of an electronic device provided in an embodiment of this application. Detailed implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will further describe the specific technical solutions of the embodiments of the application in detail with reference to the accompanying drawings in the embodiments of this application. The following embodiments are used to illustrate this application but are not used to limit the scope of this application.

[0030] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0031] In the following description, the terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.

[0033] An embodiment of this application provides a process migration method, as Figure 1 shown, the method includes:

[0034] Step S110, obtain a video memory snapshot and a second mapping relationship of a target process in kernel mode, where the video memory snapshot includes the process data of the target process and the storage address of the process data, and the second mapping relationship is created based on a first mapping relationship and a second video memory, and the first mapping relationship is the mapping relationship between a first video memory and the address space of the target process;

[0035] Here, kernel mode is the mode in which the operating system kernel runs and has the highest privilege level. In kernel mode, the system storage, external devices can be accessed without restriction, and various privileged operations can be performed, such as memory management, device driver, process scheduling, etc. Video memory is the memory on the graphics card and is used to store graphic data. The operating system and the graphics card driver jointly manage the video memory.

[0036] During implementation, a kernel mode virtualization module can be set in kernel mode to obtain a video memory snapshot and a second mapping relationship of the target process.

[0037] The target process is the process to be migrated, that is, the process is migrated from the first video memory to the second video memory. Here, the first video memory and the second video memory can be the video memories respectively corresponding to different instances obtained by multi-instance (Multi-Instance GPU, MIG) partitioning. During implementation, when the process is migrated from one instance to another instance, the process is determined as the target process.

[0038] The video memory snapshot of the target process refers to capturing and saving the state of the video memory (RAM) on the graphics card at a specific moment. Using the video memory snapshot, the data stored in the video memory can be copied or an image can be generated so that it can be restored to this specific state when needed. During implementation, the video memory snapshot can be generated based on the process data used by the target process and the storage address of the process data.

[0039] The second mapping relationship is created based on the first mapping relationship and the second video memory and is the mapping relationship between the second video memory and the address space of the target process. During implementation, the first mapping relationship can be generated first based on the mapping relationship between the first video memory and the address space of the target process, and then after the second video memory is determined, the second mapping relationship is created based on the first mapping relationship and the second video memory.

[0040] Step S120: Migrate the process data to the second video memory based on the storage address of the process data;

[0041] During implementation, the kernel-mode virtualization module can apply for the storage address in the video memory snapshot through the graphics driver to migrate the process data to the second video memory. Through the graphics driver, sufficient space is allocated in the second video memory to store the process data. The data address of the second video memory is set based on the storage address of the process data so that it points to a new address in the second video memory that is the same as the address in the first video memory, ensuring that the target process can correctly access and use the process data in the second video memory.

[0042] Step S130: Map the command written to the address space to the second video memory using the second mapping relationship.

[0043] Here, the target process has a corresponding virtual address space in the operating system for isolating and protecting the memory between processes. Memory mapping is a mechanism provided by the operating system that allows the video memory to be mapped into the address space of the process so that the process can access the video memory as if it were ordinary memory. The second mapping relationship is the mapping relationship between the second video memory and the address space of the target process.

[0044] During implementation, the second mapping relationship can be used to map the command written by the target process to the address space to the second video memory, realizing the migration of the command mapping relationship of the target process from being mapped to the first video memory to the second video memory.

[0045] When the process data of the target process is migrated to the second video memory and the commands of the target process are correspondingly mapped to the second video memory, the target process can be migrated from the first video memory to the second video memory for execution, that is, the instructions at the kernel level of the target process can be submitted to a new Graphics Processing Unit (GPU) split instance for continuous execution, and the process migration can be completed. Here, kernel is a proprietary term in GPU programming, referring to a code segment running on the GPU.

[0046] In the embodiments of the present application, first, the video memory snapshot and the second mapping relationship of the target process are obtained in the kernel mode, and then the process data is migrated to the second video memory based on the storage address of the process data in the video memory snapshot; the command written to the address space of the target process is mapped to the second video memory using the second mapping relationship. In this way, it is ensured that the target process can be migrated from the first video memory to the second video memory without restarting the target process, effectively improving the efficiency of process migration.

[0047] In some embodiments, the embodiments of the present application also provide a method for generating and storing a video memory snapshot, which can be implemented through the following steps:

[0048] Step S140: When it is determined that the target process is to be migrated, obtain the process data and the storage address of the process data;

[0049] Here, the target process is migrated from the first video memory to the second video memory. During implementation, the process data and the storage address of the process data can be obtained from the first video memory.

[0050] Step S150: Generate a video memory snapshot based on the process data and the storage address of the process data;

[0051] During implementation, the kernel-mode virtualization module can generate a video memory snapshot based on the obtained process data and the storage address of the process data.

[0052] Step S160: Store the video memory snapshot to the hard disk.

[0053] During implementation, since the video memory snapshot includes the process data and the storage address of the process data, the process data in the video memory mainly includes rendering data, vertex data, working instruction streams, and temporary data, etc. There is a high possibility that the data volume of this video memory snapshot is large. During implementation, the kernel-mode virtualization module can save this video memory snapshot to the hard disk.

[0054] In the embodiment of the present application, when it is determined that the target process is to be migrated, first obtain the process data and the storage address of the process data; then generate a video memory snapshot based on the process data and the storage address of the process data; finally, store the video memory snapshot to the hard disk. In this way, the process data and the storage address of the process data can be stored to the hard disk by using the generated video memory snapshot.

[0055] In some embodiments, the above step S110 of "obtaining the video memory snapshot and the second mapping relationship of the target process in the kernel mode" can be implemented through the following steps:

[0056] Step 111: Read the video memory snapshot from the hard disk in the kernel mode;

[0057] During implementation, the kernel-mode virtualization module reads the video memory snapshot from the hard disk to migrate the process data to the second video memory by using this video memory snapshot. Set the data address of the second video memory based on the storage address of the process data so that it points to a new address in the second video memory that is the same as the address in the first video memory.

[0058] Step 112: Obtain the first mapping relationship;

[0059] During implementation, the first mapping relationship can be generated based on the mapping relationship between the first video memory and the address space of the target process.

[0060] Step 113: Establish the second mapping relationship based on the first mapping relationship and the second video memory.

[0061] Here, the second mapping relationship is the mapping relationship between the second video memory and the address space of the target process.

[0062] During the process migration, in order to enable the instructions sent by the target process to the address space to be mapped to the second video memory, and the address information mapped to the second video memory to be consistent with the address information of the first video memory, the second mapping relationship can be established based on the first mapping relationship and the second video memory.

[0063] During the implementation process, there is no fixed order for obtaining the video memory snapshot and establishing the second mapping relationship, that is, the video memory snapshot can be obtained first, or the second mapping relationship can be established first.

[0064] In the embodiments of the present application, the video memory snapshot is read from the hard disk in the kernel state; the first mapping relationship is obtained to establish the second mapping relationship based on the first mapping relationship and the second video memory. In this way, it is possible to obtain the pre-generated video memory snapshot from the hard disk and establish the second mapping relationship based on the first mapping relationship and the second video memory.

[0065] In some embodiments, the above step 112 "obtain the first mapping relationship" can be implemented through the following steps:

[0066] Step 1121: Obtain the address space corresponding to the command written to the user process in the user state;

[0067] Here, the application (APP) running in the user state is correspondingly set to write to the command buffer area (Command buff), that is, the address space of the user process. The command buffer area is a memory area used to temporarily store command data for subsequent processing. In graphics rendering, the command buffer area is used to store a series of rendering commands, such as draw calls, setting rendering states, updating buffer data, etc. Through the command buffer, the rendering commands can be efficiently submitted to the GPU for execution, thereby improving the rendering performance. Mapping the content of the command buffer to the video memory enables the GPU to execute the rendering operations in the order of the commands. For example, the user state address of the first command can be 0x77.

[0068] Step 1122: Obtain the command address corresponding to the command in the first video memory;

[0069] During the implementation process, the command address stored in the first video memory corresponding to the command can be obtained, that is, before the process migration, the command address stored in the first video memory set for the pre-migration instance corresponding to the target process. For example, the first video memory address of the first command can be 0x55.

[0070] Step 1123: Generate the first mapping relationship based on the address space of the user process and the command address of the first video memory.

[0071] During implementation, the first mapping relationship can be generated based on the address space of the user process and the command address of the first video memory. For example, the mapping relationship of the first command is the mapping relationship between the user-mode address 0x77 and the first video memory address 0x55.

[0072] In the embodiments of the present application, first, obtain the address space corresponding to the command written to the user process in the user mode; then, obtain the command address of the command corresponding to the first video memory; finally, the first mapping relationship can be generated based on the address space of the user process and the command address of the first video memory.

[0073] In some embodiments, the above step 113, "Establish the second mapping relationship based on the first mapping relationship and the second video memory", can be implemented through the following steps:

[0074] Step 1131: Create the command address of the second video memory based on the command address of the first video memory.

[0075] For example, if the command address of the first command in the first video memory is 0x55, then the address of the first command can be created as 0x55 in the second video memory.

[0076] Step 1132: Create the second mapping relationship based on the address space of the user process and the command address of the second video memory.

[0077] For example, if the address of the first command in the address space of the user process is 0x77, then the second mapping relationship can represent the mapping relationship between the address 0x77 of the first command in the address space of the user process and the address 0x55 of the first command in the second video memory.

[0078] In the embodiments of the present application, create the command address of the second video memory based on the command address of the first video memory; create the second mapping relationship based on the address space of the user process and the command address of the second video memory. In this way, the created second mapping relationship can, in the case of process migration, not affect the mapping between the command in the address space and the second video memory, that is, the process directly uses the command stored in the address space of the user process, that is, the command can be mapped to the second video memory.

[0079] Figure 2A For a schematic diagram of a process migration scenario provided by the embodiments of the present application, as Figure 2A shown, the schematic diagram includes Instance 21, Instance 22, and Instance 23, where

[0080] Instance 21 is an instance obtained by splitting a graphics processing unit with a computing power of 1 gigaflop, i.e., 1 billion floating-point operations per second (Giga Floating-point Operations Per Second, GFLOPS), and a video memory of 2 GB, that is, an instance obtained by multi-instance splitting. This MIG allows a single physical GPU to be divided into multiple isolated GPU instances at the hardware level. Among them, the computing power of 1 g means that this instance can perform 1 billion floating-point operations per second. The 2 GB video memory is the memory capacity used by the GPU to store graphic data and temporary data required for performing operations. The size of the video memory directly affects the ability of the GPU to process complex graphics and large data sets.

[0081] Instance 22 is an instance with a computing power of 1 g and a video memory of 5 GB.

[0082] Instance 23 is an instance with a computing power of 4 g and a video memory of 30 GB.

[0083] During the implementation process, as Figure 2A shown, before process migration, Process 1 runs in Instance 21, Process 2 runs in Instance 22, and Process 3 runs in Instance 23; after process migration, Process 1, Process 2, and Process 3 all run in Instance 23. In this way, the resources of Instance 23 can be fully utilized, and Instances 21 and 23 can be freed up to provide to other processes.

[0084] Figure 2B This is a schematic diagram of a process migration process provided by an embodiment of the present application. As Figure 2B shown, this schematic diagram includes a user mode 24 and a kernel mode 25, where

[0085] In the user mode 24, a video memory snapshot file is saved in the disk.

[0086] The application running in the user mode is correspondingly set to write to a command buffer area (Command buff), that is, the address space of the user process.

[0087] In kernel mode 25, a kernel-mode virtualization module is set up to block the kernel from emitting instructions when process migration needs to be performed, so as to block the process. Here, blocking the kernel from emitting instructions means that the instruction emission at the kernel level is blocked; when the process migration is completed, the blocking of the process is lifted, and the kernel is submitted to the new GPU split instance for continued execution. Using this kernel-mode virtualization module, a video memory snapshot file is obtained from the disk, and the process data is migrated to the target video memory (the second video memory) based on the storage address of the process data in the video memory snapshot file. In this process, the kernel-mode driver of the graphics processor (the graphics card driver program) can be used to apply for the storage address based on the video memory snapshot to save the process data to the second video memory; and using this kernel-mode virtualization module, a first mapping relationship (command buff address mapping) is generated based on the address space of the user process and the command address of the first video memory, and a second mapping relationship is established using this first mapping relationship and the second video memory, so as to remap the video memory area of the command cache area of the new GPU split instance to the process address space according to the source address mapping relationship (the second mapping relationship).

[0088] Figure 2C This is a schematic diagram of the implementation process of process migration provided by an embodiment of the present application. As Figure 2C shown, it can be implemented through the following steps:

[0089] Step S201, intercept the kernel emission instruction;

[0090] Here, the kernel emission instruction refers to an instruction at the kernel level.

[0091] Intercepting the kernel emission instruction refers to the behavior of blocking the upcoming instruction or system call at the operating system kernel level.

[0092] Step S202, determine whether the process is to be migrated;

[0093] If it is determined that the process is to be migrated, step S203 is executed; if it is determined that the process is not to be migrated, step S214 is executed.

[0094] Step S203, block the current kernel emission instruction and block the current process at the same time;

[0095] Here, when the instruction emission at the kernel level is blocked, the current process can be blocked.

[0096] Step S204, generate a video memory snapshot of the process;

[0097] Here, the video memory snapshot includes process data and corresponding address relationship information, that is, the process data and the storage address of the process data.

[0098] Step S205: Save the video memory snapshot to the disk of the node;

[0099] Here, since the amount of process data may be large, the video memory snapshot can be saved to a disk on the node that can store a large amount of data. Among them, the node can be a hardware device used to run the process, such as a computer, a server, or a cloud service device, etc.

[0100] Step S206: Save the command cache mapping relationship of the process;

[0101] Here, the command cache (command buff) mapping relationship refers to the mapping relationship from the video memory to the user process address space. Using this mapping relationship, the commands written to the user process address space can be mapped to the video memory.

[0102] Step S207: Completely exit the process from the GPU;

[0103] In the implementation process, it can be set that the process exits from the GPU, that is, terminate the execution of the process on the GPU while ensuring that the process no longer occupies GPU resources.

[0104] Step S208: Create a new environment for the process on the migrated target GPU split instance;

[0105] Here, creating a new environment for the process on the migrated target GPU split instance means performing initialization settings on the split instance, including operations such as installing software for initializing the environment and configuring environment variables.

[0106] Step S209: Read the process snapshot saved on the disk;

[0107] In the implementation process, as Figure 2B shown, the kernel-mode virtualization module can be used to read the process snapshot saved on the disk.

[0108] Step S210: The kernel-mode virtualization module uses the function in the GPU kernel-mode driver to apply for video memory with a fixed address, and applies for video memory consistent with the address information saved in the snapshot;

[0109] In the implementation process, the kernel-mode virtualization module uses the function in the GPU kernel-mode driver to apply for a video memory address based on the storage address in the video memory snapshot, that is, it can realize applying for a video memory address consistent with the address information saved in the snapshot.

[0110] Step S211: Save the data in the snapshot to the new video memory according to the corresponding address;

[0111] During the implementation process, as Figure 2B shown, the kernel-mode virtualization module is used to migrate the process data to the new video memory (the second video memory) based on the video memory address that is the same as the address information saved in the snapshot applied for.

[0112] Step S212: Remap the command buffer video memory area of the new GPU split instance to the process address space according to the source address mapping relationship;

[0113] During the implementation process, the kernel-mode virtualization module is used to generate a first mapping relationship (command buff address mapping) based on the address space of the user process and the command address of the first video memory, and then a second mapping relationship is established using the first mapping relationship and the second video memory, so as to realize remapping the video memory area of the command buffer area of the new GPU split instance to the process address space according to the source address mapping relationship (the second mapping relationship).

[0114] Step S213: Release the blocking of the process, and submit the kernel to the new GPU split instance to continue execution;

[0115] During the implementation process, by releasing the blocking of the process, realizing the submission of the kernel command to the new GPU split instance to continue execution, the migration of the process can be completed.

[0116] Step S214: Normally submit the kernel for execution.

[0117] In the embodiment of the present application, when the process is migrated, the kernel-mode virtualization module records the current command buff correspondence relationship, then takes a snapshot of the video memory of the current process in the GPU card and saves it to the hard disk, then suspends the process until before the next kernel execution, and makes the process exit the GPU card. Then, the GPU card is re-segmented in space and time, and the tasks are redeployed to the newly allocated instances, and the video memory snapshot is restored from the previously saved file. And create the same command buff address mapping as before, and then continue to execute the process. In this way, by using the function of the kernel virtualization module to provide the function of saving / restoring the process command buffer (command buff) mapping relationship and the video memory snapshot, the function of process migration is realized without exiting the process. It is ensured that the process can be deployed to the re-segmented space-time resource instances without restarting the process.

[0118] Currently, in the scenario of GPU computing, in order to improve the utilization rate of GPU resources, the GPU sharing technology is adopted, that is, multiple tasks are allowed to use a single GPU card at the same time. In order to avoid resource competition and meet the target throughput of each task, the GPU virtualization technology is usually adopted.

[0119] GPU virtualization solutions are mainly divided into two categories: time-division multiplexing and space-division multiplexing. Among them,

[0120] Time-division multiplexing means slicing the GPU resources in time. Only one process is allowed to use the GPU resources within a time slice. Multiple processes switch between time slices and take turns using the GPU to achieve the purpose of sharing the GPU resources. By sharing the GPU through time slice switching, only one process can use the GPU at the same moment, and the Streaming Multiprocessor (SM) units in the GPU cannot be fully utilized. For example, the minimum time slice for Load 1 to meet the target throughput is 40%. Load 1 monopolizes all SM resources within the time slice, but in fact, only 4 / 7 of the SM resources are needed to meet the target throughput. Then, 3 / 7 of the SM resources are wasted.

[0121] Space-division multiplexing means dividing the GPU SM resources. Each process can only use a part of the resources to achieve the purpose of multiple processes using the GPU simultaneously. Space-division multiplexing includes hardware partitioning: setting dedicated circuits and chips on the GPU hardware to partition the SM, which is provided natively by the manufacturer. Based on hardware partitioning, such as MIG. There is a problem of coarse partitioning granularity. MIG supports a maximum of 7 instances for partitioning, and the partitioning granularity is an integer multiple of 1 / 7 (1 / 8 of the video memory), resulting in time slice waste for loads that meet the target throughput. For example, assume that the minimum MIG instance for Load 1 to meet the target throughput is 3 / 7. Only 62% of the time slice is needed within this MIG instance to meet the throughput target. Then, 38% of the time slice is wasted.

[0122] In summary, current GPU virtualization cannot balance the utilization of time slices and SM, and cannot balance aspects such as fault isolation and performance at the same time. It is impossible to maximize the GPU utilization while ensuring the target throughput of each load.

[0123] Based on the above, loads (processes) may have different characteristics. For example, some loads may have relatively large kernels for each one, but the frequency of kernel calls is relatively low. As the number of available SMs decreases, the performance will drop sharply. Therefore, a larger number of SMs are needed to ensure its target throughput. However, if time-division multiplexing is used, only a relatively small number of time slices may be needed to ensure its target throughput. On the contrary, some loads are basically composed of small kernels, but the call frequency is relatively high. Such loads only need a relatively small number of SMs to ensure their target throughput. However, if time-division multiplexing virtualization is used, due to the waste of SMs, a relatively large number of time slices may be needed to ensure its target throughput.

[0124] In this way, some loads are suitable for slicing using time-division multiplexing virtualization, and some loads are suitable for slicing using space-division multiplexing virtualization. Figure 3AThis application embodiment provides a spatio-temporal hybrid virtualization system. As Figure 3A shown, the spatio-temporal hybrid virtualization system includes: an offline load throughput collector (profiler) 31, a spatio-temporal orchestration module (Spatio-temporal orchestrator, orchestrator) 32, a multi-instance GPU (MIG) virtualization module 33, a multi-processor sharing (MPS) virtualization module 34, and a kernel-based virtual GPU (kvGPU) virtualization module 35. Among them,

[0125] The offline load throughput collector 31 is used to measure the maximum throughput that can be achieved for all loads to be deployed by the user on each MIG partition, and save the obtained information as a file.

[0126] The spatio-temporal orchestration module 32 is used to compile the best spatio-temporal partition configuration (MIG + MPS / kvGPU) that meets the target throughput according to the maximum throughput of each load measured by the Profiler on each MIG partition, so as to maximize resource utilization.

[0127] The multi-instance GPU (MIG) virtualization module 33 selects MIG as the basis for system virtualization. Because MIG partitions the SM and L2 cache (Cache) from the hardware, ensuring strict isolation and negligible performance loss.

[0128] The multi-processor sharing (MPS) virtualization module 34 selects MPS as fine-grained multi-processor sharing virtualization, which is superimposed on top of MIG instances to provide finer-grained partitioning for loads.

[0129] The kernel-based virtual GPU (kvGPU) virtualization module 35 selects a self-developed kernel-state virtualization system, which is superimposed on top of MIG instances to provide finer-grained partitioning for loads. When a process migrates, it provides the function of saving / restoring the mapping relationship of the process command buffer (commandbuff) and the video memory snapshot, achieving the function of migrating without exiting the process.

[0130] As Figure 3A shown in the spatio-temporal hybrid virtualization system, before deploying a task, first use the offline load throughput collector 31 to pre-collect the performance metrics of the task, obtain the maximum throughput of the load on each (Multi-Instance GPU, MIG) partition instance, and deduce the time slice quota and space quota required for the load to achieve the target throughput on each MIG instance. Then use the spatio-temporal orchestration module 32 to use the orchestration algorithm, comprehensively consider the three virtualization technologies of MIG, MPS, and kernel-based virtual GPU (kvGPU) virtualization, and finally determine the spatio-temporal optimal partition hybrid virtualization system, achieving the maximization of GPU utilization and saving GPU resources while ensuring the target throughput of the load.

[0131] In the implementation process, the time-division and space-division multiplexing are effectively combined through the time-space-division scheduling method, achieving the highest resource utilization rate while ensuring the throughput of each load target. It is not required that all tasks be completed and submitted before searching for the global optimal solution and deploying it to the final time-space-division slicing instance for execution. Instead, after each task is submitted, the time-space-division slicing of the local optimal solution can be obtained through the scoring algorithm and then directly deployed for execution. When subsequent tasks are submitted, the global optimal solution changes.

[0132] For the first process, each tGPU and sGPU can be scored through the following formula (1):

[0133] Score = p (or t) × s (1);

[0134] Among them, s represents the MIG SM resource quota, p represents the MPS SM resource quota, and t represents the time slice quota. The smaller the score, the less resources are used. The one with the smallest score is selected as the slicing instance, and the task is deployed on it for running.

[0135] In the embodiment of the present application, before task deployment, the offline performance collector measures and collects the throughput metrics of the load, and calculates the minimum space-division and time-division resources required for each load to reach the target throughput through the collected data. Through the time-space-division scheduling algorithm, the optimal time-space hybrid slicing ratio that meets the target throughput is calculated, and the hybrid virtualization instance is created according to the slicing ratio to achieve the maximum GPU resource utilization rate. The kernel-state virtualization module is used as the time-division resource management base, fully considering the accuracy of time-division resource isolation, and ensuring the correctness of the results to the greatest extent. The adaptive time-space slicing adjustment does not wait for all tasks to be submitted to calculate the optimal slicing. Instead, for the first arriving task, a current optimal time-space slicing instance is allocated for deployment and running according to the self-created scoring algorithm. For subsequent arriving tasks, they are re-planned and arranged in combination with the tasks that have already been run to ensure that the global optimal division is always achieved in real time. For the real-time adjustment of the division, it is ensured that the tasks can continue to run on the re-divided slicing instances without restarting.

[0136] The embodiment of the present application provides a method for determining the second video memory for time-space multiplexing, as Figure 3B shown, which can be implemented through the following steps:

[0137] Step S310: Determine the target instance among all image processor instances based on the maximum video memory used by the target process and the target throughput;

[0138] Here, the maximum video memory refers to the maximum amount of graphics processing unit (GPU) video memory that the target process can use during operation; the actual video memory used by the process depends on the requirements of the application. For example, some graphics-intensive or compute-intensive applications may require a large amount of video memory to store and process data.

[0139] Throughput refers to the amount of data or the number of tasks that the target process can process per unit time. This metric reflects the processing capacity and efficiency of the process.

[0140] The image processor instance, i.e., the Multi-Instance GPU (MIG). MIG allows a single GPU to be safely partitioned into multiple independent GPU instances that can run different workloads in parallel to achieve optimal GPU utilization. When configuring MIG, different numbers and sizes of GPU instances can be flexibly selected according to the requirements of the workload. For example, on an A100 GPU, up to 7 independent MIG instances can be created, and the memory and computing capabilities of each instance can be allocated as needed. Among them, each image processing instance corresponding to the target process can determine the peak throughput.

[0141] During implementation, for the target process, the peak throughput corresponding to each image processor instance can be determined respectively among all the image processor instances partitioned by the image processor. For example, for the target process, the peak throughputs of 7 image processor instances can be determined to be 1.53 requests per second (rps), 2.87 rps, 4.19 rps, 4.83 rps, 7.18 rps, and 7.32 rps respectively. The target throughput is preset based on the actual processing requirements. When the target throughput is determined to be 2 rps, it can be determined that the image processing instance with a throughput of 1.53 rps cannot meet the processing requirements.

[0142] During implementation, each image processing instance corresponds to a video memory, that is, the video memory capacity of each image processing instance. For example, when the video memories of 6 image processor instances are 5 GB, 10 GB, 10 GB, 20 GB, 20 GB, and 40 GB respectively, and the maximum video memory required by the target process is 10 GB, the image processing instance with a video memory capacity of 5 GB cannot meet the processing requirements.

[0143] Step S320: When it is determined to use the target instance for time-domain multiplexing to execute the target process, obtain the minimum time slice quota of each target instance in the first splitting method;

[0144] Here, time-domain multiplexing in time-division multiplexing means slicing the image processing instance resources in time. Only one process is allowed to use the image processing instance resources within a time slice, and multiple processes switch between time slices to take turns using the image processing instance, achieving the purpose of sharing the image processing instance resources.

[0145] The first splitting method is to use the target instance for time-domain multiplexing to execute the target process, that is, perform time-domain multiplexing on the divided image processor instances. Since the image processor instances are spatially divided, it is a splitting method that superimposes time domain and spatial domain.

[0146] tGPU(s, t, m) represents the splitting of GPU resources using kernel-mode virtualization plus MIG. Where s represents the MIG SM resource quota, t represents the time slice quota, and m represents the MIG video memory quota.

[0147] During the implementation process, determining to perform time-domain multiplexing on the basis of MIG partitioning can obtain the minimum time quota for each target instance in this first splitting method, that is, the minimum time quota t for each target instance.

[0148] Step S330, in the case of determining to use the target instance for multi-process service to execute the target process, obtain the minimum resource quota for each target instance in the second splitting method;

[0149] Here, multi-process service (MPS) at the software level combines multiple processes into one context. Utilizing the characteristic that multiple streams within the same context can concurrently use the SM, it achieves the purpose of multi-process concurrent use of the SM. Among them, context represents the environmental information and context state during program execution, and is the basis for realizing multi-process concurrent execution and resource sharing.

[0150] The second splitting method is to use the target instance plus multi-process service to execute the target process, that is, use multi-process service on the divided image processor instances. Since the image processor instances are spatially divided, it is a splitting method that superimposes time domain and time domain.

[0151] sGPU(s, p, m) represents the splitting of GPU resources using MPS plus MIG. Where s represents the MIG SM resource quota, p represents the MPS SM resource quota, and m represents the MIG video memory quota.

[0152] During the implementation process, determining to use multi-process service on the basis of MIG partitioning can obtain the minimum resource quota p for each target instance in this second splitting method.

[0153] Step S340: Based on the minimum time slice quota and the minimum resource quota of each of the target instances, determine to migrate the target process to the second video memory.

[0154] During the implementation process, the minimum partitioning of tGPU(s, t, m) and sGPU(s, p, m) corresponding to all the target instances where the target process meets the target throughput (Thrtar) can be obtained.

[0155] First, for each mig instance, determine whether to use tGPU or sGPU by determining the minimum time slice quota and the minimum resource quota through the previous comparison.

[0156] During the implementation process, for the first process, each tGPU and sGPU can be scored through the following formula (1):

[0157] Score = p (or t) × s (1);

[0158] Where s represents the MIG SM resource quota, p represents the MPS SM resource quota, and t represents the time slice quota. The smaller the score, the less resources are used. The one with the smallest score is selected as the partitioning instance, and the task is deployed to run on it.

[0159] For the subsequent arriving processes, the preset orchestration method and the tasks that have been deployed and run before can be used to perform unified orchestration together to obtain the optimal partitioning.

[0160] Based on the latest obtained partitioning, determine to migrate the target process to the second video memory. This target process is a task that has been running, and the functions provided by kernel virtualization can be used, that is, execute the above steps S110 to S130 to achieve migration without restarting the process.

[0161] In the embodiments of the present application, the first partitioning method of time division plus space division is used to determine the minimum time slice quota of each target instance; the second partitioning method of time division plus time division is used to determine the minimum resource quota of each target instance; finally, based on the minimum time slice quota and the minimum resource quota of each target instance, determine to migrate the target process to the second video memory. In this way, the time-space division hybrid virtualization system can superimpose time division slices and space division slices according to the characteristics of each process, and simultaneously determine the optimal time-space division partitioning ratio, so as to pack more loads in a single GPU card, while ensuring the target throughput of each load, achieving the purpose of maximizing the utilization rate of GPU resources and reducing the usage cost. The combination of kernel virtualization and space division isolation technology is used to achieve complete isolation of resources, and at the same time provide high-quality performance and isolation accuracy.

[0162] In some embodiments, the above step S320, "When it is determined to use the target instance to perform the target process in time-domain multiplexing, obtain the minimum time slice quota for each target instance", can be implemented through the following process:

[0163] Based on the target throughput and the peak throughput on each target instance, determine the minimum time slice quota for the target process to run on each target instance.

[0164] During implementation, an offline load throughput collector 31 as shown in Figure 3A can be used to collect the full-load throughput (peak throughput) of the load (process) on each MIG configuration. tThr(s, t, m) represents the throughput corresponding to the tGPU(s, t, m) split of the load, and sThr(s, p, m) represents the throughput corresponding to the sGPU(s, t, m) split of the load.

[0165] That is, obtain tThr(s, 100%, m) = sThr(s, 100%, m), and at the same time record the maximum memory m used by the load max .

[0166] Deduction of the minimum tGPU(s, t, m) that meets the target throughput: Due to the precise isolation of kernel-level virtualization (error < 5%), the following formula (2) can be obtained:

[0167] tThr(s,t,m) = tThr(s, 100%, m) × t(2);

[0168] From this, the following formula (3) can be used to calculate the minimum time slice quota that meets the target throughput:

[0169]

[0170] where tThr(s, 100%, m) represents the peak throughput, Thr tar represents the target throughput, and T tar represents the minimum time slice quota that meets the target throughput. The target throughput is a throughput value preset based on the target process.

[0171] For example, set the target throughput of the target process to 2 rps. The peak throughputs of the target process measured on 6 MIG configurations and the whole card of A100 GPU are 1.53 rps, 2.87 rps, 4.19 rps, 4.83 rps, 7.18 rps, and 7.32 rps respectively.

[0172] Based on the above formula (3), the tGPU that meets the target throughput is obtained as follows:

[0173] tGPU (2g, 70%, 10gb), tGPU (3g, 48%, 20gb), tGPU (4g, 42%, 20gb), tGPU (7g, 28%, 40gb), tGPU (entire, 27%, 40gb).

[0174] Figure 3C A time-sharing (tGPU) schematic diagram for meeting the target throughput provided by an embodiment of the present application is shown in Figure 3C As shown, the horizontal axis represents peak throughput, target throughput, and tGPU, and the vertical axis represents the computing power and video memory capacity of GPU instances. As Figure 3C shown, it can be seen the peak throughput, target throughput, and tGPU corresponding to GPU instances with different computing powers and video memory capacities respectively. For example, for an instance with a computing power of 2g and a video memory capacity of 10gb, the corresponding peak throughput is 2.87rps, the target throughput is 2rps, and then t is 70%.

[0175] In the embodiment of the present application, based on the target throughput and the peak throughput on each target instance, the minimum time slice quota for the target process to run on each target instance can be determined.

[0176] In some embodiments, the above step S330 "when it is determined to use the target instance to execute the target process for multi-process service, obtain the minimum resource quota for each target instance" can be implemented through the following steps:

[0177] Step 331, determine a second instance whose computing power is less than that of the target instance;

[0178] Here, the second instance may not be the target instance, that is, the second instance may be an instance that does not meet the target throughput. The second instance may also be an instance that meets the target throughput and whose computing power is less than that of the target instance.

[0179] Step 332, determine the minimum resource quota for the target process on each target instance based on the first peak throughput of each target instance and the second peak throughput of the second instance.

[0180] The calculation of the minimum sGPU (s, p, m) that meets the target throughput is as follows:

[0181] Figure 3D A schematic diagram of peak throughput and instances provided by an embodiment of the present application is shown in Figure 3D As shown, the horizontal axis represents each instance, and the vertical axis represents the peak throughput corresponding to each instance respectively.

[0182] Assume that the throughput of the target process linearly increases with the SM between the minimum MIG segmentation that meets the target throughput and the previous MIG segmentation, as Figure 3DAs shown in the figure. The linearly increasing trend is expressed by the following formula (4):

[0183] y = kx + b (4);

[0184] where y represents the throughput as shown in Figure 3D the figure, x represents the computing power of the instance, and k and b are two linear parameters respectively.

[0185] The minimum MIG partition that meets the target throughput is denoted as sGPU(s i , 100%, m), and the throughput corresponding to the previous MIG partition denoted as sGPU(s i-1 , 100%, m) is denoted as sThr(s i , 100%, m). Substituting sThr(s i-1 , 100%, m) into the above formula (4) gives the following system of equations (5):

[0186]

[0187] Solving this system of equations gives the following formulas (6) and (7):

[0188]

[0189] From this, the minimum MPS SM quota p tar that meets the target throughput can be obtained using the following formula (8) as:

[0190]

[0191] where Thr tar represents the target throughput, and p tar represents the minimum MPS SM quota that meets the target throughput, i.e., the minimum resource quota.

[0192] The minimum sGPU that meets the target throughput can be obtained as sGPU(s, p tar , m).

[0193] As Figure 3D shown in the figure, the corresponding video memory capacity can be determined to be 1.35 GB using the 1g, 10gb instance (the second instance) and the 2g, 10gb instance (the target instance) when meeting the target throughput of 2 rps.

[0194] For example, if the target throughput of the target process is set to 2 rps, the peak throughputs of the target process measured on 6 MIG configurations and the whole card of the A100 GPU are 1.53 rps, 2.87 rps, 4.19 rps, 4.83 rps, 7.18 rps, and 7.32 rps respectively.

[0195] Based on the above formulas (6), (7), and (8), the sGPUs that meet the target throughput are as follows:

[0196] sGPU(2g, 68%, 10gb), sGPU(3g, 45%, 20gb), sGPU(4g, 34%, 20gb), sGPU(7g, 19%, 40gb), sGPU(entire, 17%, 40gb).

[0197] For example, sGPU(3g, 45%, 20gb) means that this virtual resource is a 3g MIG instance and is restricted to providing only 45% of the SM resources of this instance, and 20gb of video memory.

[0198] In the embodiments of the present application, first, a second instance with computing power less than the target instance is determined; then, the minimum resource quota of the target process on each target instance can be determined based on the first peak throughput of each target instance and the second peak throughput of the second instance.

[0199] The embodiments of the present application provide a space-time hybrid bin-packing scheduling algorithm:

[0200] The preconditions are as follows: tGPU and sGPU are in an opposing relationship (only one type of partitioning can exist on one MIG). Without resource conflicts, the performance of space-division Mps should be higher than that of time-division multiplexing, so sGPU scheduling is preferred first.

[0201] Figure 4A For a schematic diagram of space-time hybrid scheduling provided by the embodiments of the present application, as Figure 4A shown, the hybrid scheduling at this moment can be achieved through the following steps:

[0202] Step 1: For each workload, that is, each process, compare the corresponding tGPU and sGPU respectively. If the corresponding t < p, it is considered that time-slicing is better than SM partitioning under this MIG instance, and the corresponding tGPU is selected; otherwise, sGPU is selected.

[0203] Step 2: Each workload selects the partitioning (sGPU or tGPU) with the smallest resource occupancy.

[0204] Step 3: Merge sGPUs with the same configuration (the merging rule is that p added together is less than or equal to 100%, and m max added together is less than or equal to m). Among them, m max represents the maximum memory used by the workload.

[0205] Step 4: Merge small sGPUs into large sGPUs (using the same merging rules as above). For sGPUs with the same configuration, start merging from the one with the smallest p. If the corresponding large sGPU configuration of the small sGPU to be merged was selected as tGPU in Step 1, it will not participate in the merging.

[0206] Step 5: After all mergings are completed, attempt to fill the remaining load into the merged sGPUs (convert all to sGPUs during filling).

[0207] Step 6: For the remaining unmerged load after filling, perform tGPU merging (this type of load is not suitable for sGPU merging). The merging steps are the same as those for sGPU merging above (the merging rule is that the sum of t is less than or equal to 100%, and the sum of m max is less than or equal to m, and no longer judge whether it is suitable for tGPU merging).

[0208] Step 7: Perform GPU card bin-packing integration on all merged sGPUs / tGPUs. Among them, GPU card bin-packing integration means distributing multiple MIG instances to 1 GPU card as much as possible (because multiple MIG instances can be created on 1 GPU card) to achieve the purpose of saving GPU card usage.

[0209] For example, the environment is an A100 card with 7 types of MIG partitions (including the whole card): {1g.5gb, 1g.10gb, 2g.10gb, 3g.20gb, 4g,20gb, 7g.40gb, entire}; assume there are 5 loads denoted as w1, w2, w3, w4, w5 respectively. The offline load throughput collector obtains the maximum video memory of each load as 3G, 4G, 3G, 12G, 1G. All sGPUs / tGPUs that meet the throughput are as shown in Table 1 below:

[0210] Table 1

[0211] Load Segmentation method 1g 2g 3g w1 Spatial division sGPU(1g, 45%, 10gb) sGPU(2g, 23%, 10gb) sGPU(3g, 15%, 20gb) Time division tGPU(1g, 45%, 10gb) tGPU(2g, 35%, 10gb) tGPU(3g, 30%, 20gb) w2 Spatial division sGPU(1g, 40%, 10gb) sGPU(2g, 20%, 10gb) sGPU(3g, 13%, 20gb) Time division tGPU(1g, 40%, 10gb) tGPU(2g, 29%, 10gb) tGPU(3g, 25%, 20gb) w3 Spatial division sGPU(2g, 30%, 10gb) sGPU(3g, 20%, 20gb) Time division tGPU(2g, 42%, 10gb) tGPU(3g, 32%, 20gb) w4 Spatial division sGPU(3g, 60%, 20gb) Time division tGPU(3g, 47%, 20gb) w5 Spatial division sGPU(2g, 73%, 10gb) sGPU(3g, 47%, 20gb) Time division tGPU(2g, 59%, 10gb) tGPU(3g, 41%, 20gb) Load Segmentation method 4g 7g entire GPU w1 Spatial division sGPU(4g, 11%, 20gb) sGPU(7g, 6%, 40gb) sGPU(entire, 5%, 40gb) Time division tGPU(4g, 27%, 20gb) tGPU(7g, 22%, 40gb) tGPU(entire, 19%, 40gb) w2 Spatial division sGPU(4g, 10%, 20gb) sGPU(7g, 6%, 40gb) sGPU(entire, 5%, 40gb) Time division tGPU(4g, 19%, 20gb) tGPU(7g, 17%, 40gb) tGPU(entire, 16%, 40gb) w3 Spatial division sGPU(4g, 15%, 20gb) sGPU(7g, 9%, 40gb) sGPU(entire, 8%, 40gb) Time division tGPU(4g, 25%, 20gb) tGPU(7g, 15%, 40gb) tGPU(entire, 10%, 40gb) w4 Spatial division sGPU(4g, 45%, 20gb) sGPU(7g, 26%, 40gb) sGPU(entire, 30%, 40gb) Time division tGPU(4g, 37%, 20gb) tGPU(7g, 29%, 40gb) tGPU(entire, 25%, 40gb) w5 Spatial division sGPU(4g, 35%, 20gb) sGPU(7g, 20%, 40gb) sGPU(entire, 18%, 40gb) Time division tGPU(4g, 34%, 20gb) tGPU(7g, 27%, 40gb) tGPU(entire, 15%, 40gb)

[0212] The merging steps are as shown in Table 2 below:

[0213] Table 2

[0214]

[0215] Figure 4B A schematic diagram of a merging result provided by an embodiment of the present application is as Figure 4B shown. This schematic diagram is the merging result obtained based on Table 2, including w1, w2, and w3 running in space-division on MIG0 of GPU0; w5 and w4 running in space-division on MIG1; and MIG2 is in an idle state.

[0216] In some embodiments, the above step S340, "determine to migrate the target process to the second video memory based on the minimum time slice quota and the minimum resource quota of each of the target instances" can be implemented through the following steps:

[0217] Step 341: Based on the minimum time slice quota and the minimum resource quota of each of the target instances, determine the slicing method with a smaller quota as the target slicing method among the first slicing method and the second slicing method, so as to merge instances with the same slicing method as other instances executing other processes based on the target slicing method of each of the target instances;

[0218] Here, based on the minimum time slice quota and the minimum resource quota of each target instance corresponding to the target process, determine a slicing method with a smaller quota. For example, Figure 4A In the steps shown, for the instance sGPU with a computing power of 1g occupied by Load (target process) 1, the resources it occupies are less than those of tGPU, so sGPU is determined as the target slicing method; for the instance tGPU with a computing power of 4g occupied by Load 2, the resources it occupies are less than those of sGPU, so tGPU is determined as the target slicing method.

[0219] During the implementation process, the merging of instances corresponding to the execution of at least one process can be completed according to Figure 4A the steps 1 to 7 shown.

[0220] Step 342: Determine the second video memory based on the merged instances, so as to migrate the target process to the second video memory.

[0221] For example, Figure 4B In the schematic diagram of the merged instances shown, if Load (target process) 1 is migrated to the MIG1 instance of GPU0, then the video memory corresponding to the MIG1 instance can be determined as the second video memory, and executing the above steps S110 to S130 can achieve migrating the target process to the second video memory.

[0222] Figure 4C This is a schematic diagram of resource utilization for a different usage scenario provided by the embodiments of the present application. As Figure 4C shown, the horizontal axis of this schematic diagram is different GPU usage scenarios, and the vertical axis is the minimum number of GPUs occupied by each scenario to meet the target throughput of all loads. Among them,

[0223] MIG: 1.29 pieces; MIG plus MPS: 0.86 pieces; MIG plus kernel virtualization (time division): 0.86 pieces; MIG plus MPS plus kernel virtualization: 0.71 pieces.

[0224] As can be seen from the above data, the GPU usage solution of MIG + MPS + kernel-mode virtualization (time-sharing) provided by the embodiments of the present application occupies the smallest number of GPUs, which is 45% less than the MIG solution and 17% less than MIG + MPS and MIG + kernel-mode virtualization.

[0225] In the embodiments of the present application, first, the segmentation method with a smaller quota is determined as the target segmentation method, and then the segmentation methods of each target instance are merged with other instances executing other processes in the same segmentation method; finally, the second video memory is determined based on the merged instances, so as to migrate the target process to the second video memory. In this way, it is possible to calculate the optimal spatio-temporal hybrid segmentation ratio that meets the target throughput, and create hybrid virtualization instances according to the segmentation ratio, achieving the maximum GPU resource utilization.

[0226] Based on the foregoing embodiments, the embodiments of the present application provide a process migration device. The device includes each module included, each module includes each sub-module, and each sub-module includes units, which can be implemented by a processor in an electronic device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0227] Figure 5 It is a schematic structural diagram of the process migration device provided by the embodiments of the present application. As Figure 5 shown, the device 500 includes:

[0228] A first acquisition module 510, configured to acquire a video memory snapshot and a second mapping relationship of a target process in the kernel mode, where the video memory snapshot includes the process data of the target process and the storage address of the process data, and the second mapping relationship is created based on a first mapping relationship and a second video memory, and the first mapping relationship is a mapping relationship between a first video memory and the address space of the target process;

[0229] A migration module 520, configured to migrate the process data to the second video memory based on the storage address of the process data;

[0230] A mapping module 530, configured to map the command written from the address space to the second video memory by using the second mapping relationship.

[0231] In some embodiments, the process migration device further includes a second acquisition module, a generation module, and a storage module. Among them, the second acquisition module is configured to acquire the process data and the storage address of the process data when it is determined that the target process is to be migrated; the generation module is configured to generate a video memory snapshot based on the process data and the storage address of the process data; and the storage module is configured to store the video memory snapshot in the hard disk.

[0232] In some embodiments, the first acquisition module 510 includes a reading sub-module, a first acquisition sub-module, and a establishing sub-module. Among them, the reading sub-module is configured to read the video memory snapshot from the hard disk in the kernel mode; the first acquisition sub-module is configured to acquire the first mapping relationship; and the establishing sub-module is configured to establish the second mapping relationship based on the first mapping relationship and the second video memory.

[0233] In some embodiments, the first acquisition sub-module includes a first acquisition unit, a second acquisition unit, and a generation unit. Among them, the first acquisition unit is configured to acquire the address space corresponding to the command written to the user process address space in the user mode; the second acquisition unit is configured to acquire the command address of the command corresponding to the first video memory; and the generation unit is configured to generate the first mapping relationship based on the address space of the user process and the command address of the first video memory.

[0234] In some embodiments, the establishing sub-module includes a first creation unit and a second creation unit. Among them, the first creation unit is configured to create the command address of the second video memory based on the command address of the first video memory; and the second creation unit is configured to create the second mapping relationship based on the address space of the user process and the command address of the second video memory.

[0235] In some embodiments, the process migration device further includes a first determination module, a third acquisition module, a fourth acquisition module, and a second determination module. Among them, the first determination module is configured to determine a target instance among all the graphics processor instances based on the maximum video memory and the target throughput corresponding to the use of the target process; the third acquisition module is configured to acquire the minimum time slice quota of each target instance in the first segmentation mode when it is determined to use the target instance for time-domain multiplexing to execute the target process; the fourth acquisition module is configured to acquire the minimum resource quota of each target instance in the second segmentation mode when it is determined to use the target instance for multi-process service to execute the target process; and the second determination module is configured to determine to migrate the target process to the second video memory based on the minimum time slice quota and the minimum resource quota of each target instance.

[0236] In some embodiments, the third acquisition module is further configured to determine the minimum time slice quota for the target process to run on each target instance based on the target throughput and the peak throughput on each target instance.

[0237] In some embodiments, the second acquisition module includes a first determination sub-module and a second determination sub-module. The first determination sub-module is configured to determine a second instance whose computing power is less than that of the target instance. The second determination sub-module is configured to determine the minimum resource quota for the target process on each target instance based on the first peak throughput of each target instance and the second peak throughput of the second instance.

[0238] In some embodiments, the second determination module includes a third determination sub-module and a fourth determination sub-module. The third determination sub-module is configured to determine the slicing method with a smaller quota as the target slicing method from the first slicing method and the second slicing method based on the minimum time slice quota and the minimum resource quota of each target instance, so as to perform merging with other instances executing other processes in the same slicing method according to the target slicing method of each target instance. The fourth determination sub-module is configured to determine the second video memory based on the instances after completion of the merging, so as to migrate the target process to the second video memory.

[0239] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to those of the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0240] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the related art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a mobile phone, a tablet computer, a notebook computer, a desktop computer, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0241] Correspondingly, an embodiment of the present application provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the process migration method provided in the above embodiments are implemented.

[0242] Correspondingly, an embodiment of the present application provides an electronic device, Figure 6 which is a schematic diagram of a hardware entity of the electronic device provided by the embodiment of the present application. As Figure 6 shown, the hardware entity of the device 600 includes: a memory 601 and a processor 602. The memory 601 stores a computer program that can run on the processor 602. When the processor 602 executes the program, it implements the steps in the process migration method provided in the above embodiment.

[0243] The memory 601 is configured to store instructions and applications executable by the processor 602, and can also cache data to be processed or already processed by the processor 602 and each module in the electronic device 600 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0244] It should be noted here that: the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments, and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0245] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The sequence numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0246] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0247] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be electrical, mechanical, or other forms.

[0248] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0249] In addition, each functional unit in the embodiments of the present application can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0250] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs.

[0251] Alternatively, if the above-mentioned integrated units of the present application are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application essentially or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable an electronic device (which can be a mobile phone, a tablet computer, a notebook computer, a desktop computer, etc.) to execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, ROM, magnetic disks, or optical discs.

[0252] The methods disclosed in several method embodiments provided by this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0253] The features disclosed in several product embodiments provided by this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0254] The features disclosed in several method or device embodiments provided by this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0255] As mentioned above, the above are only the implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.

Claims

1. A process migration method, the method comprising: Obtaining a video memory snapshot and a second mapping relationship of a target process in the kernel mode, wherein the video memory snapshot includes the process data of the target process and the storage address of the process data, and the second mapping relationship is created based on a first mapping relationship and a second video memory, and the first mapping relationship is a mapping relationship between a first video memory and the address space of the target process; Migrating the process data to the second video memory based on the storage address of the process data; Mapping the command written to the address space from the second mapping relationship to the second video memory.

2. The method according to claim 1, the method further comprising: When it is determined that the target process is to be migrated, obtaining the process data and the storage address of the process data; Generating a video memory snapshot based on the process data and the storage address of the process data; Storing the video memory snapshot to a hard disk.

3. The method according to claim 2, the obtaining the video memory snapshot and the second mapping relationship of the target process in the kernel mode includes: Reading the video memory snapshot from the hard disk in the kernel mode; Obtaining the first mapping relationship; Establishing the second mapping relationship based on the first mapping relationship and the second video memory.

4. The method according to claim 3, the obtaining the first mapping relationship includes: Obtaining the address space corresponding to the command written to the user process in the user mode; Obtaining the command address of the command corresponding to the first video memory; Generating the first mapping relationship based on the address space of the user process and the command address of the first video memory.

5. The method according to claim 4, the establishing the second mapping relationship based on the first mapping relationship and the second video memory includes: Creating the command address of the second video memory based on the command address of the first video memory; Creating the second mapping relationship based on the address space of the user process and the command address of the second video memory.

6. The method according to any one of claims 1 to 5, the method further comprising: Determining a target instance among all graphics processor instances based on the maximum video memory and target throughput used by the target process; When it is determined to use the target instance for time-domain multiplexing to execute the target process, obtaining the minimum time slice quota of each target instance in the first segmentation method; When it is determined to use the target instance for multi-process service to execute the target process, obtaining the minimum resource quota of each target instance in the second segmentation method; Determining to migrate the target process to the second video memory based on the minimum time slice quota and the minimum resource quota of each target instance.

7. The method according to claim 6, the obtaining the minimum time slice quota of each target instance when it is determined to use the target instance for time-domain multiplexing to execute the target process includes: Determining the minimum time slice quota for the target process to run on each target instance based on the target throughput and the peak throughput on each target instance.

8. The method according to claim 6, wherein when it is determined to use the target instance to execute the target process for multi-process service, obtaining the minimum resource quota of each of the target instances includes: Determining a second instance with computing power less than that of the target instance; Determining the minimum resource quota of the target process on each of the target instances based on the first peak throughput of each of the target instances and the second peak throughput of the second instance.

9. The method according to claim 6, wherein determining to migrate the target process to the second video memory based on the minimum time slice quota and the minimum resource quota of each of the target instances includes: Based on the minimum time slice quota and the minimum resource quota of each of the target instances, determining the slicing method with a smaller quota as the target slicing method among the first slicing method and the second slicing method, so as to merge the instances with the same slicing method as other instances executing other processes based on the target slicing method of each of the target instances; Determining the second video memory based on the instances after the merging is completed, so as to migrate the target process to the second video memory.

10. A process migration device, the device includes: A first acquisition module, configured to acquire a video memory snapshot and a second mapping relationship of a target process in the kernel state, wherein the video memory snapshot includes the process data of the target process and the storage address of the process data, and the second mapping relationship is created based on a first mapping relationship and a second video memory, and the first mapping relationship is a mapping relationship between a first video memory and the address space of the target process; A migration module, configured to migrate the process data to the second video memory based on the storage address of the process data; A mapping module, configured to map the command written to the address space to the second video memory by using the second mapping relationship.