Methods for virtual machine migration

By recording and transmitting the status of the AI ​​accelerator during the virtual machine migration, the problem of AI application failure or interruption caused by virtual machine migration is solved, and the continuity and stability of AI tasks are achieved.

CN114721769BActive Publication Date: 2025-05-23BAIDU USA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111655441.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-01-06
Filing Date
2021-12-30
Publication Date
2025-05-23
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

The prior art cannot effectively protect and capture the status of AI accelerators during virtual machine migration, resulting in the possibility of failure or interruption of AI applications.

Method used

Ensure that the migrated virtual machine can successfully restart the AI ​​task by recording and stopping the execution of AI tasks, generating or selecting the AI ​​accelerator status associated with the virtual function, and transmitting it to the target host's hypervisor.

Benefits of technology

It effectively avoids the problem of failure or interruption of AI applications after migration to another host, ensuring the continuity and stability of AI tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114721769B_ABST
    Figure CN114721769B_ABST
Patent Text Reader

Abstract

Disclosed are methods and systems for migrating a virtual machine (VM) having resources of an artificial intelligence (AI) accelerator mapped to a virtual function of the VM. A driver for the AI ​​accelerator can generate a checkpoint of a VM process that calls the AI ​​accelerator, and the checkpoint can include a list and configuration of resources mapped to the AI ​​accelerator by the virtual function. The driver can also access the code, data, and memory of the AI ​​accelerator to generate a checkpoint of the AI ​​accelerator state. When the VM migrates to a new host, one or both of these checkpoint frames can be used to ensure that restoring the VM on a new host with appropriate AI accelerator resources can be successfully restored on the new host. One or two checkpoint frames can be captured based on an event that anticipates the need to migrate the VM.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to one or more artificial intelligence accelerators coupled to a host of virtual machines. More particularly, embodiments of the present disclosure relate to migrating virtual machines that use artificial intelligence accelerators. Background Art

[0002] AI models (also called "machine learning models") have recently become widely used as artificial intelligence (AI) technologies have been deployed in various fields, such as image classification, medical diagnosis, or autonomous driving. Similar to an executable image or binary image of a software application, an AI model can be trained to reason based on a set of attributes to be classified as features. The training of AI models may require a significant investment in collecting, organizing, and filtering data to produce an AI model that produces useful predictions. In addition, the predictions derived using AI models may contain personal sensitive data that users wish to protect.

[0003] Generating predictions from an AI model can be a computationally intensive process. To provide sufficient computing power for one or more users, one or more AI accelerators can be coupled to a host of one or more virtual machines. To provide sufficient computing power for computationally intensive tasks, such as training an AI model, the AI ​​accelerators can be organized into clusters and then into multiple groups, each of which can be assigned to a single virtual machine. For less intensive tasks, a single virtual machine can be assigned a single AI accelerator.

[0004] A virtual machine may need to be migrated to a different host for several well-known reasons. Prior art virtual machine migration does not protect the state of one or more AI accelerators during the migration process. An AI application that generates one or more artificial intelligence tasks, at least some of which are executed on an AI accelerator, may fail or interrupt after migration to another host. Failures may include failure to capture the configuration, memory contents, and computational state of the AI ​​accelerator and failure to capture the computational state of the AI ​​tasks within the virtual machine. Summary of the invention

[0005] In a first aspect, a method for migrating a source virtual machine (VM-S) is provided, wherein the source virtual machine is executing an application accessing a virtual function of an artificial intelligence (AI) accelerator, the method comprising:

[0006] In response to receiving a command to migrate the VM-S and virtual functions, and receiving a selection to checkpoint the VM-S and virtual functions used in performing the migration:

[0007] Record, and then stop, one or more AI tasks being executed by an application.

[0008] Generate or select the state of an AI accelerator associated with a virtual function, and

[0009] Transferring the checkpoints and states of the AI ​​accelerator to the hypervisor of the target host to generate a migrated target virtual machine (VM-T); and

[0010] In response to receiving notification that the target host has verified the checkpoint and AI state, and has generated and configured resources for generating VM-T, and has loaded the AI ​​accelerator on the target host with data from the AI ​​accelerator state: migrate VM-S and virtual functions to VM-T.

[0011] In a second aspect, a computer-readable medium programmed with executable instructions is provided, which, when executed by a processing system having at least one hardware processor communicatively coupled to an artificial intelligence (AI) processor, implements the operation of migrating a source virtual machine (VM-S) of an application executing a virtual function of an artificial intelligence (AI) accelerator of the system as described in the first aspect.

[0012] In a third aspect, a system is provided, comprising at least one hardware processor coupled to a memory programmed with instructions, which, when executed by the at least one hardware processor, enables the system to implement operations as described in the first aspect for migrating a source virtual machine (VM-S) that is executing an application that is accessing a virtual function of an artificial intelligence (AI) accelerator.

[0013] In a fourth aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the operation of migrating a source virtual machine (VM-S) that is executing an application that accesses a virtual function of an artificial intelligence (AI) accelerator as described in the first aspect.

[0014] According to the embodiments of the present invention, it is possible to avoid the failure or interruption of an AI application after it is migrated to another host. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Embodiments of the present disclosure are illustrated by way of example and not limitation in the figures of the accompanying drawings in which like references indicate similar elements.

[0016] Figure 1 is a block diagram illustrating a secure processing system that can migrate virtual machines with checkpoint authentication and / or artificial intelligence (AI) accelerator state verification according to one embodiment.

[0017] Figure 2A and 2B is a block diagram illustrating a secure computing environment between one or more hosts and one or more artificial intelligence accelerators according to one embodiment.

[0018] Figure 3 is a block diagram illustrating a host controlling a cluster of artificial intelligence accelerators according to an embodiment, each cluster having a virtual function for mapping resources of a group of AI accelerators within the cluster to a virtual machine, each artificial intelligence accelerator having secure resources and non-secure resources.

[0019] Figure 4A is a block diagram illustrating components of a data processing system with an artificial intelligence (AI) accelerator that implements a method of virtual machine migration with checkpoint authentication in a virtualized environment according to an embodiment.

[0020] Figure 4B is a block diagram illustrating components of a data processing system with an artificial intelligence (AI) accelerator that implements a method for virtual machine migration with AI accelerator state verification in a virtualized environment according to an embodiment.

[0021] Figure 5A A method for virtual machine migration of a data processing system with an AI accelerator with checkpoint authentication in a virtualized environment from the perspective of a hypervisor of a host of a source virtual machine to be migrated is shown according to an embodiment.

[0022] Figure 5B A method for virtual machine migration of a data processing system having an AI accelerator with AI accelerator status verification in a virtualized environment from the perspective of a hypervisor of a host of a source virtual machine to be migrated according to an embodiment is shown.

[0023] Figure 6 A method for generating a checkpoint used in a virtual machine migration method with checkpoint authentication in a virtualization environment from the perspective of a source hypervisor of a host of a virtual machine to be migrated according to an embodiment is shown.

[0024] Figure 7 A method for determining whether to migrate a virtual machine of a data processing system having an AI accelerator with checkpoint authentication in a virtualized environment from the perspective of a source hypervisor hosting the virtual machine to be migrated according to an embodiment is shown.

[0025] Figure 8 A method for migrating a virtual machine of a data processing system having an AI accelerator with checkpoint authentication in a virtualized environment from the perspective of a source hypervisor hosting the virtual machine to be migrated is shown according to an embodiment.

[0026] Fig. 9 A method of performing post-migration cleanup of a source host computing device after migrating a virtual machine having a data processing system with an AI accelerator with checkpoint authentication in a virtualized environment according to an embodiment is shown.

[0027] Fig.10 A method for migrating a virtual machine of a data processing system having an AI accelerator with checkpoint authentication and optional AI accelerator state verification in a virtualized environment from the perspective of a target hypervisor on a host that will host the migrated virtual machine is shown in accordance with some embodiments. DETAILED DESCRIPTION

[0028] Various embodiments and aspects of the present disclosure will be described with reference to the details discussed below, and the accompanying drawings will illustrate various embodiments. The following description and the accompanying drawings are illustrative of the present disclosure and should not be construed as limiting the present disclosure. Many specific details are described to provide a comprehensive understanding of various embodiments of the present disclosure. However, in some cases, in order to provide a brief discussion of embodiments of the present disclosure, known or conventional details are not described.

[0029] References in the specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in conjunction with the embodiment may be included in at least one embodiment of the present disclosure. The phrase "in one embodiment" appearing in various places in the specification does not necessarily refer to the same embodiment.

[0030] The following embodiments relate to using an artificial intelligence (AI) accelerator to increase the processing throughput of certain types of operations that can be offloaded (or delegated) from a host device to an AI accelerator. The host device hosts one or more virtual machines (VMs). At least one VM on the host can be associated with a virtual function, and the resources of the AI ​​accelerator are mapped to the VM via the virtual function. The virtual function enumerates the resources within the AI ​​accelerator mapped to the VM and the configuration of these resources within the accelerator. A driver within the VM can track the scheduling and computing status of tasks to be processed by the AI ​​accelerator. The driver can also obtain the code, data, and memory of the AI ​​accelerator mapped to the VM.

[0031] As used herein, a "virtual function" is a mapping of a set of resources within an artificial intelligence (AI) accelerator or a set of AI accelerators in an AI accelerator cluster to a virtual machine. Resource sets are referred to herein individually and collectively as "AI resources". An AI accelerator or an AI accelerator cluster is referred to herein as an "AI accelerator" unless the distinction between an AI accelerator and an AI accelerator cluster is described.

[0032] AI accelerators can be general purpose processing units (GPUs), artificial intelligence (AI) accelerators, math coprocessors, digital signal processors (DSPs), or other types of processors. AI accelerators can be proprietary designs, such as those from Baidu. AI accelerator or other GPU, etc. Although the embodiments are shown and described in the context of a host device securely coupled to one or more AI accelerators, the concepts described herein may be more generally implemented as a distributed processing system.

[0033] Multiple AI accelerators can be linked in a cluster, which is managed by a host device having a driver that converts application processing requests into processing tasks for one or more AI accelerators. The host device can support one or more virtual machines (VMs), each of which has a user associated with a corresponding VM. The driver can implement a virtual function that maps the resources of the AI ​​accelerator to the VM. The driver may include a scheduler that schedules application processing requests from multiple VMs to be processed by one or more AI accelerators. In one embodiment, the driver can analyze the processing requests in the scheduler to determine how to group one or more AI accelerators in the cluster, and whether to instruct one or more AI accelerators to disconnect from the group and enter a low power state to reduce heat and save energy.

[0034] The host device and the AI ​​accelerator may be interconnected via a high-speed bus, such as a peripheral component interconnect express (PCIe) or other high-speed bus. The host device and the AI ​​accelerator may exchange keys and initiate a secure channel via the PCIe bus prior to performing the operations of aspects of the present invention described below. Some of the operations include the AI ​​accelerator using an artificial intelligence (AI) model to perform reasoning using data provided by the host device. Before the host device trusts the AI ​​model inference, the host device may cause the AI ​​accelerator to perform one or more validation tests, as described below, including determining a watermark for the AI ​​model. In some embodiments and operations, the AI ​​accelerator is unaware that the host device is testing the validity of the results produced by the AI ​​accelerator.

[0035] The host device may include a central processing unit (CPU) coupled to one or more AI accelerators. Each AI accelerator may be coupled to the CPU via a bus or interconnect. The AI ​​accelerator may be implemented in the form of an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) device, or other forms of integrated circuits (ICs). Alternatively, the host processor may be part of a main data processing system, and the AI ​​accelerator may be one of many distributed systems (e.g., a cloud computing system, a software as a service (SaaS) system, or a platform as a service (PaaS) system) as an auxiliary system, and the main system may remotely offload its data processing tasks to the auxiliary system via a network. The link between the host processor and the AI ​​accelerator may be a peripheral component interconnect express (PCIe) link or a network connection, such as an Ethernet connection. Each AI accelerator may include one or more link registers for enabling (connecting) or disabling (disconnecting) a communication link with another AI accelerator.

[0036] In a first aspect, a computer-implemented method for migrating a source virtual machine (VM-S) that is executing an application that accesses a virtual function of an artificial intelligence (AI) accelerator may include storing a checkpoint of the state of the VM-S in a storage of multiple states of the VM-S. Each state of the VM-S may include the state of the resources of the VM-S, the state of the application, and the state of the virtual function of the AI ​​accelerator that maps the AI ​​resources to the VM-S. In response to receiving a command to migrate the VM-S and the virtual function, and receiving a checkpoint of the state of the VM-S selected to be used in performing the migration, the method may further include recording, then stopping one or more AI tasks being executed, and migrating the VM-S, the application, the one or more AI tasks, and the virtual function to a target VM (VM-T) on a target host using the selected checkpoint. In response to receiving a notification from the target hypervisor that the checkpoint has been successfully verified by the target hypervisor and the migration has been successfully completed, the recorded one or more AI tasks and applications may be restarted on the VM-T. The virtual function maps the resources of the AI ​​accelerator to the VM-S, and the user of the VM-S is the only user who can access the resources of the AI ​​accelerator whose resources are mapped to the VM-S by the virtual function. In an embodiment, the virtual function maps the resources of multiple AI accelerators to the VM-S, the checkpoint includes a communication configuration between the multiple AI accelerators, and the user of the VM-S is the only user who can access the resources of the multiple AI accelerators mapped to the VM-S by the virtual function. In an embodiment, the method further includes receiving a notification from the target hypervisor that the migration of the VM-S is completed and one or more recorded AI tasks have been successfully restarted. In response to this notification, a post-migration cleanup of the VM-S can be performed. Post-migration cleanup may include at least erasing the secure storage of the AI ​​accelerator, including any AI reasoning, AI model, secure computing, or part thereof, and erasing the storage of the VM-S associated with the AI ​​virtual function, as well as any calls to the virtual function by the application. Verifying the signature and freshness of the checkpoint may include decrypting the signature of the checkpoint using the public key of the VM-S, determining that the date and time stamp of the checkpoint are within the threshold date and time range, and verifying the hash of the checkpoint of the VM-S. In an embodiment, a checkpoint may include a record of one or more AI tasks being executed, configuration information of resources within one or more AI accelerators communicatively coupled to the VM-S, a date and timestamp of the checkpoint, and a snapshot of the storage of the VM-S, including virtual functions, scheduling information, and communication caches within one or more AI accelerators.

[0037] In a second aspect, a method for migrating a source virtual machine (VM-S) that is executing an application that accesses a virtual function (VF) of an artificial intelligence (AI) accelerator includes receiving, through a hypervisor of a target host, a checkpoint from a source virtual machine (VM-S) associated with mapping artificial intelligence (AI) processor resources to a virtual function (VF) of the VM-S, and receiving a request to host the VM-S as a target virtual machine (VM-T). The hypervisor of the target host allocates and configures resources for hosting the VM-S and the VF of the VM-S as VM-T according to the checkpoint. The hypervisor of the target host receives a data frame of the VM-S and stores the data frame to generate the VM-T. The hypervisor of the target host receives the recorded status of the unfinished AI task of the VM-S and restarts the unfinished AI task on the VM-T. In an embodiment, verifying the checkpoint of the VM-S and the VF includes decrypting the signature of the checkpoint with the public key of the VM-S, determining that the date and time stamp of the checkpoint fall within a pre-set range, and recalculating the hash of the checkpoint and determining whether the recalculated hash matches the hash stored in the checkpoint. Responsive to successful verification of the checkpoint, migration of VM-S to the hypervisor of the target host proceeds, generating VM-T on the target host.

[0038] In a third aspect, a computer-implemented method for migrating a source virtual machine (VM-S) executing an application accessing a virtual function of an artificial intelligence (AI) accelerator includes: in response to receiving a command to migrate the VM-S and the virtual function, and in response to receiving a selection for performing a checkpoint of the migrated VM-S and the virtual function, recording and then stopping one or more executing AI tasks of the application. The method further includes generating or selecting a state of the AI ​​accelerator associated with the virtual function, and then transmitting the checkpoint and the state of the AI ​​accelerator to a hypervisor of a target host to generate a migrated target virtual machine (VM-T).

[0039] In response to receiving a notification that the target host has verified the checkpoint and the AI ​​accelerator state, and that the target host has generated and configured resources for generating VM-T, the target host migrates the VM-S and the virtual functions to the VM-T. The migration includes the target host loading data in the AI ​​accelerator state frame for the AI ​​accelerator. In an embodiment, the method further includes performing post-migration cleanup of the VM-S and the virtual functions in response to receiving a notification that the VM-T has restarted the application and the AI ​​task. The post-migration cleanup of the VM-S may include (1) erasing at least the secure storage of the AI ​​accelerator, including any AI reasoning, AI model, intermediate results of secure computation, or part thereof; (2) erasing the storage of the VM-S associated with the virtual function, and any call to the virtual function by the application. In an embodiment, storing the checkpoint of the state of the VM-S and the virtual function may include storing the checkpoint of the state of the VM-S and the VF in the storage of multiple checkpoints of the VM-S. Each checkpoint of the VM-S may include the state of the resources of the VM-S, the state of the application, and the state of the virtual function associated with the resources of the AI ​​accelerator. In an embodiment, the checkpoint may further include a record of one or more AI tasks being executed, configuration information of resources within the AI ​​accelerator communicatively coupled to the VM-S, and a snapshot of the storage of the VM-S. The checkpoint may further include virtual function scheduling information and communication caches within one or more AI accelerators, as well as a date and timestamp of the checkpoint. In an embodiment, generating the state of the AI ​​accelerator may include: (1) storing a date and timestamp of the state in the AI ​​accelerator state, (2) storing the contents of a memory within the AI ​​accelerator in the AI ​​accelerator state, including one or more registers associated with a processor of the AI ​​accelerator, and a cache, queue, or pipeline of pending instructions to be processed by the AI ​​accelerator, and (3) generating a hash of the AI ​​accelerator state and digitally signing the state, hash, date, and timestamp. In an embodiment, the AI ​​accelerator state may further include one or more register settings indicating one or more other AI accelerators in a cluster of AI accelerators with which the AI ​​accelerator is configured to communicate. In an embodiment, verifying the signature and freshness of the AI ​​accelerator state may include decrypting the signature of the AI ​​state using the public key of the VM-S, determining that the date and time stamp of the AI ​​accelerator state is within a threshold date and time range, and verifying the hash of the AI ​​accelerator state.

[0040] Any of the above functions may be programmed as executable instructions onto one or more non-transitory computer readable media. When the executable instructions are executed by a processing system having at least one hardware processor, the processing system causes the functions to be implemented. Any of the above functions may be implemented by a processing system having at least one hardware processor, the hardware processor being coupled to a memory programmed with the executable instructions, which, when executed, causes the processing system to implement the functions.

[0041] Figure 1 is a block diagram illustrating a secure processing system 100 that can migrate virtual machines and have checkpoint authentication and / or artificial intelligence (AI) accelerator state verification according to one embodiment. Figure 1 , the system configuration 100 includes, but is not limited to, one or more client devices 101-102 communicatively coupled to a source data processing (DP) server 104-S (e.g., a host) and a target data DP server 104-T via a network 103. The DP server 104-S may host one or more clients. One or more clients may be virtual machines. As described herein, any virtual machine on the DP server 104-S may be migrated to the target DP server 104-T.

[0042] The client devices 101-102 may be any type of client device, such as a personal computer (e.g., desktop, laptop, and tablet), a "thin" client, a personal digital assistant (PDA), a web-enabled device, a smart watch, or a mobile phone (e.g., a smartphone), etc. Alternatively, the client devices 101-102 may be virtual machines on the DP servers 104-S or 104-T. The network 103 may be any type of network, such as a local area network (LAN), a wide area network (WAN) such as the Internet, a high-speed bus, or a combination thereof, wired or wireless.

[0043] Servers (e.g., hosts) 104-S and 104-T (collectively referred to as DP servers 104, unless otherwise specified) can be any kind of server or server cluster, such as a Web or cloud server, an application server, a back-end server, or a combination thereof. Server 104 further includes an interface (not shown) to allow clients such as client devices 101-102 to access resources or services provided by server 104 (such as resources and services provided by AI accelerators via server 104). For example, server 104 can be a cloud server or a server in a data center that provides various cloud services to clients, such as, for example, cloud storage, cloud computing services, artificial intelligence training services, data mining services, etc. Server 104 can be configured as part of a software as a service (SaaS) or platform as a service (PaaS) system on the cloud, and the cloud can be a private cloud, a public cloud, or a hybrid cloud. The interface can include a web interface, an application programming interface (API), and / or a command line interface (CLI).

[0044] For example, the client may be a user application (e.g., a web browser, an application) of the client device 101. The client may send or transmit instructions for execution (e.g., AI training, AI reasoning instructions, etc.) to the server 104, and this instruction is received by the server 104 via an interface on the network 103. In response to this instruction, the server 104 communicates with the AI ​​accelerators 105-107 to complete the execution of the instruction. The source DP server 104-S may be communicatively coupled to one or more AI accelerators. The client virtual machine hosted by the DP server 104-T running the application using one or more AI accelerators 105-T..107-T may be migrated to the target DP server 104-T to run on the corresponding AI accelerator 105-T..107-T. In some embodiments, the instruction is a machine learning type instruction, in which an AI accelerator as a dedicated machine or processor can execute the instruction many times faster than a general-purpose processor. The server 104 can therefore control / manage the execution jobs of one or more AI accelerators in a distributed manner. The server 104 then returns the execution results to the client devices 101-102 or the virtual machines on the server 104. The AI ​​accelerator may include one or more specialized processors, such as those available from Baidu Inc. Obtained from Baidu The artificial intelligence (AI) chipset, or alternatively, the AI ​​accelerator may be an AI chipset from another AI chipset vendor.

[0045] According to one embodiment, each of the applications accessing any of the AI ​​accelerators 105-S..107-S or 105-T..107-T (collectively referred to as 105..107 unless otherwise stated) hosted by a data processing server 104 (also referred to as a host) can verify that the application is provided by a trusted source or vendor. Each application can be launched and executed within a user storage space and executed by the central processing unit (CPU) of the host 104. When the application is configured to access any of the AI ​​accelerators 105-107, an obfuscated connection can be established between the host 104 and a corresponding one of the AI ​​accelerators 105-107, so that data exchanged between the host 104 and the AI ​​accelerator 105-107 is protected from malware / intrusion attacks.

[0046] Figure 2A20 is a block diagram illustrating a secure computing environment 200 between one or more hosts and one or more artificial intelligence (AI) accelerators according to some embodiments. In one embodiment, the system 200 provides a protection scheme for obfuscated communications between a host 104 and an AI accelerator 105-107, with or without hardware modifications to the AI ​​accelerators 105-107. The host or server 104 can be described as a system having one or more layers to be protected from intrusion, such as user applications 205, runtime libraries 206, drivers 209, operating systems 211, hypervisors 212, and hardware 213 (e.g., central processing unit (CPU) 201 and storage devices 204). Below the applications 205 and runtime libraries 206, one or more drivers 209 can be installed to interface to the hardware 213 and / or AI accelerators 105-107.

[0047] The driver 209 may include a scheduler 209A that schedules processing tasks requested by one or more user applications 205. The driver 209 may further include an analyzer 209B having logic that analyzes the processing tasks scheduled for execution on the AI ​​accelerators 105-107 to determine how to best configure the AI ​​accelerators 105-107 based on scheduling criteria such as processing throughput, energy consumption, and heat generated by the AI ​​accelerator. The driver 209 may further include one or more policies designed to configure the AI ​​accelerators to achieve the scheduling criteria. Configuring the AI ​​accelerators may include grouping the AI ​​accelerators into one or more groups, removing one or more AI accelerators from one or more groups. The driver 209 may further include a checkpoint 209C. The checkpoint 209C may snapshot the state of the user application 205, the storage within the VM 201, the scheduler 209A state, the analyzer 209B state, and the configuration of the virtual function within the VM 201. As used herein, a virtual function is a mapping of a resource set within an artificial intelligence (AI) accelerator, such as 105, or a cluster of AI accelerators 105..107 to a virtual machine. Reference below Figure 3 , 4A and 4B describe virtual functions.

[0048] AI accelerators that are not assigned to a group of AI accelerators within an AI accelerator cluster may be placed in a low power state to save energy and reduce heat. The low power state may include reducing the clock speed of the AI ​​accelerator or entering a standby state, in which the AI ​​accelerator is still communicatively coupled to the host device, and may enter a run state, in which the AI ​​accelerator is ready to receive processing tasks from the host device. AI accelerators that are not assigned to a group in a cluster may alternatively remain powered on so that the driver 209 can assign work to a single AI accelerator that is not a member of a group of AI accelerators.

[0049] Configuring the AI ​​accelerator may further include instructing one or more AI accelerators to generate a communication link (connection) with one or more other AI accelerators to form a group of AI accelerators within the AI ​​accelerator cluster. Configuring the AI ​​accelerator may further include instructing one or more DP accelerators to disconnect a communication link (disconnection) between the AI ​​accelerator and one or more other AI accelerators. The connection and disconnection of the AI ​​accelerator may be managed by one or more link registers in each AI accelerator.

[0050] In a policy-based partitioning embodiment, the AI ​​accelerator configuration policy is a single policy that describes the communication link (connected or disconnected) of each AI accelerator. Although the configuration of each AI accelerator can (and usually will) be different from other AI accelerators, the configuration of each AI accelerator is included in a single policy, and each AI accelerator in the cluster receives the same policy. Then, each AI accelerator configures itself according to the policy portion that describes the configuration of the AI ​​accelerator. Policy-based partitioning can be based on the analysis of processing tasks in the scheduler 209A. The analysis can determine the best allocation for dividing the AI ​​accelerator into groups. In one embodiment, time-sharing processing tasks within a group of processors or across multiple groups of processors (to optimize throughput) minimizes energy consumption and heat generated. The advantages of policy-based division of AI accelerators into groups include fast partitioning of AI accelerators, flexible scheduling of processing tasks within or across groups, time-sharing of AI accelerators, and time-sharing of groups.

[0051] In a dynamic partitioning embodiment, an AI accelerator policy is generated for each AI accelerator. Driver 209 can dynamically change the configuration of each AI accelerator, including reorganizing the groups of AI accelerators, removing one or more AI accelerators from all groups, and setting these AI accelerators to a low power state. In a dynamic partitioning embodiment, each group of AI accelerators is assigned to a single user, rather than sharing the AI ​​accelerators between users in a time-sharing manner. Driver 209 may include an analyzer 209B that analyzes the processing tasks within scheduler 209A to determine the optimal grouping of AI accelerators. The analysis can generate a configuration for one or more AI accelerators, and the configuration can be deployed to each such AI accelerator for reconfiguration. The advantages of dynamic partitioning include saving energy by setting one or more processors to a low power state, and user-specific processing for an AI accelerator or a group of AI accelerators, rather than time slicing between users.

[0052] The hardware 213 may include a processing system 201 having one or more processors 201. The hardware 213 may further include a storage device 204. The storage device 204 may include one or more artificial intelligence (AI) models 202 and one or more kernels 203. The kernel 203 may include a signature kernel, a watermark-enabled kernel, an encryption and / or decryption kernel, etc. The signature kernel may digitally sign any input according to the kernel's programming when executed. The watermark-enabled kernel may extract a watermark from a data object (e.g., an AI model or other data object). The watermark-enabled kernel may also implant a watermark into an AI model, an inference output, or other data object.

[0053] A watermark kernel (e.g., a watermark inheritance kernel) can inherit a watermark from another data object and implant this watermark into a different object, such as an inference output or an AI model. As used herein, a watermark is an identifier associated with an AI model or an inference generated by an AI model and can be implanted therein. For example, a watermark can be implanted in one or more weight variables or bias variables. Alternatively, one or more nodes (e.g., dummy nodes that are not used or unlikely to be used by the artificial intelligence model) can be created to implant or store a watermark.

[0054] The host 104 may be a CPU system that may control and manage the execution of jobs on the host 104 and / or the AI ​​accelerators 105-107. To protect / shield the communication channel 215 between the AI ​​accelerators 105-107 and the host 104, different components may be required to protect different layers of the host system that are susceptible to data intrusion or attack.

[0055] According to some embodiments, system 200 includes host system 104 and AI accelerators 105-107. There may be any number of AI accelerators. The AI ​​accelerator may include Baidu AI chipset or other AI chipset, such as a graphics processing unit (GPU) that can perform artificial intelligence (AI) intensive computing tasks. In one embodiment, the host system 104 includes hardware having one or more CPUs 213, which is optionally equipped with a security module (such as an optional trusted platform module (TPM)) within the host 104. The optional TPM is a dedicated chip on an endpoint device that stores encryption keys specific to the host system (e.g., RSA encryption keys) for hardware verification. Each TPM chip can contain one or more RSA key pairs (e.g., public and private key pairs), called endorsement keys (endorsement key, EK) or endorsement credentials (endorsement credential, EC), i.e., root keys. The key pairs are stored in the optional TPM chip and cannot be accessed by software. Critical parts of the firmware and software can be hashed by the EK or EC before execution to protect the system from unauthorized firmware and software modifications. Therefore, the optional TPM chip on the host can be used as a root of trust for secure boot.

[0056] The optional TPM chip can also protect the driver 209 and the operating system (OS) 211 in the working kernel space to communicate with the AI ​​accelerator 105-107. Here, the driver 209 is provided by the AI ​​accelerator vendor and can be used as a driver 209 for the user application 205 to control the communication channel 215 between the host and the AI ​​accelerator. Since the optional TPM chip and the secure boot processor protect the OS 211 and the driver 209 in its kernel space, the TPM also effectively protects the driver 209 and the OS 211.

[0057] Because the communication channel 215 for the AI ​​accelerator 105-107 can be dedicated to the OS 211 and the driver 209, the communication channel 215 can be protected by an optional TPM chip. In one embodiment, the communication channel 215 includes a peripheral component interconnect (PCI) or a peripheral component interconnect express (PCIE) channel. In one embodiment, the communication channel 215 is an obfuscated communication channel.

[0058] The runtime library 206 may convert application programming interface (API) calls into commands for execution, configuration, and / or control of the AI ​​accelerator. In one embodiment, the runtime library 206 provides a set of predetermined (e.g., predefined) kernels for user applications to execute. In one embodiment, the kernels may be stored as kernels 203 in the storage device 204.

[0059] The operating system 211 may be any Release, OS or OS, or other operating system.

[0060] The system can be booted through an optional TPM-based secure boot. The optional TPM secure boot ensures that only the signed / authenticated operating system 211 and accelerator driver 209 are booted in the kernel space that provides accelerator services. In one embodiment, the operating system 211 can be loaded through a hypervisor (212). A hypervisor or virtual machine manager 212 is computer software, firmware, or hardware that creates and runs virtual machines. The kernel space is a declarative area or scope in which the kernel (i.e., a set of predetermined (e.g., predefined) functions for execution) is identified to provide functions and services to user applications. If the integrity of the system is compromised, the optional TPM secure root may not be able to boot and instead will shut down the system.

[0061] After booting, the runtime library 206 runs the user application 205. In one embodiment, the user application 205 and the runtime library 206 are statically linked and started together. In another embodiment, the runtime library 206 is started first, and then the user application 205 is dynamically loaded. A statically linked library is a library that is linked to an application at compile time. Dynamic loading can be performed by a dynamic linker. The dynamic linker loads and links shared libraries so that the user application can run at runtime. Here, the user application 205 and the runtime library 206 are visible to each other at runtime, e.g., all process data is visible to each other.

[0062] In one embodiment, the user application 205 can only call kernels from a set of kernels predetermined by the runtime library 206. On the other hand, the user application 205 and the runtime library 206 are fortified with side-channel resistant algorithms to defend against side-channel attacks, such as cache-based side-channel attacks. A side-channel attack is any attack based on information obtained from the implementation of a computer system, rather than a weakness in the implemented algorithm itself (e.g., cryptanalysis and software bugs). Examples of side-channel attacks include cache attacks, which are attacks based on the attacker's ability to monitor the cache of a shared physical system in a virtualized environment or cloud environment. Fortification can include caching, masking of outputs generated by algorithms to be placed on the cache. Next, when the user application finishes execution, the user application terminates its execution and exits.

[0063] In one embodiment, the set of kernels 203 includes obfuscated kernel algorithms. In one embodiment, the obfuscated kernel algorithm can be a symmetric or asymmetric algorithm. A symmetric obfuscation algorithm can use the same algorithm to obfuscate and de-obfuscate data communication. An asymmetric obfuscation algorithm requires a pair of algorithms, where the first in the pair is used for obfuscation and the second in the pair is used for de-obfuscation, and vice versa. In another embodiment, the asymmetric obfuscation algorithm includes a single obfuscation algorithm for obfuscating a data set but is not intended to de-obfuscate the data set, e.g., there is no corresponding de-obfuscation algorithm.

[0064] Obfuscation refers to obscuring the intended meaning of a communication by making the communication message difficult to understand, often using confusing and ambiguous language. Obfuscated data is more difficult and complex to reverse engineer. Obfuscation algorithms can be applied to obfuscate (encrypt / decrypt) data communications prior to data communication, thereby reducing the chance of eavesdropping. In one embodiment, the obfuscation algorithm can further include an encryption scheme to further encrypt the obfuscated data for an additional layer of protection. Unlike encryption, which can be computationally intensive, obfuscation algorithms can simplify computations.

[0065] Some obfuscation techniques may include, but are not limited to, letter obfuscation, name obfuscation, data obfuscation, control flow obfuscation, etc. Letter obfuscation is the process of replacing one or more letters in the data with specific replacement letters, making the data meaningless. Examples of letter obfuscation include letter rotation functions, where each letter is moved or rotated a predetermined number of positions along the alphabet. Another example is reordering letters or scrambling the order of letters according to a specific pattern. Name obfuscation is the process of replacing a specific target string with a meaningless string. Control flow obfuscation can use additional code to change the order of control flow in the program (inserting dead code, inserting uncontrolled jumps, inserting alternative structures) to hide the true control flow of the algorithm / AI model.

[0066] In summary, system 200 provides multiple layers of protection for AI accelerators (for data transmission, including machine learning models, training data, and inference outputs) to prevent loss of data confidentiality and integrity. System 200 may include an optional TPM-based secure boot protection layer and a kernel verification / authentication layer. System 200 may include applications that use side-channel protection algorithms to defend against side-channel attacks, such as cache-based side-channel attacks.

[0067] The runtime library 206 may provide a fuzzy kernel algorithm to fuzzify data communications between the host 104 and the AI ​​accelerators 105-107. In one embodiment, fuzzification may be paired with a cryptographic scheme. In another embodiment, fuzzification is the only protection scheme and cryptographic-based hardware is rendered unnecessary for the AI ​​accelerator.

[0068] Figure 2B2 is a block diagram illustrating a secure computing environment between one or more hosts and one or more artificial intelligence (AI) accelerators according to one embodiment. In one embodiment, a host channel manager (HCM) 250 includes an optional authentication module 251, an optional termination module 252, an optional key manager 253, an optional key storage 54, and an optional cryptographic engine 255. The optional authentication module 251 can authenticate a user application running on the host server 104 to obtain permission to access or use resources of the AI ​​accelerator 105. The HCM 250 can communicate with an accelerator channel manager (ACM) 280 of the AI ​​accelerator 215 via a communication channel 215.

[0069] The optional termination module 252 can terminate the connection (e.g., the channel associated with this connection will be terminated). The optional key manager 253 can manage (e.g., create or destroy) asymmetric key pairs or symmetric keys used for encryption / decryption of one or more data packets of different secure data exchange channels. Here, each user application (as Figure 2A The user applications 205 (part of) can correspond or map to different secure data exchange channels in a one-to-many relationship, and each data exchange channel can correspond to the AI ​​accelerator 105. Each application can use multiple session keys, where each session key is used for a secure channel corresponding to an AI accelerator (e.g., accelerators 105-107). The optional key storage 254 can store encrypted asymmetric key pairs or symmetric keys. The optional cryptographic engine 255 can encrypt or decrypt data packets of data exchanged over any secure channel. Note that some of these modules can be integrated into fewer modules.

[0070] In one embodiment, the AI ​​accelerator 105 includes an ACM 280, non-sensitive resources 290, and sensitive resources 270. The ACM 280 is a corresponding module corresponding to the HCM 250 and is responsible for managing communications between the host 104 and the AI ​​accelerator 105, such as, for example, resource access control. The ACM 280 includes a link configuration module 281 that cooperates with the HCM 250 of the host server 104 to establish a communication channel 215 between the host server 104 and the AI ​​accelerator 105. The ACM 280 further includes a resource manager 282. The resource manager 282 enforces restricted access to the sensitive resources 270 and the non-sensitive resources 290. In one embodiment, the sensitive resources 270 occupy a first address space range within the AI ​​accelerator 105. The non-sensitive resources 290 occupy a second address space range within the AI ​​accelerator 105. In one embodiment, the first address space and the second address space are mutually exclusive and non-overlapping. In one embodiment, resource manager 282 further includes logic (e.g., access control logic) that allows host server 104 to access both sensitive resources 270 and non-sensitive resources 280. In one embodiment, resource manager 282 enforces access and configuration policies received from host server 104, as further described below.

[0071] Sensitive resources 270 may include an optional key manager 271, an optional key storage 272, a true random number generator 273, an optional cryptographic engine 274, and memory / storage 277. The optional key manager 271 may manage (e.g., generate, securely retain, and / or destroy) asymmetric key pairs or symmetric keys. The optional key storage 272 may store encrypted asymmetric key pairs or symmetric keys in a secure storage within the sensitive resources 270. The true random number generator 273 may generate seeds for key generation and cryptographic engine 274 use, such as an AI accelerator for verifying a link. The optional cryptographic engine 274 may encrypt or decrypt key information or data packets for data exchange. The memory / storage 277 may include storage for AI models 275 and kernels 276. The kernels 276 may include watermark kernels (including inherited watermark kernels, watermark-enabled kernels, watermark signature kernels, etc.), encryption and decryption kernels, and related data.

[0072] The AI ​​accelerator 105 may further include non-sensitive resources 290. The non-sensitive resources 290 may include one or more processors or processing logic 291 and memory / storage 292. The processor or processing logic 192 is capable of executing instructions or programs to perform various processing tasks, such as AI tasks (e.g., machine learning processes).

[0073] The link configuration module 281 is responsible for establishing or connecting a link or path from one AI accelerator to another AI accelerator, or terminating or disconnecting a link or path from one AI accelerator to another AI accelerator. In one embodiment, in response to a request to join a group of AI accelerators (e.g., from a host), the link configuration module 281 establishes a link or path from the corresponding AI accelerator to at least some of the AI ​​accelerators in the group or cluster so that the AI ​​accelerator can communicate with other AI accelerators, such as accessing the resources of other AI accelerators for AI processing. Similarly, in response to a request to switch from a first group of AI accelerators to a second group of AI accelerators, the link configuration module 281 terminates the existing link from the corresponding AI accelerator of the first group and establishes a new link to the second group of AI accelerators.

[0074] In one embodiment, the AI ​​accelerator 105 further includes an AI processing unit (not shown), which may include an AI training unit and an AI reasoning unit. The AI ​​training and reasoning units may be integrated into a single unit in the sensitive resource 270. The AI ​​training module is configured to train the AI ​​model using a set of training data. The AI ​​model and training data to be trained may be received from the host system 104 via the communication link 215. In one embodiment, the training data may be stored in the non-sensitive resource 290. The AI ​​model reasoning unit may be configured to execute a trained artificial intelligence model on a set of input data (e.g., a set of input features) to infer and classify the input data. For example, an image may be input to the AI ​​model to classify whether the image contains a person, a landscape, etc. The trained AI model and input data may also be received from the host system 104 via the communication link 215 via the interface 140.

[0075] In one embodiment, the watermark unit (not shown) in the sensitive resource 270 may include a watermark generator and a watermark burner (also referred to as a "watermark implanter"). The watermark unit (not shown) may include a watermark kernel executor or kernel processor (not shown) of the sensitive resource 270 to execute the kernel 276. In one embodiment, the kernel may be received from the host 104, or retrieved from a persistent or non-persistent storage, and executed in the kernel memory 276 of the sensitive resource 270 of the AI ​​accelerator 105. The watermark generator is configured to generate a watermark using a predetermined watermark algorithm. Alternatively, the watermark generator may inherit a watermark from an existing watermark or extract a watermark from another data structure or data object, such as an artificial intelligence model or an input data set that may be received from the host system 104. The watermark implanter is configured to burn or implant a watermark in a data structure, such as an artificial intelligence model, output data generated by the artificial intelligence model. The artificial intelligence model or output data in which the watermark is implanted may be returned from the AI ​​accelerator 105 to the host system 104 via the communication link 215. Note that the AI ​​accelerators 105 - 107 have the same or similar structures or components and the description regarding the AI ​​accelerator will apply to all AI accelerators in the entire application.

[0076] Figure 3 is a block diagram showing a host 104 controlling an artificial intelligence accelerator cluster 310 according to one embodiment, each cluster having a virtual function for mapping resources of an AI accelerator group 311 within the cluster to a virtual machine on the host, each artificial intelligence accelerator having secure resources and non-secure resources.

[0077] The host 104 may include an application 205, such as an artificial intelligence (AI) application, a runtime library 206, one or more drivers 209, an operating system 211, and hardware 213, as described above with reference to Figure 2A and 2B Each of these is described in detail and will not be repeated here. In a virtual computing embodiment, the host 104 may further include a hypervisor 212, such as or Hypervisor 212 may be a type 1 "bare metal" or "native" hypervisor that runs directly on the physical server. In one embodiment, hypervisor 212 may be a type 2 hypervisor that is loaded inside and managed by operating system 211 like any other application. In either case, hypervisor 212 may support one or more virtual machines (not shown) on host 104. In such an aspect, a virtual machine (not shown) may be considered to be a virtual machine that is a part of a server. Figure 1 Client devices 101 and 102.

[0078] The artificial intelligence (AI) accelerator cluster 310 may include the above reference Figure 2A and 2B AI accelerators 105-107 are described. Figure 3 , the AI ​​accelerator cluster 310 may include, for example, eight (8) AI accelerators labeled A through H. Each AI accelerator in the accelerator cluster 310 may have one or more communication links 215 to one or more other AI accelerators in the accelerator cluster 310. Figure 2A and 2B AI accelerator communication link 215 is depicted. Each AI accelerator in cluster 310 is configured according to policies received from host 104 driver 209. Each AI accelerator in cluster 310 may have sensitive resources 270 and non-sensitive resources 290.

[0079] exist Figure 3 In the example shown in , AI accelerator AD is configured as four (4) AI accelerators of the first group 311. The resources of the AI ​​accelerators in the first group 311 are configured and managed by virtual function VF1 and are associated with the first virtual machine. AI accelerator EH is configured in the second group 312 of four (4) AI accelerators. The resources of the AI ​​accelerators in the second group 312 are configured and managed by virtual function VF2 and are associated with the second virtual machine. The resources of the two groups 311 and 312 are mutually exclusive, and users of either group cannot access the resources of the other of the two groups. In the first group 311 of AI accelerators, each AI accelerator has a communication link directly to another accelerator, such as AB, AC, BD, and CD, or has a communication path to another accelerator via one or more intermediate accelerators, such as ABD, ACD, etc. The second group 312 is shown as having a direct communication link between each AI accelerator in the second group 312 and each other AI accelerator in the second group 312. Driver 209 can generate a policy where each AI accelerator in a group has a direct communication link with each or some other AI accelerators in the group. In the case of a first group 311, driver 209 can generate a policy that further includes, for example, instructions for AI accelerators A and D to generate communication links with each other and AI accelerators B and C to generate communication links with each other. There can be any number of AI accelerators in cluster 310, configured into any number of groups.

[0080] In a static policy-based embodiment, a single policy defines the configuration of each AI accelerator and is transmitted from the driver 209 to all AI accelerators in the cluster 310. In an embodiment, the driver 209 can transmit the policy to all AI accelerators in the cluster in a single broadcast message. Each AI accelerator reads the policy and establishes (connects) or cuts off (disconnects) a communication link with one or more AI accelerators in the cluster 310, thereby configuring the AI ​​accelerators into one or more groups. Figure 3 In the embodiment, there are eight (8) AI accelerators configured into two groups of four (4) AI accelerators in each group. Each AI accelerator in the group either has a direct communication link with each AI accelerator in the group, or has an indirect communication path with each AI accelerator in the group via one or more AI accelerators, where the AI ​​accelerator has a direct communication link to the one or more AI accelerators. In a static policy based environment, the scheduler 209A of the driver 209 can schedule processing tasks on one or more groups of the cluster 310 using time slicing between applications 205 and / or users of virtual machines. In an embodiment, each group of accelerators in the accelerator cluster 310 can have a different and independent scheduler 209A. The static policy can be changed by the driver 209 to generate a new policy describing the configuration of each AI accelerator in the cluster 310.

[0081] Each AI accelerator in cluster 310 (e.g., link configuration module 281 and / or resource manager 282) reconfigures itself according to the policy, establishing (connecting) or severing (disconnecting) a communication link between the AI ​​accelerator and one or more other AI accelerators in cluster 310. Static policy-based configuration is fast because the configuration is transmitted in a single message, such as a broadcast message, and each AI accelerator configures itself substantially in parallel with the other AI accelerators in cluster 310. Since the policies for all AI accelerators are transmitted to all AI accelerators at the same time, the configuration can occur very quickly. For example, if the policy includes an instruction to AI accelerator "A" to generate a link to AI accelerator "B", then the policy also has an instruction for AI accelerator B to generate a link to AI accelerator A. Each AI accelerator can open the port of its own link substantially at the same time, thereby opening the link between AI accelerator A and AI accelerator B very quickly. In one embodiment, a single policy can be represented as an adjacency list of AI accelerators.

[0082] Configuration based on static policies is also effective because it supports time-slicing scheduling between different users and supports allocating a user's processing tasks to multiple AI accelerator groups in cluster 310. Static policies can be generated by determining the characteristics of the processing tasks in scheduler 209A from analyzer 209B. For example, scheduler 209A may include a large number of tasks that use the same AI model to perform reasoning or further train the AI ​​model. The analyzer can generate a strategy for configuring multiple AI accelerators to prepare for reasoning or training of the AI ​​model. Configuration can include identifying groupings of AI accelerators and loading one or more AI models into sensitive memories of one or more AI accelerators to prepare for processing tasks in scheduler 209A.

[0083] In an embodiment based on a dynamic policy, the driver 209 can configure each AI accelerator in the cluster 310 individually to implement the configuration of the AI ​​accelerator. The policy is transmitted to each AI accelerator individually. In fact, in an embodiment based on a dynamic policy, the policies transmitted to each AI accelerator are usually different from each other. The AI ​​accelerator receives the policy and configures itself according to the policy. The configuration includes the AI ​​accelerator configuring itself into or out of a group in the cluster 310. The AI ​​accelerator configures itself into the group by establishing a communication link (connection) with at least one AI accelerator in the group according to the policy. The AI ​​accelerator leaves the group by cutting off (disconnecting) the communication link between the AI ​​accelerator and all AI accelerators in the group. After the configuration is completed, if the AI ​​accelerator is not a member of any group of AI accelerators, the AI ​​accelerator can be set to a low power consumption model to reduce heat and save energy. In one embodiment, the scheduler 209A assigns an AI accelerator or an AI accelerator group to each user or application of the cluster 310 for which the scheduler 209A schedules processing tasks.

[0084] Figure 4A is a block diagram illustrating components of a data processing system having an artificial intelligence (AI) accelerator to implement a method of virtual machine migration with checkpoint authentication in a virtualized environment according to an embodiment.

[0085] The source host (HOST-S) 401 may support multiple virtual machines (VMs), such as a first (source) VM (VM1-S) to be migrated to a target host (HOST-T) 451 via network 103. Network 103 may be any network, as described above with reference to Figure 1 As described. HOST-S 401 may also support additional VMs, such as VM2 and VM3. Virtual machines VM1-S, VM2, and VM3 (each labeled "402") may each include at least one application 403 and at least one driver 404. Driver 404 may include one or more function libraries and application programming interfaces (APIs) that enable VM 402 including driver 404 to communicate with one or more artificial intelligence (AI) accelerators 410 that are communicatively coupled to VM 402 via hypervisor 405, CPU 406, and bus 407.

[0086] Hypervisor X 405 may be any type of hypervisor, including a "bare metal" hypervisor running on the hardware of HOST-S 401, or the hypervisor may run on an operating system (not shown) of HOST-S 401 executing on the hardware of the host, such as CPU 406 and memory (not shown). CPU 406 may be any type of CPU, such as a general purpose processor, a multi-core processor, a pipeline processor, a parallel processor, etc. Bus 407 may be any type of high-speed bus, such as Peripheral Component Interconnect Express (PCIe), a fiber optic bus, or other type of high-speed bus. As described above with reference to Figure 2A , Figure 2B and Figure 3 As described, the communication channel 215, i.e., the communication on the bus 407, can be encrypted. The bus 407 communicatively couples the CPU 406 to one or more artificial intelligence (AI) accelerators 410. Each VM can have a separate encrypted communication channel 215 that uses one or more keys that are different from the encrypted communication channel 215 of each other VM.

[0087] Each AI accelerator 410 can host one or more virtual functions, such as VF1, VF2, ... VFn, each of which is marked with reference numeral 411 in FIG. 4. Virtual functions 411 map resources 412, such as RES1, RES2, ... RESn of accelerator ACC1 410 to a specific host virtual machine 402. Each virtual machine 402 has users. A virtual function 411 associated with a specific VM 402 (e.g., VM1-S) can only be accessed by a user of the specific VM 402 (e.g., VM1-S). Virtual machine resources are each marked with reference numeral 412 in FIG. Virtual machine resources 412 are referred to above. Figure 2B 451, and includes resources such as non-sensitive resources 290 (including processing logic 291 and memory / storage 292), accelerator channel management 280 (including link configuration 281 and resource manager 282), and sensitive resources 270 (including AI model 275, kernel 276, memory / storage 277 and key manager 271, key storage 272, true random number generator 273, and cryptographic engine 274). As described more fully below, after migrating a virtual machine, such as VM1-S, to a target host, such as HOST-T 451, at least the sensitive resources should be erased so that after migrating the migrated virtual functions of VM1-S to the target host HOST-T 451 and allocating the now unused resources of the migrated virtual functions of VM1-S to the new VM, the sensitive data of the migrated VM1-S and the sensitive data associated with the virtual functions associated with VM1-S will not be accessible to the new VM.

[0088] The target host, such as HOST-T 451, can have the same or similar hardware and software configuration as HOST-S 401. Accelerator 410 and accelerator 460 should be of the same or similar type, such as having compatible instruction sets for their respective processors. HOST-T 451 should have sufficient available resources that VM-S may need in quantity so that VM1-S can migrate to VM1-T. Qualitatively, HOST-S 401 and HOST T-451 should have compatible operating hardware and software. For example, HOST-S401 accelerator 410 can belong to the same manufacturer as accelerator ACC2 460 on HOST-T 451 and have a compatible model, otherwise the migration may not be successful.

[0089] Checkpoint 420 is a snapshot of the state of VM1-S, up to and including virtual functions 411 (e.g., VF1) that are migrated as part of the migration of VM1-S from HOST-S 401 to HOST-T 451. The checkpoint of VM1-S and the associated virtual functions may include the following information. In an embodiment, the checkpoint does not include information within resources 412 contained within accelerator 410. The following list of information included in the checkpoint is for illustration and not limitation. One skilled in the art may add or delete information from the checkpoint 420 of virtual machines and virtual functions to be migrated in the following table.

[0090]

[0091]

[0092] Checkpoint 209C can be based on Figure 6420. A checkpoint frame 420 may be generated, for example, at a specified time increment, upon detection of a system anomaly or failure, or upon receipt of an instruction to take a checkpoint frame 420. Such instructions may come from a user, such as an administrator or an end user. The size of each checkpoint frame 420 may be approximately 1 gigabyte (GB), for example. In an embodiment, the checkpoint 209 may include a circular buffer that stores up to a specified number k of checkpoint frames 420. When the buffer is full, the next added frame overwrites the oldest checkpoint frame 420. When it is time to migrate virtual machines and virtual functions, the user may select a specific checkpoint frame 420 to use for performing the migration, indicating that the user prefers a known state of the running application 403 for migration. In an embodiment, the migration uses the most recent checkpoint frame 420 by default. In an embodiment, during migration of source VM1-S, the checkpoint frame 420, a hash of the checkpoint frame 420, and a date and time stamp of the checkpoint frame 420 may be digitally signed before transmitting the checkpoint frame 420 from the source VM1-S to the hypervisor of the target host HOST-T 451.

[0093] When the hypervisor 455 of the target host HOST-T 451 receives the checkpoint frame 420, the hypervisor 455 may decrypt the checkpoint frame 420 using the public key of VM1-S, verify that the date and time stamp fall within a predetermined time window, and verify the hash of the checkpoint frame. Verifying the date and time stamp confirms the freshness of the checkpoint frame 420. If the hypervisor 455 of the target HOST-T 451 confirms the checkpoint frame 420, the hypervisor 455 of HOST-T 451 may allocate resources on HOST-T 451 for the source VM1-S to generate VM1-T 452.

[0094] Reference now Figure 4B, checkpoint 209 can further obtain an AI accelerator state frame 421. The difference between the AI ​​accelerator state frame 421 and the checkpoint frame 420 is that the AI ​​accelerator state frame 421 captures information inside the AI ​​accelerator 410. The captured content of the AI ​​accelerator state frame may include the contents of one or more registers inside the AI ​​accelerator, the contents of the secure memory, and the contents of the non-secure memory, including, for example, AI models, kernels, intermediate reasoning calculations, etc. The AI ​​accelerator state frame 421 can be obtained synchronously with the checkpoint frame 420, so that the information obtained by the AI ​​accelerator state frame 421 is "fresh" (current) relative to the most recent checkpoint frame 420 of the VM1-S to be migrated and its associated virtual function (mapping the allocation of AI accelerator 410 resources to the virtual machine, such as VM1-S). In an embodiment, the AI ​​accelerator state frame 421 can be obtained after the checkpoint frame 420 and after the AI ​​task to be executed of the executed application 403 has stopped. Such an embodiment avoids the AI ​​accelerator state frame 421 storing the state of the AI ​​accelerator, which corresponds to a partially ongoing process or thread that may be difficult to reliably restart after migration.

[0095] The AI ​​accelerator state frame 421 may include the following information. The following information is provided as an example and not as a limitation. A person skilled in the art may add or delete information in the table for a particular system installation. During the migration of VM1-S, the AI ​​accelerator state frame 421, the hash of the frame, and the date and timestamp of the frame may be digitally signed with the private key of the AI ​​accelerator 410 or the private key of the virtual machine VM1-S before the frame is transmitted to the hypervisor 455 of the target host HOST-T 451. When it is time to migrate the virtual machine VM1-S and the virtual function, the user may select a specific AI accelerator state frame 421, or may generate a frame 421 in response to the selection of the checkpoint frame 420 and in response to receiving an instruction to migrate the source VM1-S to the target HOST-T 451. In an embodiment, the migration defaults to using the AI ​​accelerator state frame 421 associated with the most recent checkpoint frame 420. In an embodiment, during the migration of the source VM1-S, the AI ​​accelerator state frame 421, the hash of the AI ​​accelerator state frame 421, and the date and time stamp of the AI ​​accelerator state frame 421 may be digitally signed before the AI ​​accelerator state frame 421 is transmitted from the source VM1-S to the hypervisor 455 of the target host HOST-T 451.

[0096] When the hypervisor 455 of the target host receives the AI ​​accelerator state frame 421, the hypervisor may use the public key of VM1-S, or in an embodiment, the public key of the AI ​​accelerator 410 of VM1-S, to decrypt the AI ​​accelerator state frame 421 to verify that the date and time stamps fall within a predetermined time window, and to verify the hash of the AI ​​accelerator state frame 421. The check of the date and time stamps verifies the freshness of the AI ​​accelerator state frame 421. If the hypervisor 455 of the target HOST-T 451 verifies the AI ​​accelerator state frame 421, the hypervisor 455 of HOST-T 451 may copy the contents of the AI ​​accelerator state frame to the AI ​​accelerator ACC2 460 on VM1-T 452.

[0097]

[0098] Figure 5A A method 500 of virtual machine migration of a data processing system with an AI accelerator using checkpoint authentication in a virtualized environment from the perspective of a source hypervisor hosting a virtual machine to be migrated according to an embodiment is shown. The method 500 can be practiced on a source virtual machine to be migrated to a target host, such as VM1-S, and a target host such as HOST-T451, as a migrated virtual machine VM1-T.

[0099] In operation 600, the logic of VM1-S may determine whether to store a checkpoint frame 420 of VM1-S that is running an application 403 that utilizes one or more artificial intelligence (AI) accelerators, such as ACC1 410. The checkpoint frame 420 includes a snapshot of VM1-S, including the application 403, the execution threads of the application, the scheduler 209A containing the execution threads, the memory allocated by VM1-S related to the application, and the mapping of resources of the one or more AI accelerators to the virtual functions of VM1-S, as described above with reference to Figure 4A In an embodiment, optionally, generating the checkpoint frame 420 may also trigger obtaining the AI ​​accelerator status frame 421. In an embodiment, the AI ​​accelerator status frame 421 may be generated and stored after one or more AI tasks associated with the application 403 are paused or stopped in the following operation 800. Figure 6 Operation 600 is described in detail.

[0100] In operation 700, VM1-S may determine whether to migrate VM1-S. The decision may be based on receipt of a user command, such as a user command from an administrator or an end user. In an embodiment, the decision to migrate VM1-S may be based on an abnormality or failure threshold above a threshold. Figure 7 Operation 700 is described in detail.

[0101] In operation 800, in response to receiving a command to migrate VM1-S, the application, and the virtual functions for the associated AI accelerator to the target host 451, and in response to receiving a selection of a checkpoint frame 420 to be used when performing the migration, the checkpointer 209C records the state of one or more executing AI tasks associated with the running application, and then stops or pauses the one or more executing AI tasks. VM1-S then begins process 800 for migrating VM1-S and the virtual functions to the target host. Figure 8 Operation 800 is described.

[0102] In operation 900, in response to VM1-S receiving a notification from the hypervisor 455 of the target host 451 that the hypervisor 455 has successfully verified the checkpoint 420 and the migration is complete, the hypervisor of the source host instructs the hypervisor 455 on the target host 451 to restart the migrated application and the tasks recorded in VM1-T. Optionally, VM1-S performs post-migration cleanup of VM1-S and another AI accelerator associated with VM1-S through a virtual function. Fig. 9 Method 900 is described. Method 500 ends.

[0103] Figure 5B A method 550 of virtual machine migration on a data processing system with an AI accelerator using AI accelerator state verification in a virtualized environment is shown from the perspective of a source hypervisor hosting a source virtual machine to be migrated according to an embodiment. The method 550 can be implemented on a source virtual machine to be migrated to a target host, such as VM1-S, and a target host such as HOST2 451, as a migrated virtual machine VM1-T.

[0104] In operation 800, in response to receiving a command to migrate VM1-S, the application running on VM1-S, and the virtual functions of the associated AI accelerator to the target host 451, and in response to receiving a selection of a checkpoint frame 420 to be used when performing the migration, the checkpointer 209C records the state of one or more executing AI tasks associated with the running application, and then stops or pauses the one or more executing AI tasks. VM1-S then begins process 800 for migrating VM1-S and the virtual functions to the target host. Figure 8 Operation 800 is described.

[0105] In operation 551, after selecting the checkpoint frame 420, VM1-S generates or selects a state frame of the AI ​​accelerator 421 associated with the virtual function of VM1-S. Figure 4BThe AI ​​accelerator state frame 421 is described. A hash of the AI ​​accelerator state frame 421 is generated, a date and time stamp of the AI ​​accelerator state frame 421 is generated, and the AI ​​accelerator state frame 421, the hash, and the date and time stamp are digitally signed using a private key of VM1-S, or in an embodiment, using a private key of the AI ​​accelerator 410 associated with the virtual function that maps the AI ​​resource to VM1-S. The digitally signed AI accelerator state frame 421 is transmitted to the hypervisor 455 of the target host 451.

[0106] In operation 900, in response to receiving a notification from the hypervisor 455 of the target host 451 that the checkpoint frame 420 and the AI ​​accelerator status frame 421 are successfully verified and the migration is complete, the hypervisor 455 on the target host 541 restarts the application and the recorded AI task within the migrated virtual machine VM1-T. Optionally, VM1-S can perform post-migration cleanup. Operation 900, including post-migration cleanup of VM1-S and one or more AI accelerators associated with VM1-S by virtual functions, is described below with reference to Fig. 9 The method 550 ends.

[0107] Figure 6 A method 600 of generating a checkpoint frame for use in a virtual machine migration method with checkpoint authentication in a virtualized environment according to an embodiment is shown from the perspective of a source hypervisor hosting a virtual machine to be migrated.

[0108] In operation 601, the hypervisor 405 in the host 401 monitors the status of the source virtual machine (e.g., VM1-S), the network status, the AI ​​accelerator status, and the job completion progress.

[0109] In operation 602, it is determined whether a time increment for generating a checkpoint frame 420 has expired. The time increment may be set by a user or administrator, and may be adjusted dynamically depending on circumstances. In an embodiment, the user adjusts the time increment, such as in anticipation of the need to migrate VM1-S, such as if an application running on VM1-S is not making sufficient progress, or for other reasons. In an embodiment, the time increment is fixed. In an embodiment, the time increment is dynamically increased or decreased relative to the frequency of failures or the lack of failures, such that checkpoint frames 420 are generated more frequently if failures increase, or less frequently if failures decrease. If a checkpoint frame 420 should be generated, the method 600 continues at operation 605, otherwise the method 600 continues at operation 603.

[0110] In operation 603, it is determined whether an exception or fault has occurred. The fault counter can be configured to have different importance for one or more different types of faults. Processor exceptions are far more important than network faults in a network that supports retrying transmission or reception failures, for example. Therefore, processor faults can trigger the generation of checkpoint frames 420 with a count lower than the network fault count. If the exception or fault occurs above the fault count configured for the exception or fault type, the method 600 continues at operation 605, otherwise the method 600 continues at operation 604.

[0111] In operation 604, it is determined whether the job progress is less than a threshold process percentage completed. In an embodiment, a job process can have multiple types of job process counters. For example, each job process counter type can be triggered by calling a specific source code or by calling a specific AI function within the AI ​​accelerator, such as a job process counter for training an AI model or a counter for AI reasoning. The counter can be based on expected execution time relative to actual execution time or other measurements. If the job process counter indicates that the progress is less than a threshold percentage of the process counter type, method 600 continues at operation 605, otherwise method 600 ends.

[0112] In operation 605, VM1-S generates a checkpoint frame 420 of VM1-S, a running application, and maps AI accelerator resources to virtual functions of VM1-S.

[0113] In operation 606 , optionally, an AI accelerator status frame 421 may be generated after the checkpoint frame 420 is generated. The method 600 ends.

[0114] Figure 7 A method 700 of determining whether to migrate a virtual machine of a data processing system having an AI accelerator with checkpoint authentication and / or AI accelerator status verification in a virtualized environment according to an embodiment is shown from the perspective of a source hypervisor hosting the virtual machine to be migrated.

[0115] In operation 701 , a flag indicating whether to migrate a virtual machine (VM) is set to false.

[0116] In operation 702, it is determined whether the VM logic has received a user command to migrate the VM. In an embodiment, the migration command may originate from a user of the VM who may be monitoring the progress of the executing AI application. The reasons why the user may choose to migrate the VM may be known in the art: for example, the process is not making enough progress as expected, a particular host is heavily loaded or resource-limited and is causing the lack of progress, etc. If a user command to migrate the VM is received, the method 700 continues at operation 705, otherwise the method 700 continues at operation 703.

[0117] In operation 703, it may be determined whether a command to migrate the VM has been received from an administrator. The administrator may periodically monitor the load on the server, the processes of one or more applications, and the available resource levels. The administrator may choose to send a migration command in response to a user request or at the administrator's discretion. If the administrator issues a command to migrate the VM, method 700 continues at operation 705, otherwise method 700 continues at operation 704.

[0118] In operation 704, it can be determined whether the count of the exception or fault type has exceeded a threshold amount. There can be different thresholds for different types of faults. For example, the count of processor exceptions can be very low, and the count of network faults can be much higher than processor faults before triggering automatic migration based on the fault count. In an embodiment, a notification can be sent to an administrator to suggest migrating the VM based on the detected fault rather than automatically initiating the migration of the VM based on the automatically detected condition. If any type of fault or exception occurs more times than the threshold associated with the fault or exception type, the method 700 continues at operation 705, otherwise the method 700 ends.

[0119] In operation 705, the migration flag is set to true. A selection of a checkpoint for migration is also received. In the case of a user command or an administrator command to initiate the migration, the command may also include a checkpoint frame 420 for the migration. In the case of an automatically initiated migration command, the checkpoint frame 420 may be automatically generated, or the most recent checkpoint frame 420 may be selected. In an embodiment, if the most recently stored checkpoint frame 420 is older than a threshold amount of time, a new checkpoint frame 420 is generated.

[0120] In operation 706, optionally, an AI accelerator status frame 421 may be generated. In the case of automatically generating a migration command, based on the fault condition, the AI ​​accelerator status frame 421 may be automatically generated and may be used with the migration. If the AI ​​accelerator status frame is selected or generated, the method 550 ( Figure 5B ). Otherwise, execute method 500 ( Figure 5A ). Method 700 ends.

[0121] Figure 8 A method 800 of migrating a virtual machine of a data processing system having an AI accelerator with checkpoint authentication in a virtualized environment from the perspective of a source hypervisor hosting the virtual machine to be migrated is shown according to an embodiment.

[0122] In operation 801, a selection of a target (destination) server, such as host 451, which will host a virtual machine being migrated, such as VM1-S, is received.

[0123] In operation 802, one or more running AI tasks of an application running on VM1-S are stopped or paused. In an embodiment, one or more of the running AI tasks are allowed to complete, while others are paused or stopped.

[0124] In operation 803, the selected checkpoint frame 420 is transmitted to the target host 451. The hypervisor 405 of VM1-S waits for a response from the target host that the verification of the signature, date and time stamp, and hash of the checkpoint frame 420 has been verified.

[0125] In operation 804, the hypervisor 405 or driver 209 records the AI ​​applications running on VM1-S and any associated unfinished AI tasks, and all unfinished AI tasks are stopped.

[0126] In operation 805, VM1-S hypervisor 405 sends the recorded status of the unfinished AI job to hypervisor 455 of target host 451. Method 800 ends.

[0127] Fig. 9 A method 900 of performing post-migration cleanup of a source virtual machine after migrating a virtual machine having a data processing system with an AI accelerator with checkpoint authentication in a virtualized environment according to an embodiment is shown.

[0128] In operation 901, the hypervisor 405 of the source virtual machine (VM1-S) receives a notification from the hypervisor 455 of the target host 451 that the signature, date and timestamp, and hash of the checkpoint frame 420 have been authenticated. In an embodiment, the notification may also include an indication that the signature, date and timestamp, and hash of the AI ​​accelerator status frame 421 have been authenticated. The notification may further indicate that the migration of VM1-S to the target host 451 is complete, and the application and unfinished AI tasks have been restarted at the target host 451 as the migrated virtual machine of VM1-T.

[0129] In operation 902, the hypervisor 405 and / or the driver 404 of the source host 401 may erase at least the secure memory of the AI ​​accelerator used by the source VM1-S. The hypervisor 405 and / or the driver 404 may also erase the memory used by the application on VM1-S, which makes a call to an API or driver that uses the AI ​​accelerator associated with the application via a virtual function associated with VM1-S.

[0130] In operation 903 , the hypervisor 405 of the source host 401 may deallocate resources of VM1 -S, including deallocating AI accelerator resources used by VM1 -S and associated with the virtual function that mapped the AI ​​accelerator resources to VM-S. The method 900 ends.

[0131] Fig.10 A method 1000 of migrating a virtual machine having an AI accelerator with checkpoint authentication in a virtualized environment according to an embodiment is shown from the perspective of a target hypervisor of a host that will host the migrated virtual machine.

[0132] In operation 1001, the hypervisor 455 of the target host 451 receives a checkpoint frame 420 from a source virtual machine, such as VM1-S, associated with a virtual function that maps AI processor resources to VM1-S. The hypervisor 455 also receives a request to host VM1-S on the target host 451 as a migrated virtual machine (VM1-T).

[0133] In operation 1002, the hypervisor 455 on the host 451 calculates and reserves resources for generating VM1-S as VM1-T on the host 451. The hypervisor 455 allocates and configures resources for hosting VM1-S and its associated virtual functions based on the received checkpoint frame 420.

[0134] In operation 1003, the hypervisor 455 at the target host 451 receives the data frame and acknowledges to the hypervisor 405 at the source host 401 that the data frame is received as part of migrating VM1-S to VM1-T. The hypervisor 455 stores the received frame on the host 451 so that the hypervisor 455 can generate VM1-T.

[0135] In operation 1004, optionally, the hypervisor 455 at the target host 451 receives the signed AI accelerator state frame 421 from the hypervisor 505 at the source host 401. The hypervisor 455 decrypts the signed AI accelerator frame 421 using the public key of VM1-S, or using the accelerator public key of VM1-S. The hypervisor 455 verifies the date and time stamp in the frame 421, and verifies the digest of the frame 421. If the signed AI accelerator state frame 421 is successfully verified, the hypervisor 455 loads the data from the AI ​​accelerator state frame 421 into the AI ​​accelerator and configures the AI ​​accelerator according to the data in the AI ​​accelerator state frame 421.

[0136] In operation 1005, the hypervisor 455 of the target host 451 receives the recorded status of the unfinished AI tasks of the application running on VM1-S. VM1-T restarts the application and the unfinished AI tasks on VM1-T.

[0137] In operation 1006 , the hypervisor 455 on the target host 451 sends a notification to the source hypervisor 405 on the source host 401 , indicating that the restart of the application and the unfinished AI task was successful and the migration of VM1 -S to VM1 -T was successful.

[0138] Some portions of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of operations leading to a desired result. The operations are those requiring physical manipulations of physical quantities.

[0139] It should be remembered, however, that all of these and similar terms are associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise clear from the foregoing discussion, it should be understood that throughout the description, discussions using terms such as those set forth in the following claims refer to the actions and processes of a computer system or similar electronic computing device to process and transform data represented as physical (electronic) quantities in the computer system's registers and memories into other data similarly represented as physical quantities in the computer system memories or registers or other such information storage, transmission or display devices.

[0140] Embodiments of the present disclosure also relate to an apparatus for performing the operations herein. Such a computer program is stored in a non-transitory computer-readable medium. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). For example, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium (e.g., a read-only memory ("ROM"), a random access memory ("RAM"), a disk storage medium, an optical storage medium, a flash memory device).

[0141] The processes or methods depicted in the foregoing figures may be performed by processing logic including hardware (e.g., circuits, dedicated logic, etc.), software (e.g., embodied on a non-transitory computer-readable medium), or a combination of both. Although the processes or methods are described above based on some sequential operations, it should be understood that some of the operations described may be performed in a different order. In addition, some operations may be performed in parallel rather than sequentially.

[0142] The embodiments of the present disclosure are not described with reference to any particular programming language. It should be appreciated that a variety of programming languages ​​can be used to implement the teachings of the embodiments of the present disclosure as described herein.

[0143] In the foregoing description, embodiments of the present disclosure have been described with reference to specific exemplary embodiments of the present disclosure. It will be apparent that various modifications may be made thereto without departing from the broader spirit and scope of the present disclosure as set forth in the appended claims. Therefore, the description and drawings are to be regarded as illustrative rather than restrictive.

Claims

1. A method for migrating a source virtual machine (VM-S), wherein the source virtual machine is executing an application accessing a virtual function of an artificial intelligence (AI) accelerator, wherein the method include: In response to receiving a command to migrate the VM-S and virtual functions, and receiving a selection to checkpoint the VM-S and virtual functions used in performing the migration: Record and then stop one or more AI tasks being executed by the application. Generate or select the state of an AI accelerator associated with a virtual function, and Transfer the checkpoints and states of the AI ​​accelerator to the hypervisor of the target host to generate a migrated target virtual machine (VM-T); as well as In response to receiving notification that the target host has validated the checkpoint and the AI ​​state, and has generated and configured resources for generating the VM-T, and has loaded the AI ​​accelerator on the target host with data from the AI ​​accelerator state: Migrate VM-S and virtual functions to VM-T.

2. The method according to claim 1, further comprising: include: In response to receiving a notification that VM-T has restarted the application and the AI ​​task, performing post-migration cleanup of VM-S and the virtual function, including: Erasing at least the secure memory of the AI ​​accelerator, including any AI inferences, AI models, intermediate results of secure computations, or portions thereof; and The memory of the VM-S associated with the virtual function and any calls to the virtual function by the application are erased.

3. The method according to claim 1, further comprising: include: Checkpoints of the states of the virtual functions and the VM-S are stored in a storage of multiple checkpoints of the VM-S, wherein each checkpoint of the VM-S includes states of resources of the VM-S, states of applications, and states of virtual functions associated with resources of the AI ​​accelerator.

4. The method according to claim 3, wherein the checkpoint further include: A record of one or more AI tasks being executed; Configuration information of resources within the AI ​​accelerator communicatively coupled to the VM-S; A snapshot of the memory of the VM-S, including communication buffers and virtual function scheduling information within one or more AI accelerators; and The date and time stamp of the checkpoint.

5. The method according to claim 1, wherein the state of the AI ​​accelerator is generated include: Store the date and timestamp of the state in the AI ​​Accelerator State; Storing, in the AI ​​accelerator state, contents of a memory within the AI ​​accelerator, including one or more registers associated with a processor of the AI ​​accelerator, and a cache, queue, or pipeline of pending instructions to be processed by the AI ​​accelerator; and Generates a hash of the AI ​​accelerator state and digitally signs the state, hash, and date and timestamp.

6. The method according to claim 5, in, The AI ​​accelerator state further includes one or more register settings indicating one or more other AI accelerators in the AI ​​accelerator cluster with which the AI ​​accelerator is configured to communicate.

7. The method of claim 1, wherein the signature and freshness of the AI ​​accelerator state is verified include: Decrypt the signature of the AI ​​state using VM-S’s public key; Determine that the date and time stamp of the AI ​​accelerator status is within the threshold date and time range; as well as A hash that verifies the AI ​​accelerator status.

8. A computer readable medium programmed with executable instructions that, when executed by a processing system having at least one hardware processor communicatively coupled to an artificial intelligence (AI) processor, implements the operation of migrating a source virtual machine (VM-S) of any one of claims 1 to 7 that is executing an application that is accessing a virtual function of an artificial intelligence (AI) accelerator of the system.

9. A system comprising at least one hardware processor coupled to a memory programmed with instructions, which, when executed by the at least one hardware processor, cause the system to implement operations as described in any one of claims 1-7 for migrating a source virtual machine (VM-S) that is executing an application that accesses a virtual function of an artificial intelligence (AI) accelerator.

10. A computer program product, comprising a computer program, which, when executed by a processor, implements the operation of migrating a source virtual machine (VM-S) executing an application accessing a virtual function of an artificial intelligence (AI) accelerator as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Virtual machine homogenization to enable migration across heterogeneous computers

    CN102193824A

  • Software virtual machine for acceleration of transactional data processing

    CN103930875A