Attestable migration history for protected virtual machines
Patent Information
- Application Number
- PCT/CN2025/085610
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025085610_01102026_PF_FP_ABST
Abstract
Description
ATTESTABLE MIGRATION HISTORY FOR PROTECTED VIRTUAL MACHINESBACKGROUND
[0001] Many information processing systems employ disk encryption to protect data at rest. However, data in memory may be in plaintext and vulnerable to attacks. Attackers can use a variety of techniques including software and hardware-based bus scanning, memory scanning, hardware probing etc. to retrieve data from memory. This data from memory could include sensitive data; for example, privacy-sensitive data, intellectual property-sensitive data, cryptographic keys used for file encryption or communication, and the like. Moreover, a current trend in computing is the movement of data and enterprise workloads into the cloud by utilizing virtualization-based hosting services provided by cloud service providers (CSPs) . This further exacerbates the risk of exposure of the data.BRIEF DESCRIPTION OF THE DRAWINGS
[0002] The invention may best be understood by referring to the following description and accompanying drawings that are used to illustrate embodiments. In the drawings:
[0003] Figure 1 is a block diagram illustrating migration of a protected virtual machine from a source platform to a destination platform.
[0004] Figure 2 is a block diagram illustrating an example computing system that provides isolation in virtualized systems using trust domains according to one implementation.
[0005] Figure 3 is a block diagram illustrating another example computing system that provides isolation in virtualized systems using trust domains according to one implementation.
[0006] Figure 4 is a block diagram of an example of a trust domain architecture according to one implementation.
[0007] Figure 5 is a block diagram illustrating an embodiment of migration of a source trust domain from a source platform to a destination platform.
[0008] Figures 6A-C is a block diagram illustrating a detailed example of logic involved in migration of a source trust domain from a source platform to a destination platform.
[0009] Figure 7 illustrates migration with a hardcoded static migration policy.
[0010] Figure 8 illustrates an updatable migration policy according to embodiments.
[0011] Figure 9 illustrates rebinding to a new version of a migration trust domain according to embodiments.
[0012] Figure 10 shows an example of events that may be captured according to embodiments.
[0013] Figure 11 illustrates a migration event captured in an event log according to embodiments.
[0014] Figure 12 illustrates an attested migration log according to embodiments.
[0015] Figure 13 illustrates an attestable migration event log according to embodiments.
[0016] Figure 14 illustrates event log updating during migration according to embodiments.
[0017] Figure 15 illustrates migration event log updating during rebinding according to embodiments.
[0018] Figure 16 illustrates an example computing system according to an embodiment.
[0019] Figure 17 illustrates a block diagram of an example processor and / or System on a Chip (SoC) that may have one or more cores and an integrated memory controller according to an embodiment.
[0020] Figure 18A is a block diagram illustrating both an example in-order pipeline and an example register renaming, out-of-order issue / execution pipeline according to an embodiment.
[0021] Figure 18B is a block diagram illustrating both an example in-order architecture core and an example register renaming, out-of-order issue / execution architecture core to be included in a processor according to an embodiment.
[0022] Figure 19 illustrates examples of execution unit (s) circuitry according to an embodiment.
[0023] Figure 20 illustrates the use of a software instruction converter to convert binary instructions in a source instruction set architecture to binary instructions in a target instruction set architecture according to an embodiment.DETAILED DESCRIPTION OF EMBODIMENTS
[0024] The present disclosure relates to methods, apparatus, systems, and non-transitory computer-readable storage media to create and use an attestable migration history for a protected virtual machine (VM) . According to some examples, a system includes a plurality of processor cores and at least one non-transitory machine-readable storage medium storing a plurality of instructions to cause the plurality of processor cores to perform operations including to: determine to migrate a first protected VM from a first platform to a second protected VM on second platform; create a migration event log to indicate that the first protected VM has run on the first platform; determine to migrate the second protected VM from the second platform; use the migration event log as attestable proof to determine that the second protected VM has been migrated from the first platform; and update the migration event log to indicate that the second protected VM has run on the second platform.
[0025] As mentioned in the background section, the movement of data and enterprise workloads into the cloud by utilizing virtualization-based hosting services may exacerbate the risk of exposure to attacks. Therefore, the prospect of better security and isolation provided by embodiments may be desirable. For example, CSP customers and tenants may benefit from embodiments that may enable the operation of CSP-provided software outside of a Trusted Computing Base (TCB) of the tenant’s software. The TCB of a system refers to a set of hardware, firmware, and / or software components that have an ability to influence the trust for the overall operation of the system.
[0026] Virtualization-based hosting services may involve relocating or migrating a CSP-provided virtual machine (VM) across platforms, for example to meet desired or contracted customer service performance levels. Embodiments may be described with reference to VMs, containers, enclaves, etc.; protected, secure, trusted, isolated, etc. VMs, containers, enclaves, etc.; trusted execution environments or elements (TEEs) ; and / or a type of VM according to a specific architectural definition, such as a Trust Domain (TD) according to the Trust Domain Extension (TDX) architecture, etc., and / or with reference to details such as specific application binary interface (ABI) primitives, specific operations and sequences of operations, specific TDX implementation details, specific x86 architectural concepts (e.g., instructions, privilege level rings) , etc. but any such description does not limit embodiments to any particular type of VM or to any such details. Furthermore, descriptions and details referring to protected, secure, trusted, isolated, etc. VMs, containers, enclaves, etc.; TEEs; TDs; and / or any such terms may be applicable to implementations describable using any other such terms.
[0027] Migrating VMs across platforms within a cloud environment (e.g., to meet client Service Level Agreements (SLAs) ) may involve balancing cloud platform upgradability, patching, and other serviceability factors. From a security perspective, migration might only be performed if the destination platform is more secure than the source platform. A migration policy may define permissible platforms, physical server locations, and other acceptable service VM versions running on destination platform to which a client VM may be migrated. If the client VM was operating on a vulnerable platform prior to migration, the new platform may be secure, but the previous vulnerability could mean that the VM was compromised. Consequently, the client VM cannot be trusted. Therefore, embodiments may include the maintenance and use of protected VM migration histories for attestation and rebinding, as described below, to provide better security than existing approaches. Also, as further described below, embodiments may make it possible to rebind to a newer version of a migration VM or migrate to a newer platform with rollback prevention, and embodiments may make it possible for CSPs to avoid implementing a migration service, which may not be cost-effective and would increase the TCB of the client TD.
[0028] Figure 1 is a block diagram illustrating migration 107 of a source protected VM 105S from a source platform 100S to a destination platform 100D. The source and destination platforms may represent servers (e.g., virtualization servers like the virtualization server 200 of Figure 2) , or other computer systems similar to those disclosed herein. In general, the platforms may represent any of various different platforms on which virtual machines may run. The source platform has platform hardware 101S and the destination platform has platform hardware 101D. Examples of such hardware include cores, caches, registers, translation lookaside buffers (TLBs) , memory management units (MMUs) , a cryptographic engine (crypto) , other processor hardware, input / output devices, main memory, secondary memory, other system-level hardware, and other hardware found in processors and computer systems (e.g., as shown in the other processors and computer systems disclosed herein) .
[0029] The source platform has a source virtual machine manager (VMM) 102S and the destination platform has a destination VMM 102D. The VMMs are sometimes referred to as hypervisors. The source VMM may create and manage one or more VMs, including the source protected VM, in ways that are known in the art. Similarly, the destination VMM may create and manage, support, or otherwise interact with one or more VMs, including the destination protected VM, once it has been migrated.
[0030] The source platform has a source security component 103S and the destination platform has a destination security component 103D. In various embodiments, the source and destination security components may be security-services modules (e.g., TDX modules) , security processors (e.g., platform security processors) , processor microcode, security firmware components, or other security components. The security components may be implemented in hardware, firmware, software, or any combination thereof. The source and destination security components may be operative to provide protection or security to the source and destination protected VMs, respectively. In some embodiments, this may include protecting the protected VMs from software not within the trusted computing base (TCB) of the protected VMs, including the source and destination VMMs. Such approaches may rely on encryption, hardware isolation, or other approaches. As used herein, a protected VM may represent an encrypted, isolated, or otherwise protected VM whose state is not-accessible to the VMM.
[0031] The source and destination security components 103S, 103D may respectively use various different types of approaches known in the arts to protect the source and destination protected VMs 105S, 105D. One example of such an approach is a VM using Multi-Key Total Memory Encryption (MKTME) , which uses encryption to protect the data and code of the VM. MKTME supports total memory encryption (TME) in that it allows software to use one or more or separate keys for encryption of volatile or persistent memory. Another example is an TDX TD, which uses both encryption, integrity protection, and hardware isolation to protect the data and code of the VM against software and hardware attacks from elements outside the TCB. When MKTME is used with TDX, it provides confidentiality via separate keys for memory contents of different TDs such that the TD keys cannot be operated upon by the untrusted VMM. MKTME may be used with and without TDX. The TD is an example of an encrypted and hardware isolated VM. Another example is an encrypted VM in AMD Secure Encrypted Virtualization (SEV) , which uses encryption to protect the data and code of the VM. Yet another example is an encrypted VM in AMD SEV Secure Nested Paging (SEV-SNP) , which uses encryption and hardware isolation to protect the data and code of the VM. One more example is a secure enclave in Software Guard Extensions ( SGX) , which uses encryption, hardware isolation, replay protections, and other protections to protect code and data of a secure enclave. Each of MKTME, TDX, SEV, SEV-SNP, and SGX may provide encrypted workloads (VMs or applications) with pages that are encrypted with a key that is kept secret in hardware from the VMM, hypervisor, or host operating system. These are just a few examples. Those skilled in the art, and having the benefit of the present disclosure, will appreciate that the embodiments disclosed herein may also be applied to other techniques known in the art or developed in the future.
[0032] The source security component has a source migration component 104S, and the destination security component has a destination migration component 104D. The source and destination migration components may be operative to assist with the migration. The source and destination migration components may be implemented in hardware, firmware, software, or any combination thereof. The migration may move the source protected VM from the source platform to become a destination protected VM on the destination platform. Such migration may be performed for various different reasons, such as, for example, to help with capacity planning / load balancing, to allow a platform to be serviced, maintenance, or upgraded, to allow a microcode, firmware, or software patch to be installed on a platform, to help meet customer service-level agreement (SLA) objectives, or the like. In some embodiments, the source security component 103S and / or the source migration component 104S may also optionally be operative to “park” the source protected VM 105S as a parked source VM 106. To park the source protected VM broadly refers to stopping its execution and preserving it on the source platform (e.g., temporarily deactivating it or stopping it and putting it to rest on the source platform) . As another option, in some embodiments, the source protected VM may optionally be migrated more than once to allow cloning via live migration (e.g., in some cases if a migration policy of the source protected VM permits it and otherwise not) .
[0033] The migration 107 may either be a so-called cold migration or a so-called live migration. In a cold migration, the source protected VM 105S may be suspended or stopped for most if not all the duration of the migration and the destination protected VM may be resumed after the migration has completed. In live migration, the source protected VM may remain executing during much of the migration, and then suspended or stopped while the last part of the migration is completed (which may be shorter than network timeout for session-oriented protocols like Transmission Control Protocol / Internet Protocol (TCP / IP) . Commonly, the destination protected VM may be resumed at some point after the source protected VM has been stopped, but prior to completion of the migration. Live migration is sometimes favored for the aspect that the downtime may be reduced by allowing one of the protected VMs to continue to operate while much if not most of the migration progresses.
[0034] As discussed above, in some embodiments, the protected VMs may be TDX TDs. TDX is an Intel technology that extends Virtual Machines Extensions (VMX) and Multi-Key Total Memory Encryption (MKTME) with a type of virtual machine guest called a Trust Domain (TD) . A TD runs in a processor (e.g., central processing unit or CPU) mode which protects the confidentiality of its memory contents and its CPU state from any other software, including the hosting VMM, unless explicitly shared by the TD itself. TDX is built on top of Secure Arbitration Mode (SEAM) , which is a CPU mode and extension of VMX. The TDX module, running in SEAM mode, serves as an intermediary between the host VMM and the guest TDs. The host VMM is expected to be TDX-aware. A host VMM can launch and manage both guest TDs and legacy guest VMs. The host VMM may maintain legacy functionality from the legacy VMs perspective. It may be restricted mainly regarding the TDs it manages.
[0035] TDX may help to provide confidentiality (and integrity) for customer (tenant) software executing in an untrusted CSP infrastructure. The TD architecture, which may be a System-on-Chip (SoC) capability, provides isolation between TD workloads and CSP software, such as a VMM of the CSP. Components of the TD architecture may include memory encryption via a MKTME engine, a resource management capability such as a VMM, and execution state and memory isolation capabilities in the processor provided via a TDX-Module-managed Physical Address Meta-data Table (PAMT) and via TDX-Module-enforced confidential TD control structures. The TD architecture provides an ability of the processor to deploy TDs that leverage the MKTME engine, the PAMT, the Secure (integrity-protected) EPT (Extended Page Table) and the access-controlled confidential TD control structures for secure operation of TD workloads.
[0036] In one implementation, the tenant’s software may be executed in a TD, which may be referred to as a tenant TD and / or a tenant workload running in the TD (which may include an operating system (OS) alone along with other ring-3 applications running on top of the OS, or a VM running on top of a VMM along with other ring-3 applications, for example) . Each TD may operate independently of other TDs in the system and may use logical processor (s) , memory, input / output (I / O0 assigned by the VMM on the platform. Each TD may be cryptographically isolated in memory using at least one exclusive encryption key of the MKTME engine to encrypt the memory (holding code and / or data) associated with the TD.
[0037] In implementations, the VMM in the TD architecture may act as a host for the TDs and may have full control of the cores and other platform hardware. The VMM may assign software in a TD with logical processor (s) . The VMM, however, may be restricted from accessing the TD’s execution state on the assigned logical processor (s) . Similarly, the VMM assigns physical memory and I / O resources to the TDs, but is not privy to access the memory state of a TD due to the use of separate encryption keys enforced by the CPUs per TD, and other integrity and replay controls on memory. Software executing in a TD operates with reduced privileges so that the VMM can retain control of platform resources. However, the VMM cannot affect the confidentiality or integrity of the TD state in memory or in the CPU structures under defined circumstances.
[0038] Conventional systems for providing isolation in virtualized systems do not extract the CSP software out of the tenant’s TCB completely. Furthermore, conventional systems may increase the TCB significantly using separate chipset sub-systems that embodiments may avoid. In some cases, the TD architecture may provide isolation between customer (tenant) workloads and CSP software by explicitly reducing the TCB by removing the CSP software from the TCB. Implementations provide a technical improvement over conventional systems by providing secure isolation for CSP customer workloads (tenant TDs) and allow for the removal of CSP software from a customer’s TCB while meeting security and functionality requirements of the CSP. In addition, the TD architecture is scalable to multiple TDs, which can support multiple tenant workloads. Furthermore, the TD architecture described herein is generic and can be applied to any dynamic random access memory (DRAM) , or storage class memory (SCM) -based memory, such as Non-Volatile Dual In-line Memory Module (NV-DIMM) . As such, implementations of the disclosure allow software to take advantage of performance benefits, such as NV-DIMM direct access storage (DAS) mode for SCM, without compromising platform security requirements.
[0039] Figure 2 is a schematic block diagram of a computing system 208 that provides isolation in virtualized systems using TDs, according to an embodiment. The virtualization system includes a virtualization server 200 that supports a number of client devices 218A–218C. The virtualization server includes at least one processor 209 (also referred to as a processing device) that executes a root VMM 202. The root VMM may include a VMM (may also be referred to as hypervisor) that may instantiate one or more TDs 205A–205C accessible by the client devices 218A–218C via a network interface 217. The client devices may include, but are not limited to, a desktop computer, a tablet computer, a laptop computer, a netbook, a notebook computer, a personal digital assistant (PDA) , a server, a workstation, a cellular telephone, a mobile computing device, a smart phone, an Internet appliance, or any other type of computing device.
[0040] A TD may refer to a tenant (e.g., customer) workload. The tenant workload may include an OS alone along with other ring-3 applications running on top of the OS, or may include a VM running on top of a VMM along with other ring-3 applications, for example. In implementations, each TD may be cryptographically isolated in memory using a separate exclusive key for encrypting the memory (holding code and data) associated with the TD, and integrity-protecting the key against any tamper by the host software or host-controlled devices.
[0041] The processor 209 may include one or more cores 210, range registers 211, a memory management unit (MMU) 212, and output port (s) 219. Figure 2 shows processor core 210 executing root VMM 202 in communication with a PAMT 216 and secure-EPT 343 and one or more trust domain control structure (s) (TDCS (s) ) 214 and trust domain virtual-processor control structure (s) (TDVPS (s) ) 215 (TDVPS and TDVPX may be used interchangeably) . The processor 209 may be used in a system that includes, but is not limited to, a desktop computer, a tablet computer, a laptop computer, a netbook, a notebook computer, a PDA, a server, a workstation, a cellular telephone, a mobile computing device, a smart phone, an Internet appliance or any other type of computing device. In another implementation, the processor 209 may be used in a SoC system.
[0042] The computing system 208 may be a server or other computer system having one or more processors available from Corporation, although embodiments are not so limited. In one implementation, computing system 208 executes a version of the WINDOWSTMoperating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (UNIX and Linux for example) , embedded software, and / or graphical user interfaces, may also or instead be used. Thus, implementations are not limited to any specific combination of hardware circuitry and software.
[0043] The one or more processing cores 210 execute instructions of the system. The processing core 210 includes, but is not limited to, pre-fetch circuitry to fetch instructions, decode circuitry to decode the instructions, execution circuitry to execute instructions and the like. In an implementation, the computing system 208 includes a component, such as the processor 209 to employ execution units including circuitry to perform algorithms for processing data.
[0044] The virtualization server 200 includes a main memory 220 and a secondary storage 221 to store program binaries, OS drivers, etc. Data in the secondary storage 221 may be stored in blocks referred to as pages, and each page may correspond to a set of physical memory addresses. The virtualization server 200 may employ virtual memory management in which applications run by the core (s) 210, such as the TDs 205A–205C, use virtual memory addresses that are mapped to guest physical memory addresses, and guest physical memory addresses are mapped to host / system physical addresses by a MMU 212.
[0045] The core 210 may use the MMU 212 to load pages from the secondary storage 221 into the main memory 220 (which may include a volatile memory and / or a non-volatile memory) for faster access by software running on the processor 209 (e.g., on the core) . When one of the TDs 205A–205C attempts to access a virtual memory address that corresponds to a physical memory address of a page loaded into the main memory, the MMU returns the requested data. The core 210 may execute a portion of a VMM (e.g., root VMM 202) to translate guest physical addresses to host physical addresses of main memory and provide parameters for a protocol that allows the core to read, walk, and interpret these mappings.
[0046] In one implementation, processor 209 implements a TD architecture and instruction set architecture (ISA) extensions (e.g., SEAM) for the TD architecture. The SEAM architecture and the TDX-Module running in SEAM mode provides isolation between TD workloads 205A-205C and from CSP software (e.g., root VMM 202, a CSP VMM, etc. executing on the processor 209, any of which may be referred to as a VMM) . Components of the TD architecture may include memory encryption, integrity and replay-protection via an MKTME engine 213, a resource management capability referred to herein as a VMM (e.g., root VMM 202) , and execution state and memory isolation capabilities in the processor 209 provided via PAMT 216 and secure-EPT 343 and TDX module 204 and via access-controlled confidential TD control structures (e.g., TDCS 214 and TDVPS 215) . The TDX architecture provides an ability of the processor 209 to deploy TDs 205A-205C that leverage the MKTME engine 213, the PAMT 216 and secure-EPT 343, and the access-controlled TD control structures (e.g., TDCS 214 and TDVPS 215) for secure operation of TD workloads 205A-205C.
[0047] In implementations, a VMM acts as a host and has control of the cores 210 and other platform hardware. The VMM assigns software in a TD 205A-205C with logical processor (s) . The VMM, however, cannot access a TD’s execution state on the assigned logical processor (s) . Similarly, a VMM assigns physical memory and I / O resources to the TDs, but is not privy to access the memory state of the TDs due to separate encryption keys, and other integrity and replay controls on memory.
[0048] With respect to the separate encryption keys, the processor may utilize the MKTME engine 213 to encrypt (and decrypt) memory used during execution. With total memory encryption (TME) , any memory accesses by software executing on the core 210 can be encrypted in memory with an encryption key. MKTME is an enhancement to TME that allows use of multiple encryption keys. The processor 209 may utilize the MKTME engine to cause different pages to be encrypted using different MKTME keys. The MKTME engine 213 may be utilized in the TD architecture described herein to support one or more encryption keys per each TD 205A-205C to help achieve the cryptographic isolation between different CSP customer workloads. For example, when MKTME engine is used in the TD architecture, the CPU enforces by default that a TD (all pages) is to be encrypted using a TD-specific key. Furthermore, a TD may further choose specific TD pages to be plain text or encrypted using different ephemeral keys that are opaque to CSP software.
[0049] Each TD 205A-205C is a software environment that supports a software stack consisting of one or more VMMs (e.g., using virtual machine extensions (VMX) ) , one or more OSes, and / or application software (hosted by the OS) . Each TD may operate largely independently of other TDs and use logical processor (s) , memory, and I / O assigned by the VMM on the platform. Software executing in a TD operates with reduced privileges so that the VMM can retain control of platform resources; however, the VMM cannot affect the confidentiality or integrity of the TD under defined circumstances. Further details of the TD architecture and TDX are described in more detail below with reference to Figure 3.
[0050] System 208 includes a main memory 220. Main memory includes a DRAM device, a static random access memory (SRAM) device, flash memory device, or other memory device. Main memory stores instructions and / or data represented by data signals that are to be executed by the processor 209. The processor may be coupled to the main memory via a processing device bus. A system logic chip, such as a memory controller hub (MCH) may be coupled to the processing device bus and main memory. An MCH can provide a high bandwidth memory path to main memory for instruction and data storage and for storage of graphics commands, data, and textures. The MCH can be used to direct data signals between the processor, main memory, and other components in the system and to bridge the data signals between processing device bus, memory, and system I / O, for example. The MCH may be coupled to memory through a memory interface. In some implementations, the system logic chip can provide a graphics port for coupling to a graphics controller through an Accelerated Graphics Port (AGP) interconnect.
[0051] The computing system 208 may also include an I / O controller hub (ICH) . The ICH can provide direct connections to some I / O devices via a local I / O bus. The local I / O bus is a high-speed I / O bus for connecting peripherals to the memory 220, chipset, and / or processor 209. Some examples of I / O are an audio controller, firmware hub (e.g., flash BIOS) , wireless transceiver, data storage, legacy I / O controller containing user input and keyboard interfaces, a serial expansion port such as Universal Serial Bus (USB) , and a network controller. The data storage device can comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0052] With reference to Figure 3, this figure depicts a block diagram of a processor 309 suitable for the processor 209 of Figure 2, according to one implementation. In one implementation, the processor may execute a system stack 326 via a single core 210 or across several cores 210. As discussed above, the processor may provide a TD architecture and TDX to provide confidentiality (and integrity) for customer software running in the customer / tenants (i.e., TDs 205A) in an untrusted CSP’s infrastructure. The TD architecture provides for memory isolation via a PAMT 316 and secure-EPT 343, CPU state isolation that incorporates CPU key management via TDCS 314 and / or TDVPS 315, and CPU measurement infrastructure 338 for TD 205A software.
[0053] In one implementation, TD architecture provides ISA extensions (referred to as TDX) that support confidential operation of OS and OS-managed applications (virtualized and non-virtualized) . A TDX-enabled platform, such as one including processor 209, may function with multiple encrypted contexts referred to as TDs. For ease of explanation, a single TD 305 is depicted. Each TD can run VMMs, VMs, OSes, and / or applications. For example, TD 305 is depicted as hosting a VMM and two VMs. A TDX module 204 supports the TD 305.
[0054] In one implementation, the root VMM 302 may include and / or represent VMM functionality provided by software, firmware, and / or hardware to create, run, and manage one or more VMs. The VMM may create and run a VM and allocate one or more virtual processors (e.g., vCPUs) to the VM. The VM may also be referred to as guest. The VMM may allow the VM to access hardware of the underlying computing system, such as computing system 208 of Figure 2. The VM may execute a guest operating system (OS) . The VMM may manage the execution of the guest OS. The guest OS may function to control access of virtual processors of the VM to underlying hardware and software resources of the computing system. When there are numerous VMs operating on the processing device, the VMM may manage each of the guest OSes executing on the numerous guests.
[0055] TDX also provides a programming interface for a TD management layer of the TD architecture (e.g., a VMM implemented as or as part of a CSP or root VMM 302) . The VMM manages the operation of TDs 205A / B / C. While a VMM can assign and manage resources, such as CPU, memory, and I / O to TDs, the VMM is designed to operate outside of a TCB of the TDs. The TCB of a system refers to a set of hardware, firmware, and / or software component that have an ability to influence the trust for the overall operation of the system.
[0056] In one implementation, the TD architecture provides a capability to protect software running in a TD (e.g., TD 205A) . As discussed above, components of the TD architecture may include memory encryption via a TME engine having multi-key extensions to TME (e.g., MKTME engine 213 of Figure 2) , a software resource management layer (e.g., VMM) , and execution state and memory isolation capabilities in the TD architecture.
[0057] The PAMT 316 and secure-EPT 343 are structures, such as tables, managed by the TDX-Module to enforce assignment of physical memory pages to executing TDs, such as TD 205A. The processor 209 also uses the PAMT and secure-EPT to enforce that the physical addresses referenced by software operating as a tenant TD or the VMM cannot access memory not explicitly assigned to it. The PAMT and Secure-EPT enforces the following properties. First, software outside a TD should not be able to access (read / write / execute) in plain-text any memory belonging to a different TD (this includes the VMM) . Second, memory pages assigned via the PAMT and Secure-EPT to specific TDs, such as TD, should be accessible from any processor in the system (where the processor is executing the TD that the memory is assigned to) .
[0058] The PAMT and Secure-EPT structure is used to hold meta-data attributes for each page (e.g., 4KB page) of memory. Additional structures may be defined for additional page sizes (2MB, 1GB) . The meta-data for each 4KB page of memory is direct indexed by the physical page address. In other implementations, other page sizes may be supported by a hierarchical structure (like a page table) . A 4KB page referenced in the PAMT and Secure-EPT can belong to one running instance of a TD. 4KB pages referenced in the PAMT and Secure-EPT can either be valid memory or marked as invalid (hence could be I / O for example) . In one implementation, each TD instance includes one page holding a TDCS for that TD.
[0059] In one implementation, the PAMT and Secure-EPT is aligned on a 4KB boundary of memory and occupies a physically contiguous region of memory protected from access by software after platform initialization. In an implementation, the PAMT is a micro-architectural structure and cannot be directly accessed by software. The PAMT and Secure-EPT may store various security attributes for each 4KB page of host physical memory.
[0060] The PAMT and Secure-EPT may be enabled when TDX is enabled in the processor (e.g., via CPUID-based enumeration) . Once the PAMT and Secure-EPT is enabled, the PAMT and Secure-EPT can be used by the processor to enforce memory access control for all physical memory accesses initiated by software, including the VMM. In one implementation, the access control is enforced during the page walk for memory accesses made by software. Physical memory accesses performed by the processor to memory that is not assigned to a tenant TD or VMM may fail (e.g., with Abort page semantics) .
[0061] Figure 4 is a block diagram depicting an example computing system implementing TD architecture 400. A first environment 410 is one where the tenant trusts the CSP to enforce confidentiality and does not implement the TD architecture of implementations of the disclosure. This environment uses a CSP VMM-managed TCB 402. This environment may include a CSP VMM 455 managing a CSP VM 414 and / or one or more legacy tenant VMs 416A, 416B. In this case, the tenant VMs are managed by the CSP VMM that is in the VM’s TCB. In implementations of the disclosure, the tenant VMs, may still leverage memory encryption via TME or MKTME in this model.
[0062] The other type of environment is a TD where the tenant does not trust the CSP to enforce confidentiality and thus relies on a TDX module 424 and the CPU with a TD architecture of implementations. This type of TD is shown in two variants as TD2 405A and TD3 405B. The TD2 405A is shown with a virtualization mode (such as VMX) being utilized by the tenant VMM (non-root) 454 running in TD2 to manage tenant VMs 451, 452. The TD3 does not include software using a virtualization mode, but instead runs an enlightened OS 435 in the TD3 405B directly. TD2 and TD3 are tenant TDs having a hardware-enforced TCB 404 as described in implementations of the disclosure. In one implementation, TD2 and / or TD3 may be the same as TD 205A described with respect to Figure 2.
[0063] The VMM 422 manages the life cycle of the environment 410 as well as the TD2 405A and TD3 405B, including allocation of resources. However, the VMM 422 is not in the TCB for the TD2 and the TD3. The TD architecture 400 does not place any architectural restrictions on the number or mix of TDs active on a system. However, software and certain hardware limitations in a specific implementation may limit the number of TDs running concurrently on a system due to other constraints.
[0064] Example embodiments will now be described in conjunction with migration of a TD protected by TDX. The TD is an example embodiment of a protected, hardware-isolated VM. However, those skilled in the art, and having the benefit of the present disclosure, will appreciate that the embodiments disclosed herein may also be applied to migration of other types of protected VMs, such as, for example, those protected by MKTME, SEV, SEV-SNP, SGX, or other techniques.
[0065] Figure 5 is a block diagram illustrating an embodiment of migration 507 of a source trust domain (TD) 505S from a source platform 500S to a destination platform 500D. The platforms may be of the types previously described for the platforms of Figure 1 (e.g., servers, computer systems, etc. ) . The source platform has platform hardware 501S and the destination platform has platform hardware 501D. This hardware may be of the types previously described for the hardware of Figure 1 (e.g., cores, caches, etc. ) .
[0066] The source platform has a source TDX-aware host VMM 502S and the destination platform has a destination TDX-aware host VMM 502D. The VMMs are sometimes referred to as hypervisors. The source and destination VMMs are TDX aware in that they are respectively able to interact with a source TDX module 503S (which is an example embodiment of the source security component 103S) and a destination TDX module 503D (which is an example embodiment of the destination security component 103D) . One way in which they may be TDX-aware and interact with the TDX modules is through being operative to use respective TDX Module host-side application programming interfaces (APIs) 556S, 556D to interact with the TDX modules 503S, 503D. These TDX Module host-side APIs represent example embodiments of primitives or command primitives that the VMMs may use to interact with the TDX modules. Specific examples of the APIs according to a specific example embodiment will be provided further below. Other primitives or command primitives (e.g., instructions of an instruction set, API commands, writes to memory-mapped input and / or output (MMIO) registers, etc. ) may be used by VMMs to interact with other security components in other embodiments.
[0067] The source TDX-aware host VMM 502S and the source TDX module 503S may manage, support, or otherwise interact with the source TD 505S. Likewise, the destination TDX-aware host VMM 502D and the destination TDX module 503D may manage, support, or otherwise interact with the destination TD 505D. The source TD may represent software operating in a CPU mode (e.g., SEAM) that excludes the TDX-aware host VMM 502S and untrusted devices of the source platform from its operational TCB for confidentiality. The source TD may have a TDX-enlightened operating system (OS) 558S and optionally one or more applications 560S and drivers 559S. Likewise, the destination TD may have a TDX-enlightened OS 558D and optionally one or more applications 560D and drivers 559D. The OS is enlightened or paravirtualized in that it is aware of it running in a TD. In other embodiments, VMMs may optionally be nested within TDs. The operational TCB for the source TD may include one or more CPUs on which the source TD runs, the source TDX module 503S, the TDX-enlightened OS, and the applications. The source TD’s resources may be managed by the source TDX-aware host VMM, but its state protection may be managed by the source TDX module 503S and may not be accessible to the source TDX-aware host VMM.
[0068] Optionally, the source and destination VMMs may also manage, support, or otherwise interact respectively with one or more source VMs 549S and one or more destination VMs 549D. These VMs are not protected VMs in the same way that the TDs are protected VMs, such as by not having the same level of confidentiality (e.g., not being encrypted) , not being hardware isolated to the same extent (e.g., from the VMMs) , and so on.
[0069] The source platform has a source TDX module (e.g., representing an example embodiment of a source security component) and the destination platform has a destination TDX module (e.g., representing an example embodiment of a destination security component) . The TDX modules may be implemented in hardware, firmware, software, or any combination thereof. The source and destination TDX module may be operative to provide protection or security to the source and destination TDs, respectively. In some embodiments, this may include protecting the TDs from software not within the trusted computing base (TCB) of the TDs, including the source and destination TDX-aware VMMs. The various types of protections described elsewhere herein may optionally be used or a subset or superset thereof. In some embodiments, this may include assigning different host key identifiers (HKIDs) to different components. When MKTME is activated, the HKID may represent a key identifier for an encryption key used by one or more memory controller on the platform. When TDX is active, the HKID space may be partitioned into a CPU-enforced space and a VMM-enforced space. In some embodiments, the source migration service TD may be assigned an HKID 1, the source TD may be assigned an HKID 2, the destination migration service TD may be assigned an HKID 3, and the destination TD may be assigned an HKID 4, where HKID 1, HKID 2, HKID 3, HKID 4 are first through fourth HKIDs that are all different. Other technologies for other protected VMS like SEV and SEV-SNP also assign different key identifiers and different keys to different protected VMs.
[0070] The source and destination TDX-aware host VMMs via the source and destination TDX modules may perform operations to respectively export or import contents of the source TD. One challenge with migrating the source TD 505S (as well as other types of protected VMs in general) is maintaining security. For example, the source TD may run in a central processing unit (CPU) mode that protects the confidentiality of its memory contents and its CPU state from any other platform software, including the source TDX-aware host VMM. Similarly, the confidentiality of its memory contents of the source TD and its CPU state should also be protected while allowing the source TDX-aware host VMM to migrate source TD to the destination platform.
[0071] The source TDX module has a source migration component 504S, and the destination TDX module has a destination migration component 504D. The source and destination TDX modules and / or their respective source and destination migration components may be operative to perform the various operations described further below to assist with the migration. The migration components may be implemented in hardware, firmware, software, or any combination thereof. The migration may move the source TD 505S from the source platform to become the destination TD 505D on the destination platform. The source and destination TDX-aware host VMMs and existing (e.g., untrusted) software stack may be largely responsible for migrating the encrypted content of the source TD once it has been encrypted and otherwise protected. The migration may either be a cold migration or a live migration as previously described. In some embodiments, the source TDX module 503S and / or the source migration component 504S may also optionally be operative to “park” the source TD 505S as previously described.
[0072] In some embodiments, an optional source migration service TD 561S and an optional destination migration service TD 561D may also be included to assist with the migration 507. The migration service TDs are example embodiments of migration agents or engines. The migration service TDs may also be referred to herein as migration TDs (MigTDs) . In some embodiments, the migration service TDs may be example embodiments of service TDs per TDX that have migration logic to perform functions associated with the migration. As shown, in some embodiments, the source migration service TD may have migration service and policy evaluation logic 562S and a TD virtual firmware (TDVF) -shim logic 563S. Likewise, the destination migration service TD may have migration service and policy evaluation logic 562D and a TDVF-shim logic 563D. The logics 562S, 562D may be operative to evaluate and enforce a TD migration policy of the source TD. The TD migration policy may represent a TD tenant indicated set of rules that govern which platform security levels the TD is allowed to be migrated to, in some cases whether it may be replicated and if so how many times, or other migration policy. For example, the migration service and policy evaluation logic may be operative to cooperate to evaluate potential migration source platforms and destination platforms based on adherence to a TD migration policy of the source TD. If the TD migration policy permits a migration, then the migration service TDs may cooperate to securely transfer a migration capable key from the source platform to the destination platform which will be used to encrypt or protect contents of the source TD during the migration. In some embodiments, they may represent architecturally defined or architectural TDs. In some embodiments, they may be within the TCB of the TDs. In some embodiments, they may be more trusted by the TDs than the respective VMMs.
[0073] The TDX modules 503S, 503D and / or their migration components 504S, 504D may implement an architecturally defined TDX migration feature implemented on top of the SEAM VMX extension to protect the contents of the source TD during migration of the source TD, such as in an untrusted hosted cloud environment. Typically, the source TD may be assigned a different key ID, and in the case of TDX a different ephemeral key, on the destination platform to which the source TD is migrated. An extensible TD migration policy may be associated with the source TD that is used to maintain the source TD’s security posture. The TD migration policy may be enforced in a scalable and extensible manner by the source migration service TD or migration TD (MigTD) 561S, which is used to provide services for migrating TDs. In some embodiments, pages, state, or other contents of the source TD may be protected while being transferred by SEAM and through the use of a migration capable key or migration key that is used for a unique migration session for the source TD.
[0074] The migration capable key may be used to encrypt information that is migrated to help provide confidentiality and integrity to the information. The migration service TDs on the source and the destination platforms may agree on the migration key based on migration policy information of the source TD. The migration key may be generated by the source migration service TD after a successful policy negotiation for the source TD. By way of example, the migration service TDs on each side may set the migration key in the source / destination TD using a generic binding write protocol. For example, the migration key may be programmed by the migration service TDs, via the host VMM, into the source and destination TDs’ TDCS using a service TD metadata write protocol. To provide security, the migration key may optionally be accessible only by the migration service TDs and the TDX Modules. In some implementations, the key strength may be 256 bits or some other number of bits (e.g., an ephemeral AES-256-GCM key) . In some embodiments, a migration stream AES-GCM protocol may be used as discussed further below which may provide that state is migrated in-order between the source and destination platform. This protocol may help to ensure or enforce the order within each migration stream. The migration key may be destroyed when a TD holding it is torn down, or when a new key is programmed.
[0075] The TDX modules 503S, 503D may respectively interact with the TDs 505S, 505D and the migration service TDs 561S, 561D through respective TDX guest-host interfaces 557S, 557D. In some embodiments, the TD migration and the source migration service TD may not depend on any interaction with the TDX-enlightened OS 558S or any other TD guest software operating inside the source TD being migrated.
[0076] As mentioned above, TDX modules (e.g., the source and destination TDX modules 503S, 503D) may implement the TDX Module host side API 556S, 556D and the TDX guest-host interfaces 557S, 557D. In some embodiments, the TDX Module host side APIs and the TDX guest-host interfaces may be accessed through instructions of an instruction set. The instructions of the instruction set represent one suitable example embodiment of command primitives that may be used to access operations provided by the TDX modules. Alternate embodiments may use other types of command primitives (e.g., commands, codes written to MMIO registers, messages provided on an interface, etc. ) .
[0077] In some embodiments, the TDX modules may be operative to execute or perform two types of instructions, namely a SEAMCALL instruction and a TDCALL instruction. The SEAMCALL instruction may be used by the host VMM to invoke one of multiple operations of the TDX Module host side API. The TDCALL instruction may be used by guest TD software (e.g., a TD enlightened OS) in TDX non-root mode to invoke one of multiple operations of the TDX guest-host interface. Multiple operations may be supported by both the SEAMCALL and TDCALL instructions. Each particular instance of these instructions may indicate a particular one of these operations to be performed. In the case of the TDX SEAMCALL and TDCALL instructions, the particular operation may be indicated by a leaf operation value in the general-purpose x86 register RAX, which may be implicitly indicated by the instructions. The leaf value broadly represents a value that selects a particular one of the operations. Each of the operations may be assigned in any desired way a unique value and the leaf operation may be given that value to select the operation. Alternate embodiments may indicate the particular operation in different ways, such as, for example, a value in another register specified or indicated by an instruction of an instruction set, a value of an immediate of an instruction of an instruction set, an opcode or other value written to a MMIO register, and so forth.
[0078] The TDX-aware host VMM may invoke or otherwise use a SEAMCALL instruction to start execution of the TDX module on the hardware thread where SEAMCALL was invoked, or the TD may invoke or otherwise use a TDCALL instruction causing the TDX module to execute on the physical logical processor and / or hardware thread where the TD virtual logical processor was executing. The TDX module may receive the instructions. The TDX module may be operative to determine the particular operations to be performed as indicated by the leaf operations or other value. Some variants of the instructions may have one or more other input parameters or operands to be used in the performance of the operations but others may not. These input parameters or operands may be in memory, registers, or other storage locations specified (e.g., explicitly specified) or otherwise indicated (e.g., implicitly indicated) by the instructions. The TDX module may include logic to understand and perform the instructions. In some embodiments, this may include a decode unit or other logic to decode the instruction and an execution unit or other logic to execute or perform the instruction and / or its indicated operations. In one example embodiment, the SEAMCALL instruction and leaf may be decoded and executed by a CPU like other instructions of an instruction set of the CPU, and its execution may transfer control to the TDX module. In such an example embodiment, the TDX module may include software that is loaded into and operates from a range-register protected memory region such that its authenticity can be verified via a CPU-signed measurement or quote, and can be protected from tampering by other software and hardware at runtime (e.g., since the region of memory may be encrypted and integrity protected using a private ephemeral key which is different from the TD keys) . The TDX module may include a dispatcher submodule that may invoke an appropriate operation (e.g., an API) indicated by the leaf function. The indicated operation may then be performed also by the CPU (e.g., one or more cores) . This is just one possible way it may be done. In other embodiments, the SEAMCALL or other primitives may be performed in hardware, software, firmware, or any combination thereof (e.g., at least some hardware and / or firmware potentially combined with software) .
[0079] The TDX modules may perform the instruction and / or operations. In some embodiments, the TDX modules may perform one or more checks. In some cases the checks may fail or an error may occur before the instruction and / or operations completes. In other cases, the checks may succeed and the instruction and / or operations may complete without an error. In some embodiments, an optional return or status code (e.g., a value) may be stored to indicate an outcome of the performance of the instruction and / or operations (e.g., whether the instruction and / or operations completed successfully or whether there was an error, such as that the TDX Module experienced a fatal condition and has been shut down) . In the case of the TDX SEAMCALL and TDCALL instructions, the return or status code may be provided in the implicit general-purpose x86 register RAX. Alternate embodiments may indicate the return or status code in other ways (e.g., in one or more flags, registers, memory locations, or other architecturally-visible storage locations) . Some instructions and / or operations may have additional output parameters or operands that may be stored in architecturally-visible storage locations while others may not.
[0080] Table 1 lists an example set of operations of a TDX Module host side API according to one example embodiment. The set of operations may be supported by a TDX module (e.g., the source TDX module 503S and the destination TDX module 503D) . This set of operation is just one example and does not limit embodiments. Many other embodiments are contemplated in which two or more of the operations may be combined together into a single operation and / or one of the operations may be divided into two or more operations and / or various of the operations may be performed differently. In some embodiments, each of these operations may be available as a different leaf operation under a SEAMCALL instruction with the SEAMCALL instruction serving as the opcode and the leaf operation as a value indicated by the SEAMCALL instruction. The table is arranged in rows and columns. Each row represents a different operation under the SEAMCALL instruction. An opcode or mnemonics column lists a unique mnemonic for each of the operations. The mnemonics are in text to uniquely identify the different instruction variants to a human reader, although it is to be understood that there may also be numerical (e.g., binary) value for each of these different operations. Particular numerical values are not provided because they may be assigned arbitrarily or at least in various different ways. An operation column lists the particular operation to be performed. The specific operations can be read from the table for each of the mnemonics. It is to be appreciated that these operations may be performed differently in different embodiments. A method according to embodiments may include performing any one or more of these operations as part of performing an instruction (e.g., a SEAMCALL instruction) or other command primitive. Table 1: TDX Module host side API operations
[0081] The TDX module supporting each of these operations may have a corresponding set of logic to perform each of these operations. Such logic may include hardware (e.g., circuitry, transistors, etc. ) , firmware (e.g., microcode or other low-level or circuit-level instructions stored in non-volatile memory) , software, or any combination thereof (e.g., at least some hardware and / or firmware potentially combined with some software) . For convenience, in the discussion below, such logic of the TDX module for each operation is also referred to by the mnemonic for that corresponding operation. While described as being separate logic for convenience and to simplify the description, it is also to be appreciated that logic of some operations may be shared by and reused by other operations. That is, the logic need not be separate but rather may optionally overlap in order to reduce the total amount of logic that needs to be included in the TDX module.
[0082] Figures 6A-C are a block diagram illustrating a detailed example of logic involved in migration of a TD from a source platform 600S to a destination platform 600D. Each block may represent logic to perform a corresponding operation at a stage of the migration. Most of the blocks are named the same or analogously to the mnemonics of Table 1, and each such block may represent logic of the TDX module (e.g., TDX module 503S if the logic is in the source platform 500S or TDX module 503D if the logic is in the destination platform 500D) that is operative to perform the corresponding operation of that mnemonic. Several of the logic with MigTD in their names are logic of the MigTD (e.g., MigTD 561S if the logic is in the source platform or MigTD 561D if the logic is in the destination platform) . Each such logic may be implemented in hardware, firmware, software, or any combination (e.g., at least some hardware and / or firmware potentially combined with software) .
[0083] The logic may be used and their operations may be performed generally at the stage of the migration at which the logic block is shown. The migration proceeds generally from left to right in time. As shown on the bottom of the illustration, the stages include, with increasing time, initially a pre-migration stage, a reservation stage, an iterative pre-copy stage, a stop and copy control state stage, a commitment stage, and finally a post-copy stage.
[0084] Initially, as shown on the top left of the illustration, a source guest TD may be built. As a legacy TD, the source VMM may call a legacy TDH. MNG. CREATE logic 664 to perform its associated TDH. MNG. CREATE TDX operation to create a source TD. A destination guest TD may be built similarly. The destination VMM may call a legacy TDH. MNG. CREATE logic 678 to perform its associated TDH. MNG. CREATE operation to create a destination TD. The destination TD may be setup as a “template” to receive the state of the source TD.
[0085] The source VMM may call a TDH. SERVTD. BIND logic 665 to perform its associated TDH. SERVTD. BIND operation to bind a MigTD to the source TD being migrated and optionally others source TDs. Similarly, a migration TD is bound to the destination TD as well. The source VMM can then build the TDCS by adding TDCS pages using another TDX module operation. The destination VMM may call a TDH. MNG. CONFIGKEY logic 679 to perform its associated TDH. MNG. CONFIGKEY operation to program the HKID and hardware-generated encryption key assigned to the TD into the MKTME encryption engines for each package.
[0086] The source VMM may call a TDH. MNG. INIT logic 667 to perform its associated TDH. MNG. INIT operation to set immutable control state. This immutable control state may represent TD state variables that may be modified during TD build but are typically not modified after the TD’s measurement is finalized. Some of these state variables control how the TD and its memory is migrated. Therefore, as shown, the immutable TD control state or configuration may be migrated before any of the TD memory state is migrated. Since the source TD is to be migrated, it may be initialized with a migration capable attribute. The source VMM may also call the TDX module to learn that the TDX module supports TD Migration. The TDX module may responsively return information about its migration capabilities (e.g., whether live migration is supported) .
[0087] The destination VMM may call a TDH. SERVTD. BIND logic 681 to perform its associated TDH. SERVTD. BIND operation of attestation and binding of the destination MigTD. The source MigTD logic 669 and the destination MigTD logics 682 may perform their TDHSERVTD. MSG operations to perform a quote-based mutual authentication (e.g., using Diffie-Hellman exchange) of the migration TDs executing on both the source and destination platforms, verify or process TD migration policy (e.g., establish TD migration compatibility per a migration policy of the source TD) , establish a secure or protected transport session or channel between the TDX modules and MigTDs, perform migration session key negotiation, generate an ephemeral migration session key, and populate a TDCS with the migration key. Control structure resources may also be reserved for the TD template to start migration.
[0088] The source MigTD logic 668 and the destination MigTD logics 682 may perform their TDH. SERVTD. MSG and TDH. MIG. STREAM. CREATE operations to transfer the migration key from the source platform to the destination platform and write the migration session key to the target TD’s TDCS. Service TD binding allows the Migration TD to access specific target TD metadata, by exchanging messages encrypted and MAC’ ed with an ephemeral binding key.
[0089] TD global immutable control state may also be migrated. The TDX module may protect the confidentiality and integrity of a guest TD global state (e.g., TD-scope control structures) , which may store guest TD metadata, and which are not directly accessible to any software or devices besides the TDX module. These structures may be encrypted and integrity-protected with a private key and managed by the TDX module API functions. Some of these immutable control state variables control how the TD and its memory is migrated. Therefore, the immutable TD control state may be migrated before any of the TD memory state is migrated. The source VMM may call a TDH. EXPORT. STATE. IMMUTABLE logic 670 to perform its associated TDH. EXPORT. STATE. IMMUTABLE operation to export the immutable state (e.g., TD immutable configuration information) . The destination VMM may call a TDH. IMPORT. STATE. IMMUTABLE logic 684 to perform its associated TDH. IMPORT. STATE. IMMUTABLE operation to import the immutable state.
[0090] TD private memory migration can happen in the in-order migration phase and out-of-order migration phase. During the in-order phase, the source VMM may call a TDH. EXPORT. PAGE logic 671 to perform its associated TDH. EXPORT. PAGE operation to export memory content (e.g., one or more pages) . In live migration, this may be done during live migration pre-copy stage while the source TD is running. The destination VMM may call a TDH. IMPORT. PAGE logic 685 to perform its associated TDH. IMPORT. PAGE operation to import memory content (e.g., one or more pages) . Live migration is not required. Cold migration may be used instead where the source TD is paused prior to exporting memory content. Also, shared memory assigned to the TD may be migrated using legacy mechanisms used by the VMM (since shared memory is accessible to the VMM freely) .
[0091] TD migration may not migrate the HKIDs. Rather, a free HKID may be assigned to the TD created on the destination platform to receive migratable assets of the TD from the source platform. All TD private memory may be protected during transport from the source platform to the destination platform using an intermediate encryption performed (e.g., using AES-GCM 256) using the TD migration key negotiated via the migration TD and SEAM. On the destination platform the memory may be encrypted via the destination ephemeral key as it is imported into the destination platform memory assigned to the destination TD.
[0092] In live migration, the destination TD may begin execution during the migration process. In a typical live migration scenario where the source TD is expected to be migrated (without cloning) , the destination TD may begin executing after the iterative pre-copy stage imports the working set of memory pages. The source VMM may call a TDH. EXPORT. PAUSE logic 672 to perform its associated TDH. EXPORT. PAUSE operation to pause the source TD. This may include checking pre-conditions and preventing the source TD from executing any more. After the source TD has been paused, a hardware blackout period may be entered where remaining memory, final (mutable) control state of the source TD, and global control state may be exported from the source platform and imported by the destination platform.
[0093] The source VMM may call a TDH. EXPORT. PAGE logic 673 to perform its associated TDH. EXPORT. PAGE operation to export memory state. The destination VMM may call a TDH. IMPORT. PAGE logic 686 to perform its associated TDH. IMPORT. PAGE operation to import memory state. TD mutable non-memory state is a set of source TD state variables that might have changed since it was finalized. Immutable non-memory state exists for the TD scope (as part of the TDR and TDCS control structures) and the VCPU scope (as part of the TDVPS control structure) . The source VMM may call a TDH. EXPORT. STATE. VP logic 674 to perform its associated TDH. EXPORT. STATE. VP operation to export mutable TD VP state (per VCPU) . The destination VMM may call a TDH. IMPORT. STATE. VP logic 687 to perform its associated TDH. IMPORT. STATE. VP operation to import mutable TD VP state. The source VMM may call a TDH. EXPORT. STATE. TD logic 674 to perform its associated TDH. EXPORT. STATE. TD operation to export TD state (per TD) . The destination VMM may call a TDH. IMPORT. STATE. TD logic 687 to perform its associated TDH. IMPORT. STATE. TD operation to import TD state.
[0094] In some embodiments, post-copy of memory state may be used by allowing the TD to run on the destination platform. In some live migration scenarios, the host VMM may stage some memory state transfer to occur lazily after the destination TD has started execution. In this case, the host VMM will be required to fetch the required pages as accesses occur by the destination TD –this order of access is indeterminate and will likely differ from the order in which the host VMM has queued memory state to be transferred. The remaining memory state transferred in the post-copy stage via TDH. EXPORT. PAGE logic 676 and TDH. IMPORT. PAGE logic 690.
[0095] In order to support that on-demand model, the order of memory migration during post-copy is not enforced by TDX. In some embodiments, the host VMM may implement multiple migration queues with multiple priorities for memory state transfer. For example, the host VMM on the source platform may keep a copy of each encrypted migrated page until it receives a confirmation from the destination that the page has been successfully imported. If needed, that copy can be re-sent on a high priority queue. Another option is, instead of holding a copy of exported pages, to call TDH. EXPORT. PAGE again on demand. Also, to simplify host VMM software for this model, the TDX module interface API used for memory import in this post-copy stage will return additional informational error codes to indicate that a stale import was attempted by the host-VMM to account for the case where the low-latency import operation for a guest physical address (GPA) superseded the import from the higher latency import queue. Alternatively, in other embodiments, the VMM may first complete all memory migration before allowing the destination TD to run, yet benefit from the simpler and potentially higher performance operation supported during the out-of-order phase.
[0096] In some embodiments, the TDX modules may use a commitment protocol to enforce that all exported state for the source TD are imported before the destination TD may run. The commitment protocol may help to ensure that a host VMM cannot violate the security objectives of TD Live migration. For example, in some embodiments, it may be enforced that both the destination and source TDs may not continue to execute after live migration of the source TD to a destination TD, even if an error causes the TD migration to be aborted.
[0097] On the source platform, TDH. EXPORT. PAUSE logic 672 starts the blackout phase of TD live migration and TDH. EXPORT. STATE. DONE logic 691 ends the blackout phase of live migration (and marks the end of the transfer of TD memory pre-copy, mutable TD VP, and mutable TD global control state) . TDH. EXPORT. STATE. DONE logic 691 generates a cryptographically-authenticated start token to allow the destination TD to become runnable. On the destination platform, TDH. IMPORT. STATE. DONE logic 688 consumes the cryptographic start token to allow the destination TD to be un-paused.
[0098] In error scenarios, the migration process may be aborted proactively by the host on the source platform via TDH. EXPORT. ABORT logic 675 before a start token was generated. If a start token was already generated (e.g., pre-copy completed) , the destination platform may generate an abort token using TDH. IMPORT. STATE. ABORT logic 689 which generates an abort token which may be consumed by TDH. EXPORT. ABORT logic 675 by the source TD platform TDX module to abort the migration process and allow the source TD to become runnable again. Upon a successful commit, the source TD may be torn down and the migration key may be destroyed by TD teardown logic 677.
[0099] In some embodiments, the TDX modules may optionally enforce a number of security objectives. One security objective is that is that the CSP not be able to migrate a TD to a destination platform with a TCB that does not meet the minimum requirements expressed in the TD migration policy which may be configurable and attestable. As a corollary, TD attestation may enforce the base line of security. A TD may start on a stronger TCB platform and then be migrated to a weaker TCB, if they are within the security policy of the TD. Another security objective is that a CSP not be able to create a TD to be migratable without the tenant knowledge (the attestation report may include a TD migratable attribute) . Another security objective is that a CSP not be able to clone a TD during migration (only source or destination TD be executing after migration process completes) . Another security objective is that a CSP not be able to operate the destination TD on any stale state from the source TD. Another security objective is that security (confidentiality, integrity, and replay protection) of TD migration data should be agnostic of the transport mechanism between source and destination. Another security objective is that TD migration should enforce that non-migratable assets (anything that can be serialized and recreated on the target) are reset at source to allow for secure reuse and restored on destination after migration.
[0100] In some embodiments, the TDX modules may optionally enforce a number of functionality objectives. One functionality objective is that the CSP be able to migrate the tenant TD without tenant TD runtime involvement. Another functionality objective is that a TD be able to opt-in to migration at creation. Tenant software should not be part of this decision. Another functionality objective is that the CSP be able to minimize additional performance impact due to TD migration only during migrating the TD. For example, TD migration entails page fragmentation. Another functionality objective is that TD OS does not need new enlightenments to migrate TDs with TDX I / O direct assigned devices. As a corollary, the TDX I / O may allow for hot-plug device attach / detach notifications. Another functionality objective is that a TD be migratable with VMX nesting.
[0101] Existing approaches to VM or TD migration may include a migration policy concept, as described above, in which migration may happen if the policy is satisfied. However, there may be no way to know how many platforms a VM or TD has been migrated to in its history and / or for a customer or verifier to determine if the VM or TD can be trusted going forward and / or to what vulnerabilities a migrated VM or TD might have been exposed.
[0102] To overcome the limitation of not having a comprehensive migration history, one approach uses a hardcoded static migration policy in the Migration TD for the lifetime of a client TD (from the moment the Migration TD was bound to the client TD to the time the client TD is terminated) and this policy is encoded in the Migration TD code. According to this approach, migration is only allowed if the Migration Policy is the same between the source and destination (and migration to the destination platform is acceptable based on the migration policy) .
[0103] In this static hardcoded migration policy approach, the migration policy may define Minimum (Min) Secure Version Numbers (SVNs) and component measurement hashes that are acceptable, and the hardcoded migration policy itself may be used as a history.
[0104] For example, the hardcoded policy is measured and its hash is reflected in an information data structure (e.g., TDINFO_STRUCT) and the verifier assumes that the client TD ran or could be running on all platforms higher than the Min SVNs or with the hash values as defined in the migration policy, as illustrated for example in Figure 7. However, a hardcoded static migration policy approach may allow downgrading to Min SVNs defined in the migration policy because they do not provide sufficient history about what platforms a client TD ran on.
[0105] Therefore, embodiments may provide for preventing rollback to an old platform (e.g., after the platform is upgraded or patched and / or to ensure that the client TD is running on most up to date platform) . In embodiments, attestation may provide a history that the client TD previously ran on an old or vulnerable version of the platform.
[0106] Embodiments may provide for runtime updates of the migration policy (e.g., to mitigate vulnerability to rollback attacks) . For example, instead of providing one static Migration TD hash (e.g., SERVTD_HASH) for attestation (which includes Migration TD code and the migration policy, such that once the hash is calculated and bound to a user TD, it cannot be updated anymore without the CSP tearing down the current TD to bind a new Migration TD) , embodiments provide for an updateable migration policy that may be different between the source platform and the destination platform, as shown for example in Figure 8.
[0107] Embodiments provide a mechanism for comprehensive migration history logging, using the migration history as attestable proof to determine the platforms on which the client TD has run, thus breaking the dependency between migration policy and history. This approach allows the destination Migration TD to be built with a new, updated migration policy, independent of the current migration policy, as the history is reflected in the migration event log. It also enables rebinding to a newer, more up-to-date version of the Migration TD, potentially with an updated policy, without the need for actual migration, as shown for example in Figure 9.
[0108] Embodiments may include the creation of an attestable migration event log as a migration history. The migration event log may be verified by a verifier to determine the platforms on which the client TD has run. When the Migration TD performs a migration or rebinds to a newer destination Migration TD, it creates migration records and extends the hash of these records within the TD. For example, the hash (e.g., MIGRATION EVENT_HASH) may be part of the TD information data structure (e.g., TDINFO_STRUCT) , which may be a data structure containing the measurements and initial configuration of the TD that was locked at initialization and a set of measurement registers that are run-time extendable, and is reflected in the TD’s quote.
[0109] In embodiments, a public key hash is embedded in the Migration TD code. The TD owner can inspect the public key hash to confirm that the Migration TD is using a trusted signing service for updates.
[0110] During migration or rebind, the Migration TD may calculate the new event hash and provide it to the TDX Module. The TDX Module then extends this hash with the previous event hash in the client TD's TDCS.
[0111] The Migration TD owner provides the Migration TD code and its measurement to the TD owner. The TD owner verifies that the measurement is correct and publishes the measurement of the code to the verifier.
[0112] The TD then generates a quote for verification. This quote includes the migration event log (e.g., EventLoghash, a hash of the migration event log) , that contains a record (or hashed record) of all previous migration and rebinding events.
[0113] In embodiments, the TDX Module records the previous identity of the platform, the attestation components, and the old migration policy of the Migration TD in the EventLoghash. The EventLoghash is calculated by the destination migration TD when it approves the migration, or by the source migration TD when it approves the rebinding to a new migration TD and is stored in the TDX Module. Figure 10 shows an example of events that may be captured.
[0114] The integrity of the migration event log data may be protected by using a hash algorithm, as shown for example in Figure 11. If an attacker tampers the event log data, then the verifier (during verification as described below) will notice an EventLoghash mismatch.
[0115] During attestation, the client TD may request the VMM to provide the migration records corresponding to the TD. For example, when the client TD requests the quoting TD to sign the report, the VMM / CSP supplies the migration event log. The quoting TD retrieves the migration event log from the VMM and sends it, along with the quote, to the verifier, as shown for example in Figures 12 and 13.
[0116] In embodiments, the migration event log may be retrieved by the quoting TD even if the Migration TD is not running or has been terminated. Each event in the migration log is tagged with a unique identifier of the client TD (e.g., the client TD's TD_UUID) value for retrieval by the quoting TD.
[0117] The verifier checks the integrity of the migration event log through the EventLoghash. It hashes each record in the migration event log and compares it with the value in the EventLoghash. The verifier then verifies the quote and the Migration TD measurement. If everything matches, the verification is successful.
[0118] The Migration TD may also provide an event record to the VMM, allowing the VMM to add the record to the migration event log. For example, the EventLoghash may updated during migration as shown in Figure 14. After a successful key exchange (e.g., Diffie Hellman) , the destination migration TD calculates the EventLoghash for the event related to the migration from the source platform. The destination migration TD records this EventLoghash in the Client TD's TDCS. When the migration starts, the previous EventLoghash is imported from the Client TD and extended with the new EventLoghash provided by the destination migration TD. Finally, a new event record is created and provided to the VMM for the migration event log update through migration log service. In embodiments, the migration event log (e.g., the EventLoghash) is updated during rebinding, as shown for example in Figure 15. After the old migration TD accepts the new migration TD, the old migration TD performs the rebind. During this process, the TDX Module extends the EventLoghash with the measurement of the old migration TD. The old migration TD then generates a new record and an updated EventLoghash, which are provided to the VMM to send to the migration log service.
[0119] Embodiments may include the creation and use of an attestable migration event log as a migration history, with the following details (as or in addition to the details described above or below) : · Using an event log as a migration history format. · Recording Migration TD identity, TDX Module identity, platform and attestation component identity, migration policy, and action string as migration history data. · Using a TDX-module to record an EventLoghash to maintain the integrity of the history. · Using a trusted timer service to record the timestamp for rebind or pre-migration in the history.
[0120] Embodiments may include the creation and use of a migration policy rebinding mechanism, with the following details (as or in addition to the details described above or below) : · Using signed migration policy verification by the old Migration TD. · Enforcing the new Migration TD hash by the TDX-module for rebinding. · Recording the rebinding process to the migration history for evaluation.
[0121] Example apparatuses, methods, etc.
[0122] According to some examples, a system includes a plurality of processor cores and at least one non-transitory machine-readable storage medium storing a plurality of instructions to cause the plurality of processor cores to perform operations including to: determine to migrate a first protected VM from a first platform to a second protected VM on second platform; create a migration event log to indicate that the first protected VM has run on the first platform; determine to migrate the second protected VM from the second platform; use the migration event log as attestable proof to determine that the second protected VM has been migrated from the first platform; and update the migration event log to indicate that the second protected VM has run on the second platform.
[0123] Any such examples may include any or any combination of the following aspects.
[0124] According to some examples, a method includes determining to migrate a first protected virtual machine (VM) from a first platform to a second protected VM on second platform; creating a migration event log to indicate that the first protected VM has run on the first platform; determining to migrate the second protected VM from the second platform; using the migration event log as attestable proof to determine that the second protected VM has been migrated from the first platform; and updating the migration event log to indicate that the second protected VM has run on the second platform.
[0125] Any such examples may include any or any combination of the following aspects.
[0126] According to some examples, at least one non-transitory machine-readable storage medium storing a plurality of instructions, the plurality of instructions, if executed by a machine, causes the machine to perform operations including determining to migrate a first protected virtual machine (VM) from a first platform to a second protected VM on second platform; creating a migration event log to indicate that the first protected VM has run on the first platform; determining to migrate the second protected VM from the second platform; using the migration event log as attestable proof to determine that the second protected VM has been migrated from the first platform; and updating the migration event log to indicate that the second protected VM has run on the second platform.
[0127] Any such examples may include any or any combination of the following aspects. The migration event log includes a migration policy. The operations or method also include to update or updating the migration event log to update the migration policy. Updating the migration policy is in connection with a migration. Updating the migration policy is performed in connection with a rebinding without a migration. The migration event log includes an identifier of a third protected VM, the third protected VM to perform services in connection with migration. The services include enforcing the migration policy. The migration event log includes a migration timestamp. The migration event log includes a rebinding timestamp.
[0128] According to some examples, an apparatus may include means for performing any function disclosed herein; an apparatus may include a data storage device that stores code that when executed by a hardware processor or controller causes the hardware processor or controller to perform any method or portion of a method disclosed herein; an apparatus, method, system etc. may be as described in the detailed description; a non-transitory machine-readable medium may store instructions that when decoded and / or executed by a machine causes the machine to perform any method or portion of a method disclosed herein. Embodiments may include any details, features, etc. or combinations of details, features, etc. described in this specification.
[0129] Example Computer Architectures
[0130] Detailed below are descriptions of example computer architectures. Other system designs and configurations known in the arts for laptop, desktop, and handheld personal computers (PC) s, personal digital assistants, engineering workstations, servers, disaggregated servers, network devices, network hubs, switches, routers, embedded processors, digital signal processors (DSPs) , graphics devices, video game devices, set-top boxes, micro controllers, cell phones, portable media players, hand-held devices, and various other electronic devices, are also suitable. In general, a variety of systems or electronic devices capable of incorporating a processor and / or other execution logic as disclosed herein are generally suitable.
[0131] FIG. 16 illustrates an example computing system. Multiprocessor system 1600 is an interfaced system and includes a plurality of processors or cores including a first processor 1670 and a second processor 1680 coupled via an interface 1650 such as a point-to-point (P-P) interconnect, a fabric, and / or bus. In some examples, the first processor 1670 and the second processor 1680 are homogeneous. In some examples, the first processor 1670 and the second processor 1680 are heterogenous. Though the example system 1600 is shown to have two processors, the system may have three or more processors, or may be a single processor system. In some examples, the computing system is a system on a chip (SoC) .
[0132] Processors 1670 and 1680 are shown including integrated memory controller (IMC) circuitry 1672 and 1682, respectively. Processor 1670 also includes interface circuits 1676 and 1678; similarly, second processor 1680 includes interface circuits 1686 and 1688. Processors 1670, 1680 may exchange information via the interface 1650 using interface circuits 1678, 1688. IMCs 1672 and 1682 couple the processors 1670, 1680 to respective memories, namely a memory 1632 and a memory 1634, which may be portions of main memory locally attached to the respective processors.
[0133] Processors 1670, 1680 may each exchange information with a network interface (NW I / F) 1690 via individual interfaces 1652, 1654 using interface circuits 1676, 1694, 1686, 1698. The network interface 1690 (e.g., one or more of an interconnect, bus, and / or fabric, and in some examples is a chipset) may optionally exchange information with a coprocessor 1638 via an interface circuit 1692. In some examples, the coprocessor 1638 is a special-purpose processor, such as, for example, a high-throughput processor, a network or communication processor, compression engine, graphics processor, general purpose graphics processing unit (GPGPU) , neural-network processing unit (NPU) , embedded processor, or the like.
[0134] A shared cache (not shown) may be included in either processor 1670, 1680 or outside of both processors, yet connected with the processors via an interface such as P-P interconnect, such that either or both processors’ local cache information may be stored in the shared cache if a processor is placed into a low power mode.
[0135] Network interface 1690 may be coupled to a first interface 1616 via interface circuit 1696. In some examples, first interface 1616 may be an interface such as a Peripheral Component Interconnect (PCI) interconnect, a PCI Express interconnect or another I / O interconnect. In some examples, first interface 1616 is coupled to a power control unit (PCU) 1617, which may include circuitry, software, and / or firmware to perform power management operations with regard to the processors 1670, 1680 and / or co-processor 1638. PCU 1617 provides control information to a voltage regulator (not shown) to cause the voltage regulator to generate the appropriate regulated voltage. PCU 1617 also provides control information to control the operating voltage generated. In various examples, PCU 1617 may include a variety of power management logic units (circuitry) to perform hardware-based power management. Such power management may be wholly processor controlled (e.g., by various processor hardware, and which may be triggered by workload and / or power, thermal or other processor constraints) and / or the power management may be performed responsive to external sources (such as a platform or power management source or system software) .
[0136] PCU 1617 is illustrated as being present as logic separate from the processor 1670 and / or processor 1680. In other cases, PCU 1617 may execute on a given one or more of cores (not shown) of processor 1670 or 1680. In some cases, PCU 1617 may be implemented as a microcontroller (dedicated or general-purpose) or other control logic configured to execute its own dedicated power management code, sometimes referred to as P-code. In yet other examples, power management operations to be performed by PCU 1617 may be implemented externally to a processor, such as by way of a separate power management integrated circuit (PMIC) or another component external to the processor. In yet other examples, power management operations to be performed by PCU 1617 may be implemented within BIOS or other system software.
[0137] Various I / O devices 1614 may be coupled to first interface 1616, along with a bus bridge 1618 which couples first interface 1616 to a second interface 1620. In some examples, one or more additional processor (s) 1615, such as coprocessors, high throughput many integrated core (MIC) processors, GPGPUs, accelerators (such as graphics accelerators or digital signal processing (DSP) units) , field programmable gate arrays (FPGAs) , or any other processor, are coupled to first interface 1616. In some examples, second interface 1620 may be a low pin count (LPC) interface. Various devices may be coupled to second interface 1620 including, for example, a keyboard and / or mouse 1622, communication devices 1627 and storage circuitry 1628. Storage circuitry 1628 may be one or more non-transitory machine-readable storage media as described below, such as a disk drive or other mass storage device which may include instructions / code and data 1630. Further, an audio I / O 1624 may be coupled to second interface 1620. Note that other architectures than the point-to-point architecture described above are possible. For example, instead of the point-to-point architecture, a system such as multiprocessor system 1600 may implement a multi-drop interface or other such architecture.
[0138] Example Core Architectures, Processors, and Computer Architectures
[0139] Processor cores may be implemented in different ways, for different purposes, and in different processors. For instance, implementations of such cores may include: 1) a general purpose in-order core intended for general-purpose computing; 2) a high-performance general purpose out-of-order core intended for general-purpose computing; 3) a special purpose core intended primarily for graphics and / or scientific (throughput) computing. Implementations of different processors may include: 1) a CPU including one or more general purpose in-order cores intended for general-purpose computing and / or one or more general purpose out-of-order cores intended for general-purpose computing; and 2) a coprocessor including one or more special purpose cores intended primarily for graphics and / or scientific (throughput) computing. Such different processors lead to different computer system architectures, which may include: 1) the coprocessor on a separate chip from the CPU; 2) the coprocessor on a separate die in the same package as a CPU; 3) the coprocessor on the same die as a CPU (in which case, such a coprocessor is sometimes referred to as special purpose logic, such as integrated graphics and / or scientific (throughput) logic, or as special purpose cores) ; and 4) a system on a chip (SoC) that may be included on the same die as the described CPU (sometimes referred to as the application core (s) or application processor (s) ) , the above described coprocessor, and additional functionality. Example core architectures are described next, followed by descriptions of example processors and computer architectures.
[0140] FIG. 17 illustrates a block diagram of an example processor and / or SoC 1700 that may have one or more cores and an integrated memory controller. The solid lined boxes illustrate a processor 1700 with a single core 1702 (A) , system agent unit circuitry 1710, and a set of one or more interface controller unit (s) circuitry 1716, while the optional addition of the dashed lined boxes illustrates an alternative processor 1700 with multiple cores 1702 (A) - (N) , a set of one or more integrated memory controller unit (s) circuitry 1714 in the system agent unit circuitry 1710, and special purpose logic 1708, as well as a set of one or more interface controller units circuitry 1716. Note that the processor 1700 may be one of the processors 1670 or 1680, or co-processor 1638 or 1615 of FIG. 16.
[0141] Thus, different implementations of the processor 1700 may include: 1) a CPU with the special purpose logic 1708 being integrated graphics and / or scientific (throughput) logic (which may include one or more cores, not shown) , and the cores 1702 (A) - (N) being one or more general purpose cores (e.g., general purpose in-order cores, general purpose out-of-order cores, or a combination of the two) ; 2) a coprocessor with the cores 1702 (A) - (N) being a large number of special purpose cores intended primarily for graphics and / or scientific (throughput) ; and 3) a coprocessor with the cores 1702 (A) - (N) being a large number of general purpose in-order cores. Thus, the processor 1700 may be a general-purpose processor, coprocessor, or special-purpose processor, such as, for example, a network or communication processor, compression engine, graphics processor, GPGPU (general purpose graphics processing unit) , a high throughput many integrated cores (MIC) coprocessor (including 30 or more cores) , embedded processor, or the like. The processor may be implemented on one or more chips. The processor 1700 may be a part of and / or may be implemented on one or more substrates using any of a number of process technologies, such as, for example, complementary metal oxide semiconductor (CMOS) , bipolar CMOS (BiCMOS) , P-type metal oxide semiconductor (PMOS) , or N-type metal oxide semiconductor (NMOS) .
[0142] A memory hierarchy includes one or more levels of cache unit (s) circuitry 1704 (A) - (N) within the cores 1702 (A) - (N) , a set of one or more shared cache unit (s) circuitry 1706, and external memory (not shown) coupled to the set of integrated memory controller unit (s) circuitry 1714. The set of one or more shared cache unit (s) circuitry 1706 may include one or more mid-level caches, such as level 2 (L2) , level 3 (L3) , level 4 (L4) , or other levels of cache, such as a last level cache (LLC) , and / or combinations thereof. While in some examples interface network circuitry 1712 (e.g., a ring interconnect) interfaces the special purpose logic 1708 (e.g., integrated graphics logic) , the set of shared cache unit (s) circuitry 1706, and the system agent unit circuitry 1710, alternative examples use any number of well-known techniques for interfacing such units. In some examples, coherency is maintained between one or more of the shared cache unit (s) circuitry 1706 and cores 1702 (A) - (N) . In some examples, interface controller unit circuitry 1716 couples the cores 1702 to one or more other devices 1718 such as one or more I / O devices, storage, one or more communication devices (e.g., wireless networking, wired networking, etc. ) , etc.
[0143] In some examples, one or more of the cores 1702 (A) - (N) are capable of multi-threading. The system agent unit circuitry 1710 includes those components coordinating and operating cores 1702 (A) - (N) . The system agent unit circuitry 1710 may include, for example, power control unit (PCU) circuitry and / or display unit circuitry (not shown) . The PCU may be or may include logic and components needed for regulating the power state of the cores 1702 (A) - (N) and / or the special purpose logic 1708 (e.g., integrated graphics logic) . The display unit circuitry is for driving one or more externally connected displays.
[0144] The cores 1702 (A) - (N) may be homogenous in terms of instruction set architecture (ISA) . Alternatively, the cores 1702 (A) - (N) may be heterogeneous in terms of ISA; that is, a subset of the cores 1702 (A) - (N) may be capable of executing an ISA, while other cores may be capable of executing only a subset of that ISA or another ISA.
[0145] Example Core Architectures -In-order and out-of-order core block diagram.
[0146] FIG. 18A is a block diagram illustrating both an example in-order pipeline and an example register renaming, out-of-order issue / execution pipeline according to examples. FIG. 18B is a block diagram illustrating both an example in-order architecture core and an example register renaming, out-of-order issue / execution architecture core to be included in a processor according to examples. The solid lined boxes in FIGS. 18A-18B illustrate the in-order pipeline and in-order core, while the optional addition of the dashed lined boxes illustrates the register renaming, out-of-order issue / execution pipeline and core. Given that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.
[0147] In FIG. 18A, a processor pipeline 1800 includes a fetch stage 1802, an optional length decoding stage 1804, a decode stage 1806, an optional allocation (Alloc) stage 1808, an optional renaming stage 1810, a schedule (also known as a dispatch or issue) stage 1812, an optional register read / memory read stage 1814, an execute stage 1816, a write back / memory write stage 1818, an optional exception handling stage 1822, and an optional commit stage 1824. One or more operations can be performed in each of these processor pipeline stages. For example, during the fetch stage 1802, one or more instructions are fetched from instruction memory, and during the decode stage 1806, the one or more fetched instructions may be decoded, addresses (e.g., load store unit (LSU) addresses) using forwarded register ports may be generated, and branch forwarding (e.g., immediate offset or a link register (LR) ) may be performed. In one example, the decode stage 1806 and the register read / memory read stage 1814 may be combined into one pipeline stage. In one example, during the execute stage 1816, the decoded instructions may be executed, LSU address / data pipelining to an Advanced Microcontroller Bus (AMB) interface may be performed, multiply and add operations may be performed, arithmetic operations with branch results may be performed, etc.
[0148] By way of example, the example register renaming, out-of-order issue / execution architecture core of FIG. 18B may implement the pipeline 1800 as follows: 1) the instruction fetch circuitry 1838 performs the fetch and length decoding stages 1802 and 1804; 2) the decode circuitry 1840 performs the decode stage 1806; 3) the rename / allocator unit circuitry 1852 performs the allocation stage 1808 and renaming stage 1810; 4) the scheduler (s) circuitry 1856 performs the schedule stage 1812; 5) the physical register file (s) circuitry 1858 and the memory unit circuitry 1870 perform the register read / memory read stage 1814; the execution cluster (s) 1860 perform the execute stage 1816; 6) the memory unit circuitry 1870 and the physical register file (s) circuitry 1858 perform the write back / memory write stage 1818; 7) various circuitry may be involved in the exception handling stage 1822; and 8) the retirement unit circuitry 1854 and the physical register file (s) circuitry 1858 perform the commit stage 1824.
[0149] FIG. 18B shows a processor core 1890 including front-end unit circuitry 1830 coupled to execution engine unit circuitry 1850, and both are coupled to memory unit circuitry 1870. The core 1890 may be a reduced instruction set architecture computing (RISC) core, a complex instruction set architecture computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As yet another option, the core 1890 may be a special-purpose core, such as, for example, a network or communication core, compression engine, coprocessor core, general purpose computing graphics processing unit (GPGPU) core, graphics core, or the like.
[0150] The front-end unit circuitry 1830 may include branch prediction circuitry 1832 coupled to instruction cache circuitry 1834, which is coupled to an instruction translation lookaside buffer (TLB) 1836, which is coupled to instruction fetch circuitry 1838, which is coupled to decode circuitry 1840. In one example, the instruction cache circuitry 1834 is included in the memory unit circuitry 1870 rather than the front-end circuitry 1830. The decode circuitry 1840 (or decoder) may decode instructions, and generate as an output one or more micro-operations, micro-code entry points, microinstructions, other instructions, or other control signals, which are decoded from, or which otherwise reflect, or are derived from, the original instructions. The decode circuitry 1840 may further include address generation unit (AGU, not shown) circuitry. In one example, the AGU generates an LSU address using forwarded register ports, and may further perform branch forwarding (e.g., immediate offset branch forwarding, LR register branch forwarding, etc. ) . The decode circuitry 1840 may be implemented using various different mechanisms. Examples of suitable mechanisms include, but are not limited to, look-up tables, hardware implementations, programmable logic arrays (PLAs) , microcode read only memories (ROMs) , etc. In one example, the core 1890 includes a microcode ROM (not shown) or other medium that stores microcode for certain macroinstructions (e.g., in decode circuitry 1840 or otherwise within the front-end circuitry 1830) . In one example, the decode circuitry 1840 includes a micro-operation (micro-op) or operation cache (not shown) to hold / cache decoded operations, micro-tags, or micro-operations generated during the decode or other stages of the processor pipeline 1800. The decode circuitry 1840 may be coupled to rename / allocator unit circuitry 1852 in the execution engine circuitry 1850.
[0151] The execution engine circuitry 1850 includes the rename / allocator unit circuitry 1852 coupled to retirement unit circuitry 1854 and a set of one or more scheduler (s) circuitry 1856. The scheduler (s) circuitry 1856 represents any number of different schedulers, including reservations stations, central instruction window, etc. In some examples, the scheduler (s) circuitry 1856 can include arithmetic logic unit (ALU) scheduler / scheduling circuitry, ALU queues, address generation unit (AGU) scheduler / scheduling circuitry, AGU queues, etc. The scheduler (s) circuitry 1856 is coupled to the physical register file (s) circuitry 1858. Each of the physical register file (s) circuitry 1858 represents one or more physical register files, different ones of which store one or more different data types, such as scalar integer, scalar floating-point, packed integer, packed floating-point, vector integer, vector floating-point, status (e.g., an instruction pointer that is the address of the next instruction to be executed) , etc. In one example, the physical register file (s) circuitry 1858 includes vector registers unit circuitry, writemask registers unit circuitry, and scalar register unit circuitry. These register units may provide architectural vector registers, vector mask registers, general-purpose registers, etc. The physical register file (s) circuitry 1858 is coupled to the retirement unit circuitry 1854 (also known as a retire queue or a retirement queue) to illustrate various ways in which register renaming and out-of-order execution may be implemented (e.g., using a reorder buffer (s) (ROB (s) ) and a retirement register file (s) ; using a future file (s) , a history buffer (s) , and a retirement register file (s) ; using a register maps and a pool of registers; etc. ) . The retirement unit circuitry 1854 and the physical register file (s) circuitry 1858 are coupled to the execution cluster (s) 1860. The execution cluster (s) 1860 includes a set of one or more execution unit (s) circuitry 1862 and a set of one or more memory access circuitry 1864. The execution unit (s) circuitry 1862 may perform various arithmetic, logic, floating-point or other types of operations (e.g., shifts, addition, subtraction, multiplication) and on various types of data (e.g., scalar integer, scalar floating-point, packed integer, packed floating-point, vector integer, vector floating-point) . While some examples may include a number of execution units or execution unit circuitry dedicated to specific functions or sets of functions, other examples may include only one execution unit circuitry or multiple execution units / execution unit circuitry that all perform all functions. The scheduler (s) circuitry 1856, physical register file (s) circuitry 1858, and execution cluster (s) 1860 are shown as being possibly plural because certain examples create separate pipelines for certain types of data / operations (e.g., a scalar integer pipeline, a scalar floating-point / packed integer / packed floating-point / vector integer / vector floating-point pipeline, and / or a memory access pipeline that each have their own scheduler circuitry, physical register file (s) circuitry, and / or execution cluster –and in the case of a separate memory access pipeline, certain examples are implemented in which only the execution cluster of this pipeline has the memory access unit (s) circuitry 1864) . It should also be understood that where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution and the rest in-order.
[0152] In some examples, the execution engine unit circuitry 1850 may perform load store unit (LSU) address / data pipelining to an Advanced Microcontroller Bus (AMB) interface (not shown) , and address phase and writeback, data phase load, store, and branches.
[0153] The set of memory access circuitry 1864 is coupled to the memory unit circuitry 1870, which includes data TLB circuitry 1872 coupled to data cache circuitry 1874 coupled to level 2 (L2) cache circuitry 1876. In one example, the memory access circuitry 1864 may include load unit circuitry, store address unit circuitry, and store data unit circuitry, each of which is coupled to the data TLB circuitry 1872 in the memory unit circuitry 1870. The instruction cache circuitry 1834 is further coupled to the level 2 (L2) cache circuitry 1876 in the memory unit circuitry 1870. In one example, the instruction cache 1834 and the data cache 1874 are combined into a single instruction and data cache (not shown) in L2 cache circuitry 1876, level 3 (L3) cache circuitry (not shown) , and / or main memory. The L2 cache circuitry 1876 is coupled to one or more other levels of cache and eventually to a main memory.
[0154] The core 1890 may support one or more instructions sets (e.g., the x86 instruction set architecture (optionally with some extensions that have been added with newer versions) ; the MIPS instruction set architecture; the ARM instruction set architecture (optionally with optional additional extensions such as NEON) ) , including the instruction (s) described herein. In one example, the core 1890 includes logic to support a packed data instruction set architecture extension (e.g., AVX1, AVX2) , thereby allowing the operations used by many multimedia applications to be performed using packed data.
[0155] Example Execution Unit (s) Circuitry
[0156] FIG. 19 illustrates examples of execution unit (s) circuitry, such as execution unit (s) circuitry 1862 of FIG. 18B. As illustrated, execution unit (s) circuitry 1862 may include one or more ALU circuits 1901, optional vector / single instruction multiple data (SIMD) circuits 1903, load / store circuits 1905, branch / jump circuits 1907, and / or Floating-point unit (FPU) circuits 1909. ALU circuits 1901 perform integer arithmetic and / or Boolean operations. Vector / SIMD circuits 1903 perform vector / SIMD operations on packed data (such as SIMD / vector registers) . Load / store circuits 1905 execute load and store instructions to load data from memory into registers or store from registers to memory. Load / store circuits 1905 may also generate addresses. Branch / jump circuits 1907 cause a branch or jump to a memory address depending on the instruction. FPU circuits 1909 perform floating-point arithmetic. The width of the execution unit (s) circuitry 1862 varies depending upon the example and can range from 16-bit to 1, 024-bit, for example. In some examples, two or more smaller execution units are logically combined to form a larger execution unit (e.g., two 128-bit execution units are logically combined to form a 256-bit execution unit) .
[0157] Program code may be applied to input information to perform the functions described herein and generate output information. The output information may be applied to one or more output devices, in known fashion. For purposes of this application, a processing system includes any system that has a processor, such as, for example, a digital signal processor (DSP) , a microcontroller, an application specific integrated circuit (ASIC) , a field programmable gate array (FPGA) , a microprocessor, or any combination thereof.
[0158] The program code may be implemented in a high-level procedural or object-oriented programming language to communicate with a processing system. The program code may also be implemented in assembly or machine language, if desired. In fact, the mechanisms described herein are not limited in scope to any particular programming language. In any case, the language may be a compiled or interpreted language.
[0159] Examples of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of such implementation approaches. Examples may be implemented as computer programs or program code executing on programmable systems comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements) , at least one input device, and at least one output device.
[0160] One or more aspects of at least one example may be implemented by representative instructions stored on a machine-readable medium which represents various logic within the processor, which when read by a machine causes the machine to fabricate logic to perform the techniques described herein. Such representations, known as “intellectual property (IP) cores” may be stored on a tangible, machine readable medium and supplied to various customers or manufacturing facilities to load into the fabrication machines that make the logic or processor.
[0161] Such machine-readable storage media may include, without limitation, non-transitory, tangible arrangements of articles manufactured or formed by a machine or device, including storage media such as hard disks, any other type of disk including floppy disks, optical disks, compact disk read-only memories (CD-ROMs) , compact disk rewritables (CD-RWs) , and magneto-optical disks, semiconductor devices such as read-only memories (ROMs) , random access memories (RAMs) such as dynamic random access memories (DRAMs) , static random access memories (SRAMs) , erasable programmable read-only memories (EPROMs) , flash memories, electrically erasable programmable read-only memories (EEPROMs) , phase change memory (PCM) , magnetic or optical cards, or any other type of media suitable for storing electronic instructions.
[0162] Accordingly, examples also include non-transitory, tangible machine-readable media containing instructions or containing design data, such as Hardware Description Language (HDL) , which defines structures, circuits, apparatuses, processors, and / or system features described herein. Such examples may also be referred to as program products.
[0163] Emulation (including binary translation, code morphing, etc. )
[0164] In some cases, an instruction converter may be used to convert an instruction from a source instruction set architecture to a target instruction set architecture. For example, the instruction converter may translate (e.g., using static binary translation, dynamic binary translation including dynamic compilation) , morph, emulate, or otherwise convert an instruction to one or more other instructions to be processed by the core. The instruction converter may be implemented in software, hardware, firmware, or a combination thereof. The instruction converter may be on processor, off processor, or part on and part off processor.
[0165] FIG. 20 is a block diagram illustrating the use of a software instruction converter to convert binary instructions in a source ISA to binary instructions in a target ISA according to examples. In the illustrated example, the instruction converter is a software instruction converter, although alternatively the instruction converter may be implemented in software, firmware, hardware, or various combinations thereof. FIG. 20 shows a program in a high-level language 2002 may be compiled using a first ISA compiler 2004 to generate first ISA binary code 2006 that may be natively executed by a processor with at least one first ISA core 2016. The processor with at least one first ISA core 2016 represents any processor that can perform substantially the same functions as an processor with at least one first ISA core by compatibly executing or otherwise processing (1) a substantial portion of the first ISA or (2) object code versions of applications or other software targeted to run on an Intel processor with at least one first ISA core, in order to achieve substantially the same result as a processor with at least one first ISA core. The first ISA compiler 2004 represents a compiler that is operable to generate first ISA binary code 2006 (e.g., object code) that can, with or without additional linkage processing, be executed on the processor with at least one first ISA core 2016. Similarly, FIG. 20 shows the program in the high-level language 2002 may be compiled using an alternative ISA compiler 2008 to generate alternative ISA binary code 2010 that may be natively executed by a processor without a first ISA core 2014. The instruction converter 2012 is used to convert the first ISA binary code 2006 into code that may be natively executed by the processor without a first ISA core 2014. This converted code is not necessarily to be the same as the alternative ISA binary code 2010; however, the converted code will accomplish the general operation and be made up of instructions from the alternative ISA. Thus, the instruction converter 2012 represents software, firmware, hardware, or a combination thereof that, through emulation, simulation, or any other process, allows a processor or other electronic device that does not have a first ISA processor or core to execute the first ISA binary code 2006.
[0166] References to “one example, ” “an example, ” “one embodiment, ” “an embodiment, ” etc., indicate that the example or embodiment described may include a particular feature, structure, or characteristic, but every example or embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same example or embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example or embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other examples or embodiments whether or not explicitly described.
[0167] Moreover, in the various examples described above, unless specifically noted otherwise, disjunctive language such as the phrase “at least one of A, B, or C” or “A, B, and / or C” is intended to be understood to mean either A, B, or C, or any combination thereof (i.e., A and B, A and C, B and C, and A, B and C) . As used in this specification and the claims and unless otherwise specified, the use of the ordinal adjectives “first, ” “second, ” “third, ” etc. to describe an element merely indicates that a particular instance of an element or different instances of like elements are being referred to and is not intended to imply that the elements so described must be in a particular sequence, either temporally, spatially, in ranking, or in any other manner. Also, as used in descriptions of embodiments, a “ / ” character between terms may mean that what is described may include or be implemented using, with, and / or according to the first term and / or the second term (and / or any other additional terms) .
[0168] The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense. It will, however, be evident that various modifications and changes may be made thereunto without departing from the broader spirit and scope of the disclosure as set forth in the claims.
Claims
1.A system comprising:a plurality of processor cores; andat least one non-transitory machine-readable storage medium storing a plurality of instructions, the plurality of instructions, if executed by the plurality of processor cores, to cause the plurality of processor cores to perform operations including to:determine to migrate a first protected virtual machine (VM) from a first platform to a second protected VM on second platform;create a migration event log to indicate that the first protected VM has run on the first platform;determine to migrate the second protected VM from the second platform;use the migration event log as attestable proof to determine that the second protected VM has been migrated from the first platform; andupdate the migration event log to indicate that the second protected VM has run on the second platform.2.The system of claim 1, wherein the migration event log includes a migration policy.3.The system of claim 2, wherein the operations also include to update the migration event log to update the migration policy.4.The system of claim 3, wherein updating the migration policy is in connection with a migration.5.The system of claim 3, wherein updating the migration policy is performed in connection with a rebinding without a migration.6.The system of claim 2, wherein the migration event log includes an identifier of a third protected VM, the third protected VM to perform services in connection with migration.7.The system of claim 6, wherein the services include enforcing the migration policy.8.The system of claim 1, wherein the migration event log includes a migration timestamp.9.The system of claim 5, wherein the migration event log includes a rebinding timestamp.10.A method comprising:determining to migrate a first protected virtual machine (VM) from a first platform to a second protected VM on second platform;creating a migration event log to indicate that the first protected VM has run on the first platform;determining to migrate the second protected VM from the second platform;using the migration event log as attestable proof to determine that the second protected VM has been migrated from the first platform; andupdating the migration event log to indicate that the second protected VM has run on the second platform.11.The method of claim 10, wherein the migration event log includes a migration policy.12.The method of claim 11, further comprising updating the migration event log to update the migration policy.13.The method of claim 12, wherein updating the migration policy is in connection with a migration.14.The method of claim 12, wherein updating the migration policy is performed in connection with a rebinding without a migration.15.The method of claim 11, wherein the migration event log includes an identifier of a third protected VM, the third protected VM to perform services in connection with migration.16.The method of claim 15, wherein the services include enforcing the migration policy.17.The method of claim 10, wherein the migration event log includes a migration timestamp.18.The method of claim 14, wherein the migration event log includes a rebinding timestamp.19.At least one non-transitory machine-readable storage medium storing a plurality of instructions, the plurality of instructions, if executed by a machine, causes the machine to perform operations including:determining to migrate a first protected virtual machine (VM) from a first platform to a second protected VM on second platform;creating a migration event log to indicate that the first protected VM has run on the first platform;determining to migrate the second protected VM from the second platform;using the migration event log as attestable proof to determine that the second protected VM has been migrated from the first platform; andupdating the migration event log to indicate that the second protected VM has run on the second platform.20.The at least one non-transitory machine-readable storage medium of claim 19, wherein the operations also include updating the migration event log to update a migration policy.