Ownership of the password storage
By introducing CMOT and MKTME engines, combined with the use of TDX instructions and the use of key private keys, the problem of trust domain and virtual machine integrity protection in the prior art is solved, efficient encryption and signature of memory pages is realized, and security and isolation are ensured.
Patent Information
- Application Number
- CN201780094639.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2017-09-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2037-09-29
AI Technical Summary
The prior art is difficult to effectively protect the integrity of trust domains (TDs) and their virtual machines, especially when faced with attacks from malicious virtual machine managers (VMMs) or virtual machines (VMs).
Encryption and signature using TDX instructions (such as TDADDPAGE) and key private keys to ensure the integrity and isolation of memory pages.
Maintaining the security of large CMOTs in main memory, ensuring the integrity of each CMOT entry, and providing security for multiple TDs, reducing the impact of malicious VMMs or VMs on TDs.
Smart Images

Figure CN111095252B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of data center computing, and more specifically but not exclusively to systems and methods for cryptographic memory ownership. Background Art
[0002] Multi-processor systems are becoming increasingly common. In the modern world, computing resources play an increasingly integrated role in human life. As computers become more prevalent, controlling everything from the power grid to large industrial machines to personal computers to light bulbs, the demand for more capable processors increases. Brief Description of the Drawings
[0003] The present disclosure is best understood by reading the following detailed description in conjunction with the accompanying drawings. It should be emphasized that, according to standard industry practice, the various features are not necessarily drawn to scale and are for illustrative purposes only. Where scale is explicitly or implicitly shown, it provides only an illustrative example. In other embodiments, the dimensions of the various features may be arbitrarily enlarged or reduced for clarity of discussion.
[0004] Figure 1 is a block diagram of selected components of a data center with network connectivity according to one or more examples of the present specification.
[0005] Figure 2 is a block diagram of selected components of an end-user computing device according to one or more examples of the present specification.
[0006] Figure 3 is a block diagram of a computing system according to one or more examples of the present specification.
[0007] Figure 4 is a block diagram of a computing system illustrating additional aspects of the teachings of the present specification.
[0008] Figure 5 is a flowchart of a method that may be performed in conjunction with the teachings of the present specification.
[0009] Figures 6a - 6b is a block diagram illustrating a general vector-friendly instruction format and its instruction templates according to one or more examples of the present specification.
[0010] Figures 7a - 7d is a block diagram illustrating an example specific vector-friendly instruction format according to one or more examples of the present specification.
[0011] Figure 8 is a block diagram of a register architecture according to one or more examples of the present specification.
[0012] Figure 9ais a block diagram illustrating an example in - order pipeline and an out - of - order issue / execution pipeline with example register renaming, according to one or more examples of this specification.
[0013] Figure 9b is a block diagram illustrating an example in - order architectural core to be included in a processor and an out - of - order issue / execution architectural core with example register renaming, according to one or more examples of this specification.
[0014] Figures 10a - 10b is a block diagram illustrating a more specific in - order core architecture, according to one or more examples of this specification, which core would be one of several logic blocks in a chip (including other cores of the same type and / or different types).
[0015] Figure 11 is a block diagram of a processor that, according to one or more examples of this specification, may have more than one core, may have an integrated memory controller, and may have integrated graphics.
[0016] Figures 12 - 15 is a block diagram of a computer architecture, according to one or more examples of this specification.
[0017] Figure 16 is a block diagram for contrasting the conversion of binary instructions in a source instruction set to binary instructions in a target instruction set using a software instruction converter, according to one or more examples of this specification. Detailed Description
[0018] The following disclosure provides many different embodiments or examples for implementing different features of the present disclosure. Specific examples of components and arrangements are described below to simplify the present disclosure. Of course, these are merely examples and are not intended to be restrictive. Additionally, the present disclosure may repeat reference numerals and / or letters in the examples. This repetition is for simplicity and clarity only and does not in itself prescribe a relationship between the various embodiments and / or configurations discussed. Different embodiments may have different advantages, and no particular advantage is necessarily required for any embodiment.
[0019] Many existing microprocessor architectures include special instructions for provisioning memory enclaves and for establishing and utilizing a trusted execution environment (TEE). For example, Intel Software Guard Extensions (SGX) instructions can be used to establish memory enclaves such that only special SGX instructions can be used to enter, exit, or manipulate the memory within the enclave.
[0020] Existing Trusted Execution Environments (TEEs), such as SGX, can provide a Memory Encryption Engine (MEE) that encrypts memory, ensures memory integrity, and protects memory from attacks such as replay attacks. Typically, a TEE is established by provisioning a small area of memory as an enclave and using that enclave for the part of the application known as the Trusted Computing Base (TCB).
[0021] In many cases, memory encrypted within a TEE is signed using an encryption key that is used for encryption, decryption, and verification.
[0022] Embodiments of this specification include features of a device such as a microprocessor that is configured to provide not only a small memory enclave within an application, but also an entire Trust Domain (TD), which can be (or include) a virtual machine (VM) that provides fully trusted execution. Similar to Intel instructions, example processors of this specification can be provided with Trust Domain eXtension (TDX) instructions. TDX instructions can be used to provision an isolated VM that can operate as a Trust Domain protected not only from other VMs but also from a Virtual Machine Monitor (VMM), a hypervisor, or other management entity acting as a blind hypervisor in the case of a Trust Domain.
[0023] In the case of a standard untrusted VM, the VMM has full access to the underlying operations of the VM. The VMM can inspect or change the state of the VM and its memory. However, in the case of a TD, the VMM is "blind" with respect to the internal workings of the VM. The VMM only provisions the TD, allocates computing resources including memory to the TD, and provides a Host Physical Mapping (HPM) of memory locations to a Guest Physical Mapping (GPM). In other words, the VMM knows the mapping of the guest physical address (or in other words, the guest's virtualized physical address) to the physical address in the physical memory of the host device. However, the TD maintains its own Guest Virtual Mapping (GVM) to GPM table. The VMM is unaware of the GVM to GPM mapping. This reduces the ability of a compromised or malicious VMM to corrupt the integrity of the TD. It also reduces the ability of a compromised or malicious VM or TD operating on the same VMM to affect an uncompromised TD.
[0024] Because the VMM lacks visibility into the GVM to GPM mapping table, the only remaining attack vector is the extended memory page that contains the HPM to GPM mapping. This extended mapping can be protected using a Memory Ownership Table (MOT).
[0025] In an example, the MOT is a microarchitectural structure that cannot be directly accessed by software. Architecturally, it saves security attributes for each 4KB memory page. Advantageously, in one example, the MOT includes a 40-bit TD control structure (TDCS) pointer and provides a reverse mapping for the GPM to HPM mapping. Thus, when the processor "walks" through a memory page, the walk can include an integrity check in which it ensures that the mapping has not been tampered with. The CPU can use the MOT similarly to the MEE to ensure memory integrity and protect against tampering and attacks (such as replay attacks).
[0026] To ensure protection of the TD, the MOT itself is protected from tampering. In some cases, the MOT can be protected by placing it in a special memory area accessible only to the CPU or the microarchitecture (such as dedicated microarchitectural memory) or in a memory page protected by a physical memory range register (PMRR) that can only be accessed by the CPU. This ensures that it cannot be tampered with by software processes.
[0027] However, compared to the very large main memory in modern computing systems, such dedicated microarchitectural memory is relatively precious. Therefore, it is advantageous in some embodiments to move the MOT out of such dedicated memory locations and into the main memory.
[0028] To accomplish this while still maintaining the integrity of the MOT, a Cryptographic MOT (CMOT) can be used. The CMOT can be set within the memory range protected by the PMRR or can be set in unprotected memory. For further security, each trust domain is provided with its own private key for encryption and signature. In a system where multiple tenants share a common hardware platform, each tenant can have its own trust domain in which one or more Trusted Virtual Machines (TVMs) can be provisioned. Each entry in the CMOT can be encrypted using the corresponding private key of the associated TD, thus ensuring that the owner of the key "owns" the corresponding physical memory page. The CMOT is managed by a Multi-Key Total Memory Encryption (MKTME) engine that is capable of isolating tenants and VMs within the key domain, where the key domain includes at least one exclusive key for the tenant owning the TD. This cryptographically isolates the tenant's TD or key domain from other tenants as well as the CSP itself. A processor (e.g., a page miss handler (PMH) that performs a page walk) can use the CMOT to determine whether a physical memory address and client physical address memory mapping has been assigned to the correct owning TD. Each entry in the table maps a physical page to a key and is encrypted using the TD's private encryption key. Thus, the party owning the key domain can verify that the memory mapping is correct for its unique key while encrypting the verified entries using the key.
[0029] The MKTME engine can utilize physical address bits to convey which key is used to encrypt or decrypt data lines going to or coming from physical memory. In one example, there is one key and key identifier (within the address) for exclusive use by the owning trust domain (i.e., the TD's private key). This is the key or key ID used to encrypt the individual CMOT entries associated with that TD.
[0030] On memory write, the MKTME engine uses the TD's private key to encrypt the data being written, and on memory read, the MKTME engine uses the private key to decrypt the data. Since CPU / PMH access is to the CMOT mapping, cache lines belonging to CMOT entries can be encrypted or decrypted. If the integrity value for the entry is corrupted during a memory read by the PMH, the PMH rejects the mapping. The integrity value or field can be, for example, one or more bits that must be set to 0 to indicate a successful mapping.
[0031] Thus, encrypting the entry with a key or key ID associated with the entry and verifying the integrity value allows the CMOT to reside in a larger amount of main memory while preventing a hardware attacker from replaying a CMOT entry created for one TD to a different TD or an untrusted VM or the VMM itself. Additionally, any attempt by a physical attacker to modify the ciphertext of the CMOT in memory will result in memory corruption that can be detected in the processor because the integrity value of the CMOT (i.e., the integrity field containing the hash of the CMOT entry) will be corrupted.
[0032] Examples of this specification provide a new TDX instruction (such as "TDADDPAGE") that marks an idle CMOT entry corresponding to an HPA as being exclusively allocated to the TD specified by the TD identifier (TDID). The CPU writes the CMOT entry using the exclusive private key ID of the TD in the address bits. This associates the CMOT entry with the exclusive private key of the TD. Any other previous page state results in an error. The instruction enforces a cross-thread translation lookaside buffer (TLB) invalidate operation to ensure that no other TD is caching the mapping to this physical page address (PPA). The instruction leaf can be called by VMM software. The instruction specifies the initial GPA mapped to the specified HPA. The CPU verifies whether a GPA has been mapped for the HPA by walking through the extended page table (EPT) structure managed by the VMM.
[0033] As shown herein, the teachings of this specification enable a large CMOT to be maintained in a large amount of main memory while ensuring the integrity of each CMOT entry and providing security for multiple TDs, each of which may have its own one or more private keys and can be protected from each other, from untrusted VMs, and from the VMM itself, thus ensuring that the CSP cannot corrupt the integrity of the TD and its virtual machines either.
[0034] A system and method for cryptographic memory ownership will now be described more specifically with reference to the accompanying drawings. It should be noted that throughout the drawings, some reference numerals may be repeated to indicate that a particular device or block is exactly or substantially consistent across the drawings. However, this is not intended to imply any particular relationship between the disclosed embodiments. In some examples, a class of elements may be referred to by a particular reference numeral ("widget 10"), while individual species or examples within that class may be referred to by hyphenated labels ("first particular widget 10-1" and "second particular widget 10-2").
[0035] Figure 1It is a block diagram of selected components of a data center 100 of a cloud service provider (CSP) 102 with connectivity to a network, according to one or more examples of this specification. As a non-limiting example, the CSP 102 can be a traditional enterprise data center, an enterprise "private cloud", or a "public cloud" that provides services such as infrastructure as a service (IaaS), platform as a service (PaaS), or software as a service (SaaS).
[0036] The CSP 102 can provision a number of workload clusters 118, which can be clusters of stand-alone servers, blade servers, rack servers, or any other suitable server topology. In this illustrative example, two workload clusters 118-1 and 118-2 are shown, with each workload cluster providing rack servers 146 in a chassis 148.
[0037] In this illustration, the workload clusters 118 are shown as modular workload clusters that conform to the rack unit ("U") standard, where a standard 19-inch wide rack can be constructed to accommodate 42 units (42U), each unit being 1.75 inches high and approximately 36 inches deep. In this case, computing resources (such as processors, memory, storage devices, accelerators, and switches) can be incorporated into some multiple of the rack units from 1 to 42.
[0038] Each server 146 can host an independent operating system and provide server functionality, or the servers can be virtualized, in which case they can be under the control of a virtual machine manager (VMM), a hypervisor, and / or an orchestrator, and can host one or more virtual machines, virtual servers, or virtual appliances. These server racks can be co-located in a single data center or can be located in different geographical data centers. Depending on the contract agreement, some servers 146 can be dedicated to certain enterprise customers or tenants, while other servers can be shared.
[0039] The various devices in the data center can be connected to each other via a fabric 170, which can include one or more high-speed routing and / or switching devices. The fabric 170 can provide both "north-south" traffic (e.g., traffic going to and from a wide area network (WAN) such as the Internet) and "east-west" traffic (e.g., traffic across the data center). Historically, north-south traffic has accounted for a large chunk of network traffic, but as network services have become more complex and distributed, east-west traffic volume has risen. In many data centers, east-west traffic now accounts for the majority of the traffic.
[0040] In addition, as the capabilities of each server 146 increase, the traffic volume can be further increased. For example, each server 146 can provide multiple processor slots, where each slot houses a processor with 4 to 8 cores, as well as sufficient memory for the cores. Thus, each server can host multiple VMs, and each VM generates its own traffic.
[0041] To accommodate the large traffic volume in the data center, a high-capacity switching fabric 170 can be provided. The switching fabric 170 is illustrated as a "flat" network in this example, where each server 146 can have a direct connection to a top-of-rack (ToR) switch 120 (e.g., a "star" configuration), and each ToR switch 120 can be coupled to a core switch 130. This two-layer flat network architecture is shown only as an illustrative example. In other examples, other architectures can be used, as non-limiting examples, such as a three-layer star or leaf-spine (also known as "fat tree" topology) based on the "Clos" architecture, a hub-and-spoke topology, a mesh topology, a ring topology, or a 3-D mesh topology.
[0042] The fabric itself can be provided by any suitable interconnection. For example, each server 146 can include an Intel host fabric interface (HFI), a network interface card (NIC), or other host interface. The host interface itself can be coupled to one or more processors via an interconnection or bus (such as PCI, PCIe, etc.), and in some cases, this interconnection bus can be considered part of the fabric 170.
[0043] The interconnection technology can be provided by a single interconnection or a hybrid interconnection, such as in the case where PCIe provides on-chip communication, 1 Gb or 10 Gb copper Ethernet provides a relatively short connection to the ToR switch 120, and fiber optic cables provide a relatively long connection to the core switch 130. As non-limiting examples, the interconnection technologies include Intel Omni-Path TM (Omni-Path TM )、TrueScale TM (TrueScale TM )、UltraPath Interconnect (UPI) (previously known as QPI or KTI), Fibre Channel, Ethernet, Fibre Channel over Ethernet (FCoE), InfiniBand, PCI, PCIe, or fiber optic, just to name a few. Some of these interconnection technologies will be more suitable for certain deployments or functions than others, and selecting the appropriate fabric for an immediate application is the practice of those of ordinary skill in the art.
[0044] However, note that while Omni-Path TMHigh - end structures such as this, but more generally, structure 170 can be any suitable interconnection or bus for a particular application. In some cases, this may include traditional interconnections such as local area networks (LANs), token ring networks, synchronous optical networks (SONETs), asynchronous transfer mode (ATM) networks, wireless networks (such as WiFi and Bluetooth), "plain old telephone system" (POTS) interconnections, or the like. It is also explicitly anticipated that new network technologies will emerge in the future to supplement or replace some of the technologies listed here, and any such future network topologies and technologies can be part of or form part of structure 170.
[0045] In some embodiments, as originally outlined in the OSI seven - layer network model, structure 170 can provide communication services at various "layers". In contemporary practice, the OSI model is not strictly followed. Generally, layer 1 and layer 2 are often referred to as the "Ethernet" layers (although in large data centers, Ethernet has typically been replaced by newer technologies). Layer 3 and layer 4 are often referred to as the Transmission Control Protocol / Internet Protocol (TCP / IP) layers (which can be further subdivided into TCP and IP layers). Layers 5 through 7 can be referred to as the "application layer". These layer definitions are presented as a useful framework but are not intended to be restrictive.
[0046] Figure 2 is a block diagram of data center 200 according to one or more examples of this specification. In embodiments, data center 200 can be the same as Figure 1 network 100, or can be a different data center. Figure 2 Additional views are provided to illustrate different aspects of data center 200.
[0047] In this example, structure 270 is provided to interconnect aspects of data center 200, including VMM 260. VMM 260 can be a virtual machine manager, hypervisor, domain 0, or other management entity for virtual machines. Structure 270 can be the same as Figure 1 structure 170, or can be a different structure. As described above, structure 270 can be provided by any suitable interconnection technology. In this example, Intel Omni - Path TM is used as an illustrative and non - restrictive example.
[0048] As shown, data center 200 includes a plurality of logical elements that form a plurality of nodes. It should be understood that each node can be provided by a physical server, a group of servers, or other hardware. Each server can run one or more virtual machines suitable for its application.
[0049] Node 0 208 is a processing node that includes processor socket 0 and processor socket 1. The processor can be, for example, an Intel Xeon TM (Xeon TM ) processor. Node 0 208 can be configured to provide network or workload functions, for example, by hosting multiple virtual machines or virtual devices.
[0050] On-board communication between processor socket 0 and processor socket 1 can be provided via on-board uplink 278. This can provide a very high-speed, short-length interconnection between the two processor sockets so that virtual machines running on node 0 208 can communicate with each other at a very high speed. To facilitate this communication, a virtual switch (vSwitch) can be provisioned on node 0 208, which can be considered part of fabric 270.
[0051] Node 0 208 is connected to fabric 270 via HFI 272. HFI 272 can be connected to an Intel Omni-Path TM fabric. In some examples, communication with fabric 270 can be tunneled, such as by providing UPI tunneling on the Omni-Path TM fabric.
[0052] Because data center 200 can provide many functions in a distributed manner that was provided on-board in previous generations, a high-capability HFI 272 can be provided. HFI 272 can operate at speeds of multiple gigabits per second and, in some cases, can be tightly coupled to node 0 208. For example, in some embodiments, the logic for HFI 272 is integrated directly with the processors on the system-on-chip. This provides very high-speed communication between HFI 272 and the processor sockets without the need for an intermediate bus device, which might introduce additional latency in the fabric. However, this does not mean that embodiments that provide HFI 272 on a traditional bus are excluded. On the contrary, it is explicitly anticipated that in some examples, HFI 272 can be provided on a bus, such as a PCIe bus, which is a serial version of PCI that provides higher speeds than traditional PCI. Throughout data center 200, various nodes can provide different types of HFI 272, such as on-board HFI and plug-in HFI. It should also be noted that certain blocks in the system-on-chip can be provided as intellectual property (IP) blocks, which can be "dropped" into an integrated circuit as modular units. Thus, in some cases, HFI 272 can be derived from such IP blocks.
[0053] Note that in the "network as a device" approach, node 0 208 may provide limited on-board memory or storage devices or no on-board memory or storage devices. Instead, node 0 208 may rely primarily on distributed services such as memory servers and network storage servers. On-board, node 0 208 may only provide enough memory and storage devices to boot the device and enable it to communicate with fabric 270. This distributed architecture is possible because of the very high speeds of contemporary data centers, and it may be advantageous because resources do not need to be over-provisioned for each node. Instead, large amounts of high-speed or specialized memory can be dynamically provisioned among multiple nodes such that each node can access a large amount of resources, but these resources are not left idle when that particular node does not need them.
[0054] In this example, node 1 memory server 204 and node 2 storage server 210 provide the operating memory and storage capabilities for node 0 208. For example, memory server node 1 204 may provide Remote Direct Memory Access (RDMA), whereby node 0 208 can access the memory resources on node 1 204 in DMA fashion via fabric 270, similar to how it would access its own on-board memory. The memory provided by memory server 204 may be volatile conventional memory (such as double data rate type 3 (DDR3) dynamic random access memory (DRAM)) or it may be a more unique type of memory that operates at a speed similar to DRAM but is non-volatile (such as persistent fast memory (PFM), such as Intel Cross Point TM (3DXP)).
[0055] Similarly, instead of providing an on-board hard drive for node 0 208, a storage server node 2 210 may be provided. Storage server 210 may provide networked block of disks (NBOD), PFM, redundant array of independent disks (RAID), redundant array of independent nodes (RAIN), network attached storage (NAS), optical storage, tape drives, or other non-volatile memory solutions.
[0056] Thus, when performing its designated functions, node 0 208 can access memory from memory server 204 and store the results on storage provided by storage server 210. Each of these devices is coupled to fabric 270 via HFI 272, which provides the fast communication that makes these technologies possible.
[0057] By further illustration, node 3 206 is also depicted. Node 3 206 also includes an HFI 272, as well as two processor sockets connected internally via an uplink. However, unlike node 0 208, node 3 206 includes its own on-board memory 222 and storage device 250. Thus, node 3 206 can be configured to perform its functions primarily on-board and may not need to rely on the memory server 204 and storage server 210. However, in appropriate circumstances, node 3 206 can supplement its own on-board memory 222 and storage device 250 with distributed resources similar to node 0 208.
[0058] The basic building blocks of the various components disclosed herein may be referred to as "logical elements". Logical elements can include hardware (e.g., including software programmable processors, ASICs, or FPGAs), external hardware (digital signals, analog signals, or mixed signals), software, reciprocating software, services, drivers, interfaces, components, modules, algorithms, sensors, components, firmware, microcode, programmable logic, or objects that can be coordinated to achieve logical operations. Additionally, some logical elements are provided by tangible, non-transitory computer-readable media having executable instructions stored thereon for instructing a processor to perform a certain task. Such non-transitory media can include, for example, hard drives, solid state memories or disks, read-only memories (ROMs), persistent fast memories (PFMs) (e.g., Intel Intersection TM )), external storage devices, redundant arrays of independent disks (RAIDs), redundant arrays of independent nodes (RAINs), network-attached storage (NAS), optical storage, tape drives, backup systems, cloud storage, or any combination of the foregoing as non-limiting examples. Such media can also include instructions programmed into an FPGA or instructions encoded in the hardware of an ASIC or processor.
[0059] Figure 3 is a block diagram of a computing system 300 according to one or more examples of the present specification. In this example, the computing system 300 is configured to provide one or more trust domains on a host platform 302. As described above, each trust domain 312 can have its own private key 314. Each private key 314 can be used to sign, encrypt, and decrypt any memory pages "owned" by the TD.
[0060] In this example, the host platform 302 includes one or more processors 304, and the one or more processors 304 include a multi-key total memory encryption (MKTME) engine 306. As described in this specification, the MKTME engine 306 is configured to provide encrypted memory pages that are owned by the corresponding TD 312 and are invisible to the VMM 330 that manages the virtual machines.
[0061] In various embodiments, the MKTME engine 306 can be provided in hardware, such as in the form of instructions directly encoded in silicon or in microcode. In other embodiments, the MKTME engine 306 can also be provided in read-only memory, flash memory, or software running in a protected memory region (such as in a trusted execution environment (TEE)).
[0062] The host platform 302 also includes physical memory 308 that can be partitioned among the various virtual machines. Thus, the VMM 330 provisions a trust domain 312 with its own private key 314 and allocates one or more memory pages (i.e., 4KB memory pages) to the TD 312. The VMM 330 includes an extended memory mapping that includes a host physical mapping 320 that maps host physical addresses to guest physical addresses for each TD 312.
[0063] For example, TD 312-1 has private key 314-1, TD 312-2 has private key 314-2, and TD 312-3 has private key 314-3. When the VMM 330 provisions these TDs 312, it provides a match of HPM to GPM for each TD 312. For example, HPM 320-1 maps to GPM 316-1. GPM 316-1 includes a memory block owned by TD 312-1. Similarly, HPM 320-2 maps to GPM 316-2, where HPM 320-2 maps a memory block owned by TD 312-2.
[0064] HPM 320-3 maps to GPM 316-2. HPM 320-3 is a mapping of one or more memory pages owned by TD 312-3.
[0065] In this example, each TD 312 can provision one or more virtual machines residing within the memory owned by the TD. Each TD also includes its own guest physical memory to guest virtual memory mapping. For example, TD 312-1 includes GPM 316-1 mapped to GVM 322-1. TD 312-2 includes a mapping of GPM 316-2 to GVM 322-2. TD 312-3 includes GPM 316-3 mapped to GVM 322-3.
[0066] As shown here, the guest memory mapping tables are all independent of each other and are invisible to the VMM 330. Thus, the VMM 330 cannot view or interfere with the state of any TD 312. Additionally, a compromised or malicious TD or an untrusted VM on the host platform 302 cannot view or interfere with any TD 312.
[0067] Thus, the only attack vector available to a compromised host is to attack the extended memory mapping on the VMM 330.
[0068] As described above, in some prior implementations, the extended memory table that maps the HPM 320 to the GPM 316 is protected by placing it in a special memory location dedicated to the microarchitecture.
[0069] However, embodiments of the MKTME engine 306 of this specification are configured to provide a CMOT that can be securely placed in main memory and still maintain the integrity of the TD 312. This can be achieved by placing the extended memory table into a memory region controlled by the CPU via the PMRR, such that the memory range cannot be accessed by software processes. To further protect the CMOT, the CMOT is a microarchitecture structure that cannot be directly accessed by software. Additionally, the CMOT includes security attributes, such as a TDCS pointer that provides a reverse mapping of the GPM 316 to HPM mapping. Thus, when the MKTME engine 306 walks through the memory pages, it can check the TDCS pointer and ensure that the mapping remains consistent.
[0070] To further ensure that the CMOT is not corrupted, an integrity field can be provided, which can include a cryptographic hash of the entries signed by the corresponding private key 314 of the owning TD. Thus, the MKTME engine 306 can be confident that the memory pages have not been tampered with, because if the memory pages have been tampered with, the cryptographic hash will no longer be valid.
[0071] An example of the CMOT is provided below. This example should be understood as a non - limiting and illustrative example, and it should be understood that embodiments of the teachings of this specification can make different forms of CMOTs that still implement the teachings of this specification.
[0072]
[0073] Figure 4 is a block diagram of a computing system 400 that illustrates additional aspects of the teachings of this specification.
[0074] In this example, the hardware platform 404 includes a CMOT 408 and an MKTME engine 406. The hardware platform 404 provides a VMM 412, which is configured to provide one or more trusted domains 440.
[0075] When the VMM 412 provisions multiple virtual machines or trusted domains, it establishes an extended page table 416 for each virtual machine or trusted domain. In the case of an untrusted VM (such as the untrusted VM 428), in addition to or instead of the EPT 420, the VMM 412 can provision a virtual machine control structure (VMCS).
[0076] When the VMM 412 provisions an untrusted VM 428, it has visibility not only into the EPT of the untrusted VM but also into the internal memory mapping of the untrusted VM. This includes the GPM-to-GVM mapping. This gives the VMM 412 full visibility into and control over the untrusted VM 428. In the case of an untrusted VM 428, the VMM 412 provisions a VMCS and an EPT 420 for the untrusted VM 428. The EPT 420 includes the HPM-to-GPM mapping, and the VMCS can include the GPM-to-GVM mapping. Although this applies to many computing tasks, some tasks require the establishment of trust domains so that the operations of the VMM can be kept private and protected from the influence of other VMs and from the VMM itself. This ensures that the CSP cannot tamper with the trusted VMs and that it cannot access privileged information. Thus, in this example, the VMM 412 provisions trust domains, namely TD 1 440-1 and TD 2 440-2. The VMM 412 establishes an EPT 416-1 for TD 1 440-1 and an EPT 416-2 for TD 2 440-2. The VMM 412 must have visibility into and control over the EPT416. However, to ensure that a malicious VM or a corrupted VMM does not tamper with the EPT416, the CMOT 408 is provided in the memory of the hardware platform 404 and can include useful fields (such as a TDCS pointer and an integrity field) to ensure that the EPT 416 has not been modified.
[0077] Each TD 440 includes its own key domain 424. For example, TD 1 440-1 includes KD 1 424-1, and TD 2440-2 includes KD 2 424-2.
[0078] Each TD 440 can provision one or more trusted VMs (TVMs) 432. For each TVM 432, the TD 440 can also provision a trusted domain control structure (TDCS) 430 that provides the GPM-to-GVM mapping for the TVM 432.
[0079] Thus, within TD 1 440-1, TVM 1 432-1 is provisioned with TDCS 430-1.
[0080] Within TD 2 440, TVM 2 432-2, TVM 3 432-3, and TVM 4 432-4 are supplied. The TD can supply the corresponding TDCS 430 for each TVM 432. For example, TDCS 430-2 is supplied for TVM 2 432-2. TDCS 430-3 is supplied for TVM 3 432-3. TDCS 430-4 is supplied for TVM 4 432-4.
[0081] As shown in the figure, each TD 440 maintains its own separate KD 424. The separate KD 424 includes one or more private keys that the TD 440 can use to encrypt, decrypt, and sign the memory owned by the TD. As shown above, the MKTME engine 406 employs CMOT 408 to ensure that the EPT 416 is not damaged or tampered with. It also prevents attacks such as replay attacks and other interferences.
[0082] Figure 5 It is a flowchart of method 500 that can be executed in combination with the teachings of this specification.
[0083] Starting from box 504, the processor may encounter a page miss. Based on the page miss, the PMH can traverse this paging structure, for example, starting from the EPTP.
[0084] In decision box 508, a processor such as the MKTME engine can determine whether there is a paging structure configuration error. If there is, an error condition may be triggered, and for example, in box 512, the VM can exit.
[0085] In box 516, when the processor is traversing the memory page, it can read the CMOT entry for any HPM mapping it finds using the TD key ID in its physical address space.
[0086] In decision box 520, when such a CMOT entry is encountered, the MKTME engine can determine whether the CMOT GPA integrity check matches the entry in the CMOT. This can include ensuring that the TDCS pointer has the correct value and ensuring that the CMOT entry has not been tampered with via the integrity field.
[0087] If the integrity check fails, then in box 524, an error condition may be triggered, such as exiting the TD.
[0088] In box 528, the MKTME or other control structure can determine the current address space ID (ASID) label assigned to the current KD for the CMOT entry.
[0089] In decision block 532, it is determined whether CMOT specifies a key different from the key determined in block 528. If not, then in block 536, the processor fills the TLB with the key ID ASID (e.g., using a k-bit offset as specified in the CMOT entry).
[0090] Return to block 532. If CMOT specifies a different key, then in block 540, the processor replaces the upper physical address bits with the specified key domain identifier (KDID).
[0091] In block 598, the processor sets the TLB with the given address and ASID tag. And the process is complete.
[0092] Note that in the flow 500 shown above, for the TD being executed, the PMH can normally walk through the page table and extended page table. During the terminal walkthrough, the PMH accesses the CMOT entry associated with the physical address found from the page walkthrough. It accesses the CMOT entry using the private key ID of the TD being executed. Then, as in decision block 520, it checks the integrity value of the CMOT entry to ensure it has not been corrupted.
[0093] CMOT encrypts the mapping of the physical address to the guest physical address using the key of the trust domain. The host (VMM) VMX root can include instructions for granting permissions to CMOT entries, including novel instructions such as TDADDPAGE and TDREVOKEPAGE. The tenant's TD ensures that the CMOT entry encrypted with the key is correct. The TDADD instruction uses the KeyID (key ID) of the TD. If the CMOT entry integrity is corrupted, a TD exit may occur. Otherwise, it verifies the entry using the HPA to GPA mapping.
[0094] According to an embodiment of the present specification, the hardware PMH provides an EPT walkthrough for switching to the KeyID of the VMM. The PMH EPT walkthrough ends with a lookup in the CMOT table using the HPA as an index. The PMH can use the exclusive KeyID or private KeyID of the TD to access the CMOT entry (attached to the HPA). The PMH also verifies the integrity of the entry. If the CMOT entry confirms the HPA to GPA mapping, then it caches the TLB. Otherwise, it exits. If the entry is a shared page, it can indicate the key to be used, otherwise it can be mapped to a larger page. It can also verify permissions and other tenant policies.
[0095] KeyID TDCS GPA License Encryption Version Valid Status 1 1 5 R / X / W N 3 Y A 2 2 3 R / W N 2 Y A 3 3 1 R / W Y 1 Y A 1 4 4 R N 3 Y A 2 5 2 X N 2 N F
[0096] The above table illustrates how CMOT table entries can be encrypted using the private key or KeyID of the TD. The MKTME engine encrypts the memory at the cache line granularity, which suggests that the size of a CMOT entry can be 512 bits to conform to the cache line. However, since in some embodiments this structure is exclusively managed by the CPU, it is possible to use partial permissions and limit each entry to the AES block size (such as 128 bits). Thus, the table can be aligned such that each entry maps to a 128-bit aligned TME block size. Corruption of any part of the AES block will similarly corrupt all bits, resulting in corruption of the integrity value. When the PMH causes a CMOT entry to be read, the MKTME can use the private KeyID of the TD in the address to decrypt the entire 512-bit memory row. If different keys are used to encrypt and partially write different entries on the row, those entries are corrupted, but the specific entry accessed and read by the PMH with the correct KeyID is correct.
[0097] In some embodiments, in addition to or instead of the integrity value, a version number can be specified. This allows the CPU to avoid replaying entries. When a conflicting entry is observed, such as changing a shared CMOT entry or a plaintext CMOT entry to an entry using a private or exclusive TD key, the CPU can increment the version number of all CMOT entries belonging to that TD. Then, the PMH can use the current version of the TD (e.g., the protected CPU register counter of the TD) and compare this value with the value stored in the MOT entry to ensure that the MOT entry has not been replayed from some previous state.
[0098] Finally, other values can be used for integrity or to prevent replay. For example, the current context or CR3 value of the TD can be stored in the MOT entry field to bind the entry to a specific CR3 or context. Similarly, a linear address (LA) can be specified in the table, where the PMH can check whether the LA (in addition to the GPA) corresponds to the LA specified in the MOT entry and otherwise exit. Since not all CMOT entries can always have their context / LA checked, a bit field can be used to determine whether the LA or context should be checked for a particular CMOT entry. This can be useful for input / output operations.
[0099] Certain of the following figures detail example architectures and systems for implementing the embodiments above. In some embodiments, one or more of the hardware components and / or instructions described above are emulated as detailed below or implemented as software modules.
[0100] In some examples, the (multiple) instructions may be embodied in a "general vector friendly instruction format" detailed below. In other embodiments, another instruction format is used. The descriptions below of write mask registers, various data transforms (e.g., blend, broadcast, etc.), addressing, etc. generally apply to the descriptions of the embodiments of the (multiple) instructions above. Additionally, example systems, architectures, and pipelines are detailed below. Multiple embodiments of the (multiple) instructions above may execute on those systems, architectures, and pipelines, but are not limited to those detailed.
[0101] An instruction set may include one or more instruction formats. A given instruction format may define various fields (e.g., number of bits, position of bits) to specify the operation to be performed (e.g., opcode) and the (multiple) operands and / or (multiple) other data fields (e.g., mask) on which the operation is to be performed, etc. Some instruction formats are further decomposed by the definition of instruction templates (or sub-formats). For example, an instruction template of a given instruction format may be defined as having different subsets of the fields of that instruction format (the included fields are typically in the same order, but at least some fields have different bit positions as fewer fields are included), and / or defined as having a given field interpreted in a different way. Thus, each instruction of the ISA is expressed using a given instruction format (and if defined, according to a given one of the instruction templates in that instruction format) and includes fields for specifying the operation and operands. In one embodiment, an example ADD instruction has a specific opcode and instruction format (the instruction format includes an opcode field for specifying the opcode and an operand field for selecting the operands (source1 / destination and source2)); and the appearance of the ADD instruction in the instruction stream will result in specific contents in the operand field for selecting the specific operands. Advanced Vector Extensions (AVX) (AVX1 and AVX2) and a set of SIMD extensions using the Vector Extension (VEX) encoding scheme have been introduced and / or published (see, e.g., the Intel and IA-32 Architecture Software Developer's Manual of September 2014; and see the Intel Advanced Vector Extensions Programming Reference of October 2014).
[0102] Example Instruction Formats
[0103] Embodiments of the (multiple) instructions described herein can be embodied in different formats. Additionally, example systems, architectures, and pipelines are detailed below. Embodiments of the (multiple) instructions may execute on such systems, architectures, and pipelines, but are not limited to those detailed.
[0104] General Vector Friendly Instruction Format
[0105] A vector-friendly instruction format is an instruction format suitable for vector instructions (e.g., there are certain fields dedicated to vector operations). Although embodiments are described in which both vector and scalar operations are supported by the vector-friendly instruction format, alternative embodiments use only vector operations with the vector-friendly instruction format.
[0106] Figures 6a - 6b is a block diagram illustrating a general vector-friendly instruction format and its instruction templates according to an embodiment of the present specification. Figure 6a is a block diagram illustrating a general vector-friendly instruction format and its class A instruction templates according to an embodiment of the present specification; and Figure 6b is a block diagram illustrating a general vector-friendly instruction format and its class B instruction templates according to an embodiment of the present specification. Specifically, class A and class B instruction templates are defined for the general vector-friendly instruction format 600, both of which include instruction templates for memoryless access 605 and instruction templates for memory access 620. The term "general" in the context of the vector-friendly instruction format refers to an instruction format that is not tied to any specific instruction set.
[0107] Embodiments of the present specification will be described in which the vector-friendly instruction format supports the following: a 64-byte vector operand length (or size) with a 32-bit (4-byte) or 64-bit (8-byte) data element width (or size) (and thus, a 64-byte vector consists of 16 double-word-sized elements, or alternatively, 8 quad-word-sized elements); a 64-byte vector operand length (or size) with a 16-bit (2-byte) or 8-bit (1-byte) data element width (or size); a 32-byte vector operand length (or size) with a 32-bit (4-byte), 64-bit (8-byte), 16-bit (2-byte), or 8-bit (1-byte) data element width (or size); and a 16-byte vector operand length (or size) with a 32-bit (4-byte), 64-bit (8-byte), 16-bit (2-byte), or 8-bit (1-byte) data element width (or size); however, alternative embodiments may support larger, smaller, and / or different vector operand sizes (e.g., a 256-byte vector operand) with larger, smaller, or different data element widths (e.g., a 128-bit (16-byte) data element width).
[0108] Figure 6a The class A instruction templates in include: 1) within the instruction template for memoryless access 605, an instruction template showing a fully rounded control-type operation 610 without memory access and an instruction template for a data transformation-type operation 615 without memory access; and 2) within the instruction template for memory access 620, an instruction template showing the timeliness 625 of memory access and an instruction template for non-timeliness 630 of memory access. Figure 6bThe Class B instruction templates in it include: 1) Instruction templates showing partial rounding control type operations 612 with write mask control for memoryless access and instruction templates showing VSIZE type operations 617 with write mask control for memoryless access within the instruction template of memoryless access 605; and 2) Instruction templates showing write mask control 627 for memory access within the instruction template of memory access 620.
[0109] The general vector friendly instruction format 600 includes the following fields listed in the order illustrated in Figure 6a – Figure 6b as follows.
[0110] Format field 640 - The specific value (instruction format identifier value) in this field uniquely identifies the vector friendly instruction format and thereby identifies that the instruction appears in the instruction stream in vector friendly instruction format. Thus, this field is optional in the sense that it is not required for an instruction set that only has the general vector friendly instruction format.
[0111] Base operation field 642 - Its content differentiates different base operations.
[0112] Register index field 644 - Its content directly or through address generation specifies the location of source and destination operands in registers or in memory. These fields include a sufficient number of bits to select N registers from a PxQ (e.g., 32x512, 16x128, 32x1024, 64x1024) register file. Although in one embodiment N can be up to three source registers and one destination register, alternative embodiments can support more or fewer source and destination registers (e.g., can support up to two source registers, where one of these source registers also serves as the destination register; can support up to three source registers, where one of these source registers also serves as the destination register; or can support up to two source registers and one destination register).
[0113] Modifier field 646 - Its content differentiates instructions that appear in general vector instruction format specifying memory access from those that do not specify memory access in general vector instruction format; i.e., it differentiates between the instruction template of memoryless access 605 and the instruction template of memory access 620. Memory access operations read and / or write to the memory hierarchy (in some cases, using values in registers to specify source and / or destination addresses), while non-memory access operations do not (e.g., source and destination are registers). Although in one embodiment, this field also selects between three different ways to perform memory address calculation, alternative embodiments can support more, fewer, or different ways to perform memory address calculation.
[0114] Expansion operation field 650 - whose content differentiates which one of various different operations in addition to the base operation is to be performed. This field is context-dependent. In one embodiment of the present specification, this field is divided into a class field 668, an alpha field 652, and a beta field 654. The expansion operation field 650 allows multiple sets of common operations to be performed in a single instruction rather than in 2, 3, or 4 instructions.
[0115] Scale field 660 - whose content allows the content of the index field used for memory address generation (e.g., for address generation using 2 比例 * index + base address) to be scaled.
[0116] Displacement field 662A - whose content is used as part of memory address generation (e.g., for address generation using 2 比例 * index + base address + displacement).
[0117] Displacement factor field 662B (note that the juxtaposition of displacement field 662A directly above displacement factor field 662B indicates the use of one or the other) - whose content is used as part of address generation, which specifies a displacement factor scaled by the size (N) of the memory access, where N is the number of bytes in the memory access (e.g., for address generation using 2 比例 * index + base address + scaled displacement). Redundant low-order bits are ignored, and thus the content of the displacement factor field is multiplied by the total size (N) of the memory operand to generate the final displacement to be used in calculating the effective address. The value of N is determined by the processor hardware at runtime based on the full opcode field 674 (described later in this document) and the data manipulation field 654C. The displacement field 662A and the displacement factor field 662B are optional in the sense that they are not used in instruction templates without memory access 605 and / or different embodiments may implement only one of the two or neither of the two.
[0118] Data element width field 664 - whose content differentiates which one of multiple data element widths is to be used (in some embodiments for all instructions; in other embodiments only for some instructions). This field is optional in the sense that it is not needed if only one data element width is supported and / or some aspect of the opcode is used to support the data element width.
[0119] Write mask field 670 - whose content controls, on a per data element position basis, whether the data element positions in the destination vector operand reflect the results of the base operation and the expansion operation. Class A instruction templates support a merge-write mask, while class B instruction templates support both a merge-write mask and a zero-write mask. When merged, the vector mask allows any set of elements in the destination to be protected from update during the execution of any operation (specified by the base operation and the expansion operation) - in one embodiment, preserving the old value of each element of the destination where the corresponding mask bit has a value of 0. In contrast, when zeroed, the vector mask allows any set of elements in the destination to be zeroed during the execution of any operation (specified by the base operation and the expansion operation), and in one embodiment, the elements of the destination are set to 0 when the corresponding mask bit has a value of 0. A subset of this functionality is the ability to control the vector length of the operation being performed (i.e., the span from the first to the last element being modified); however, the elements being modified do not have to be contiguous. Thus, the write mask field 670 allows partial vector operations, which include loads, stores, arithmetic, logic, etc. Although embodiments of this specification have been described in which the content of the write mask field 670 selects one write mask register out of a plurality of write mask registers that contains the write mask to be used (and thus, the content of the write mask field 670 indirectly identifies the mask to be executed), alternative embodiments alternatively or additionally allow the content of the mask write field 670 to directly specify the mask to be executed.
[0120] Immediate number field 672 - whose content allows the specification of an immediate number. This field is optional in the sense that it does not exist in a general vector-friendly format that does not support immediate numbers and does not exist in instructions that do not use immediate numbers.
[0121] Class field 668 - whose content differentiates between different classes of instructions. Refer to Figures 6a - 6b , the content of this field selects between class A instructions and class B instructions. In Figures 6a - 6b , rounded rectangles are used to indicate the presence of dedicated values in the field (e.g., in Figures 6a - 6b , 668A for class A and 668B for class B, respectively, for class field 668).
[0122] Class A instruction template
[0123] In the case of the instruction template for Class A non-memory access 605, the α field 652 is interpreted as an RS field 652A whose content differentiates which of different augmentation operation types is to be performed (e.g., for the instruction templates of rounding-type operations 610 without memory access and data transformation-type operations 615 without memory access, rounding 652A.1 and data transformation 652A.2 are specified respectively), while the β field 654 differentiates which of the operations of the specified type is to be performed. In the instruction template for non-memory access 605, the scale field 660, displacement field 662A, and displacement scale field 662B do not exist.
[0124] Instruction template for non-memory access - fully rounding control type operation
[0125] In the instruction template for the fully rounding control type operation 610 without memory access, the β field 654 is interpreted as a rounding control field 654A whose content provides static rounding. Although in the described embodiments of this specification the rounding control field 654A includes a suppress all floating-point exceptions (SAE) field 656 and a rounding operation control field 658, alternative embodiments may encode these two concepts into the same field, or have only one or the other of these concepts / fields (e.g., may have only the rounding operation control field 658).
[0126] SAE field 656 - whose content differentiates whether to disable the reporting of exception events; when the content of the SAE field 656 indicates enabling suppression, a given instruction does not report any kind of floating-point exception flag and does not invoke any floating-point exception handler.
[0127] Rounding operation control field 658 - whose content differentiates which of a set of rounding operations is to be performed (e.g., round up, round down, round towards zero, and round to nearest). Thus, the rounding operation control field 658 allows the rounding mode to be changed instruction by instruction. In one embodiment of this specification in which the processor includes a control register for specifying the rounding mode, the content of the rounding operation control field 650 overrides the register value.
[0128] Instruction template for non-memory access - data transformation type operation
[0129] In the instruction template for the data transformation type operation 615 without memory access, the β field 654 is interpreted as a data transformation field 654B whose content differentiates which of multiple data transformations is to be performed (e.g., no data transformation, mix, broadcast).
[0130] In the case of the instruction template for Class A memory access 620, the α field 652 is interpreted as an eviction hint field 652B, whose content differentiates which eviction hint is to be used (in Figure 6aIn it, for the instruction template of memory access timeliness 625 and the instruction template of memory access non - timeliness 630, the timeliness 652B.1 and the non - timeliness 652B.2 are specified respectively), and the β field 654 is interpreted as a data manipulation field 654C, the content of which differentiates which one of multiple data manipulation operations (also called primitives) is to be executed (for example, no manipulation, broadcast, up - conversion of the source, and down - conversion of the destination). The instruction template of memory access 620 includes a scale field 660 and optionally includes a displacement field 662A or a displacement - scale field 662B.
[0131] Vector memory instructions use conversion support to perform vector loads from memory and vector stores to memory. Like ordinary vector instructions, vector memory instructions transfer data to / from memory in a data - element - by - data - element manner, where the actually transferred elements are specified by the content of the vector mask selected as the write mask.
[0132] Instruction template for memory access - Timeliness
[0133] Timely data is data that may be reused fast enough to benefit from cache operations. However, this is a hint, and different processors can implement it in different ways, including completely ignoring the hint.
[0134] Instruction template for memory access - Non - timeliness
[0135] Non - timely data is data that cannot be reused fast enough to benefit from cache operations in the level - 1 cache and should be given eviction priority. However, this is a hint, and different processors can implement it in different ways, including completely ignoring the hint.
[0136] Class B instruction template
[0137] In the case of the Class B instruction template, the α field 652 is interpreted as a write - mask control (Z) field 652C, the content of which differentiates whether the write - mask operation controlled by the write - mask field 670 should be merge or zero.
[0138] In the case of the instruction template for B - type non - memory access 605, a portion of the β field 654 is interpreted as the RL field 657A, the content of which differentiates which one of different extended operation types is to be performed (for example, the instruction templates for the partial rounding - control type operation 612 with write - mask control for non - memory access and the VSIZE - type operation 617 with write - mask control for non - memory access specify rounding 657A.1 and vector length (VSIZE) 657A.2 respectively), while the remaining portion of the β field 654 differentiates which one of the operations of the specified type is to be performed. In the non - memory access 605 instruction template, the scale field 660, the displacement field 662A, and the displacement - scale field 662B do not exist.
[0139] In the instruction template for the partial rounding - control type operation 610 with write - mask control for non - memory access, the remaining portion of the β field 654 is interpreted as the rounding - operation field 659A, and abnormal - event reporting is disabled (the given instruction does not report any kind of floating - point exception flag and does not invoke any floating - point exception handler).
[0140] The rounding - operation control field 659A - like the rounding - operation control field 658, the content of which differentiates which one of a set of rounding operations is to be performed (for example, rounding up, rounding down, rounding towards zero, and rounding to nearest). Thus, the rounding - operation control field 659A allows the rounding mode to be changed instruction - by - instruction. In one embodiment of this specification in which the processor includes a control register for specifying the rounding mode, the content of the rounding - operation control field 650 overrides the register value.
[0141] In the instruction template for the VSIZE - type operation 617 with write - mask control for non - memory access, the remaining portion of the β field 654 is interpreted as the vector - length field 659B, the content of which differentiates which one of multiple data vector lengths is to be performed (for example, 628 bytes, 256 bytes, or 512 bytes).
[0142] In the case of the instruction template for B - type memory access 620, a portion of the β field 654 is interpreted as the broadcast field 657B, the content of which differentiates whether a broadcast - type data - manipulation operation is to be performed, while the remaining portion of the β field 654 is interpreted by the vector - length field 659B. The instruction template for the memory access 620 includes the scale field 660 and optionally includes the displacement field 662A or the displacement - scale field 662B.
[0143] For the general vector friendly instruction format 600, the complete opcode field 674 is shown to include a format field 640, a base operation field 642, and a data element width field 664. Although one embodiment is shown in which the complete opcode field 674 includes all these fields, in embodiments that do not support all these fields, the complete opcode field 674 includes fewer than all these fields. The complete opcode field 674 provides an operation code (opcode).
[0144] The extended operation field 650, the data element width field 664, and the write mask field 670 allow these features to be specified on a per-instruction basis in the general vector friendly instruction format.
[0145] The combination of the write mask field and the data element width field creates various types of instructions, as these instructions allow the mask to be applied based on different data element widths.
[0146] The various instruction templates that occur within classes A and B are beneficial in different scenarios. In some embodiments of this specification, different processors or different cores within a processor may support only class A, only support class B, or may support both classes. For example, a high-performance general out-of-order core intended for general computing may support only class B, a core intended primarily for graphics and / or scientific (throughput) computing may support only class A, and a core intended for both general computing and graphics and / or scientific (throughput) computing may support both class A and class B (of course, with some mix of templates and instructions from both classes, but not all templates and instructions from both classes are within the scope of this specification). Similarly, a single processor may include multiple cores, all of which support the same class, or in which different cores support different classes. For example, in a processor with separate graphics and general cores, one core in the graphics core intended primarily for graphics and / or scientific computing may support only class A, while one or more in the general core may be high-performance general out-of-order cores with register renaming that support only class B for general computing. Another processor without a separate graphics core may include one or more general in-order or out-of-order cores that support both class A and class B. Of course, in different embodiments of this specification, features from one class may also be implemented in other classes. This will enable programs written in a high-level language to be (e.g., just-in-time compiled or statically compiled) into various different executable forms, including: 1) a form that has only instructions of one or more classes supported by the target processor for execution; or 2) a form that has alternative routines written using different combinations of instructions from all classes and has control flow code that selects these routines for execution based on the instructions supported by the processor currently executing the code.
[0147] Example specialized vector friendly instruction format
[0148] Figure 7a is a block diagram showing an example specialized vector-friendly instruction format according to an embodiment of this specification. Figure 7a Shows a specialized vector-friendly instruction format 700, which is specialized in the sense that it specifies the positions, sizes, interpretations, and orders of the various fields, as well as the values of some of those fields. The specialized vector-friendly instruction format 700 can be used to extend the x86 instruction set, and thus some of the fields in this format are similar to or the same as those used in the existing x86 instruction set and its extensions (e.g., AVX). This format is consistent with the prefix coding field, real opcode byte field, MOD R / M field, SIB field, displacement field, and immediate number field of the existing x86 instruction set with extensions. The figure shows fields from Figure 6a and Figure 6b The fields from Figure 2 are mapped to the fields from Figure 6a and Figure 6b among others.
[0149] It should be understood that although the embodiments of this specification are described in the context of the specialized vector-friendly instruction format 700 with reference to the general vector-friendly instruction format 600 for illustrative purposes, this specification is not limited to the specialized vector-friendly instruction format 700 unless otherwise stated. For example, the general vector-friendly instruction format 600 contemplates various possible sizes for the various fields, while the specialized vector-friendly instruction format 700 is shown with fields of specific sizes. As a specific example, although the data element width field 664 is illustrated as a one-bit field in the specialized vector-friendly instruction format 700, this specification is not limited thereto (i.e., the general vector-friendly instruction format 600 contemplates other sizes for the data element width field 664).
[0150] The general vector-friendly instruction format 600 includes the following fields listed in the order illustrated in Figure 7a as follows.
[0151] EVEX prefix (bytes 0-3) 702 - Encoded in a four-byte form.
[0152] Format field 640 (EVEX byte 0, bits [7:0]) - The first byte (EVEX byte 0) is the format field 640, and it contains 0x62 (in one embodiment, the unique value for differentiating the vector-friendly instruction format).
[0153] The second through fourth bytes (EVEX bytes 1-3) include multiple bit fields that provide specialized capabilities.
[0154] REX field 705 (EVEX byte 1, bits [7-5]) - consists of the EVEX.R bit field (EVEX byte 1, bit [7] – R), the EVEX.X bit field (EVEX byte 1, bit [6] – X), and (157BEX byte 1, bit [5] – B). The EVEX.R, EVEX.X, and EVEX.B bit fields provide the same functionality as the corresponding VEX bit fields and are encoded in one's complement form, i.e., ZMM0 is encoded as 1111B and ZMM15 is encoded as 0000B. Other fields of these instructions encode the lower three bits (rrr, xxx, and bbb) of the register index as known in the art, such that Rrrr, Xxxx, and Bbbb can be formed by adding EVEX.R, EVEX.X, and EVEX.B.
[0155] REX' field 610 - This is the first part of the REX' field 610 and is the EVEX.R' bit field (EVEX byte 1, bit [4] – R') used to encode the upper 16 or lower 16 registers of the extended 32-register set. In one embodiment, this bit is stored in bit-reversed format together with other bits indicated below to distinguish (in the well-known x86 32-bit mode) from the BOUND instruction whose real opcode byte is 62, but does not accept the value 11 in the MOD field in the MOD R / M field (described below); other embodiments do not store this indicated bit and other bits indicated below in reversed format. The value 1 is used to encode the lower 16 registers. In other words, R'Rrrr is formed by combining EVEX.R', EVEX.R, and other RRRs from other fields.
[0156] Opcode mapping field 715 (EVEX byte 1, bits [3:0] – mmmm) - its content encodes the implicit leading opcode byte (0F, 0F 38, or 0F 3).
[0157] Data element width field 664 (EVEX byte 2, bit [7] – W) - denoted by the notation EVEX.W. EVEX.W is used to define the granularity (size) of the data type (32-bit data element or 64-bit data element).
[0158] EVEX.vvvv 720 (EVEX byte 2, bits [6:3] - vvvv) - The functions of EVEX.vvvv can include the following: 1) EVEX.vvvv encodes the first source register operand specified in inverted (1's complement) form and is valid for instructions with two or more source operands; 2) EVEX.vvvv encodes the destination register operand specified in 1's complement form for a specific vector displacement; or 3) EVEX.vvvv does not encode any operand, this field is reserved and shall contain 1111b. Thus, the EVEX.vvvv field 720 encodes the 4 low-order bits of the first source register specifier stored in inverted (1's complement) form. Depending on the instruction, additional different EVEX bit fields are used to extend the specifier size to 32 registers.
[0159] EVEX.U 168 class field (EVEX byte 2, bit [2] - U) - If EVEX.U = 0, it indicates class A or EVEX.U0; if EVEX.U = 1, it indicates class B or EVEX.U1.
[0160] Prefix encoding field 725 (EVEX byte 2, bits [1:0] - pp) - Provides additional bits for the base operation field. In addition to supporting traditional SSE instructions in EVEX prefix format, this also has the benefit of compressing SIMD prefixes (the EVEX prefix only requires 2 bits instead of a byte to express the SIMD prefix). In one embodiment, to support traditional SSE instructions using SIMD prefixes (66H, F2H, F3H) in both traditional format and EVEX prefix format, these traditional SIMD prefixes are encoded as SIMD prefix encoding fields; and are extended to traditional SIMD prefixes before being provided to the PLA at runtime (thus, the PLA can execute these traditional instructions in both traditional and EVEX formats without modification). Although newer instructions can directly use the content of the EVEX prefix encoding field as an opcode extension, for consistency, a particular embodiment extends it in a similar way but allows different meanings specified by these traditional SIMD prefixes. Alternative embodiments can redesign the PLA to support 2-bit SIMD prefix encoding and thus do not require extension.
[0161] α field 652 (EVEX byte 3, bit [7] – EH, also known as EVEX.eh, EVEX.rs, EVEX.rl, EVEX.write mask control, and EVEX.n; also shown as α in the diagram) - As previously described, this field is context-specific.
[0162] β field 654 (EVEX byte 3, bits [6:4] - SSS, also known as EVEX.s 2-0, EVEX.r 2-0 , EVEX.rr1, EVEX.LL0, EVEX.LLB; also shown as βββ) - as previously described, this field is context - specific.
[0163] REX’ field 610 - This is the remainder of the REX’ field and is the EVEX.V’ bit field (EVEX byte 3, bit [3] – V’) that can be used to encode the upper 16 or lower 16 registers of the extended 32 - register set. This bit is stored in bit - reversed format. The value 1 is used to encode the lower 16 registers. In other words, V’VVVV is formed by combining EVEX.V’ and EVEX.vvvv.
[0164] Write - mask field 670 (EVEX byte 3, bits [2:0] - kkk) - Its content specifies the register index in the write - mask register, as previously described. In one embodiment, the specific value EVEX.kkk = 000 has a special behavior that implies no write - mask is used for a particular instruction (this can be implemented in various ways, including using hardware hard - wired to all - ones write - mask or bypassing the mask hardware).
[0165] The real opcode field 730 (byte 4) is also called the opcode byte. Parts of the opcode are specified in this field.
[0166] The MOD R / M field 740 (byte 5) includes the MOD field 742, the Reg field 744, and the R / M field 746. As previously described, the content of the MOD field 742 distinguishes memory - access operations from non - memory - access operations. The role of the Reg field 744 can be boiled down to two cases: encoding the destination register operand or the source register operand; or being regarded as an opcode extension and not being used to encode any instruction operand. The role of the R / M field 746 can include the following: encoding the instruction operand that references a memory address; or encoding the destination register operand or the source register operand.
[0167] Scale, Index, Base (SIB) byte (byte 6) - As previously described, the content of the scale field 650 is used for memory - address generation. SIB.xxx 754 and SIB.bbb 756 - The content of these fields has been previously mentioned for register indices Xxxx and Bbbb.
[0168] Displacement field 662A (bytes 7 - 10) - When the MOD field 742 contains 10, bytes 7 - 10 are the displacement field 662A, and it works the same as the traditional 32 - bit displacement (disp32) and works at byte granularity.
[0169] Displacement factor field 662B (byte 7) - When the MOD field 742 contains 01, byte 7 is the displacement factor field 662B. The position of this field is the same as that of the 8-bit displacement (disp8) of the traditional x86 instruction set that works at the byte granularity. Since disp8 is sign-extended, it can only address between 128 and 127 byte offsets; in terms of a 64-byte cache line, disp8 uses 8 bits that can be set to only four truly useful values -128, -64, 0, and 64; since a larger range is often needed, disp32 is used; however, disp32 requires 4 bytes. In contrast to disp8 and disp32, the displacement factor field 662B is a reinterpretation of disp8; when using the displacement factor field 662B, the actual displacement is determined by multiplying the content of the displacement factor field by the size (N) of the memory operand access. This type of displacement is called disp8*N. This reduces the average instruction length (a single byte for the displacement but with a much larger range). Such compressed displacements are based on the assumption that the effective displacement is a multiple of the granularity of the memory access, and thus the redundant low-order bits of the address offset do not need to be encoded. In other words, the displacement factor field 662B replaces the 8-bit displacement of the traditional x86 instruction set. Thus, the displacement factor field 662B is encoded in the same way as the 8-bit displacement of the x86 instruction set (therefore, there is no change in the ModRM / SIB encoding rules), the only difference being that disp8 is overloaded to disp8*N. In other words, there is no change in the encoding rules or encoding length, but only a change in the interpretation of the displacement value by the hardware (which requires scaling the displacement by the size of the memory operand to obtain a byte-wise address offset). The immediate field 672 operates as previously described.
[0170] Full opcode field
[0171] Figure 7b is a block diagram showing the fields that make up the full opcode field 674 of the dedicated vector-friendly instruction format 700 according to one embodiment. Specifically, the full opcode field 674 includes a format field 640, a base operation field 642, and a data element width (W) field 664. The base operation field 642 includes a prefix encoding field 725, an opcode mapping field 715, and a real opcode field 730.
[0172] Register index field
[0173] Figure 7cFIG. is a block diagram of fields constituting a register index field 644 of a special vector-friendly instruction format 700 according to one embodiment. Specifically, the register index field 644 includes a REX field 705, a REX' field 710, a MODR / M.reg field 744, a MODR / M.r / m field 746, a VVVV field 720, an xxx field 754, and a bbb field 756.
[0174] Expansion operation field
[0175] Figure 7d FIG. is a block diagram of fields constituting an expansion operation field 650 of a special vector-friendly instruction format 700 according to one embodiment. When the class (U) field 668 contains 0, it indicates EVEX.U0 (class A 668A); when it contains 1, it indicates EVEX.U1 (class B 668B). When U = 0 and the MOD field 742 contains 11 (indicating no memory access operation), the α field 652 (EVEX byte 3, bit [7] – EH) is interpreted as the rs field 652A. When the rs field 652A contains 1 (rounding 652A.1), the β field 654 (EVEX byte 3, bits [6:4] – SSS) is interpreted as the rounding control field 654A. The rounding control field 654A includes a one-bit SAE field 656 and a two-bit rounding operation field 658. When the rs field 652A contains 0 (data transformation 652A.2), the β field 654 (EVEX byte 3, bits [6:4] – SSS) is interpreted as a three-bit data transformation field 654B. When U = 0 and the MOD field 742 contains 00, 01, or 10 (indicating a memory access operation), the α field 652 (EVEX byte 3, bit [7] – EH) is interpreted as the eviction hint (EH) field 652B and the β field 654 (EVEX byte 3, bits [6:4] – SSS) is interpreted as a three-bit data manipulation field 654C.
[0176] When U = 1, the α field 652 (EVEX byte 3, bit [7] – EH) is interpreted as the write mask control (Z) field 652C. When U = 1 and the MOD field 742 contains 11 (indicating no memory access operation), a part of the β field 654 (EVEX byte 3, bit [4] – S0) is interpreted as the RL field 657A; when it contains 1 (rounding 657A.1), the rest of the β field 654 (EVEX byte 3, bits [6-5] – S 2-1 ) is interpreted as the rounding operation field 659A, and when the RL field 657A contains 0 (VSIZE 654.A2), the rest of the β field 654 (EVEX byte 3, bits [6-5]-S 2-1 ) is interpreted as the vector length field 659B (EVEX byte 3, bits [6-5] – L1-0 )。When U = 1 and the MOD field 742 contains 00, 01, or 10 (indicating a memory access operation), the β field 654 (EVEX byte 3, bits [6:4] – SSS) is interpreted as a vector length field 659B (EVEX byte 3, bits [6 - 5] – L 1-0 ) and a broadcast field 657B (EVEX byte 3, bit [4] – B).
[0177] Example Register Architecture
[0178] Figure 8 is a block diagram of a register architecture 800 according to one embodiment. In the illustrated embodiment, there are 32 vector registers 810 that are 512 bits wide; these registers are referred to as zmm0 through zmm31. The lower 256 bits of the lower 16 zmm registers overlay the registers ymm0 - 16. The lower 128 bits of the lower 16 zmm registers (the lower 128 bits of the ymm registers) overlay the registers xmm0 - 15. The specialized vector-friendly instruction format 700 operates on these overlaid register banks, as illustrated in the following table.
[0179]
[0180]
[0181] In other words, the vector length field 659B selects between a maximum length and one or more other shorter lengths, where each such shorter length is half of the previous length; and instruction templates that do not have a vector length field 659B operate on the maximum vector length. Additionally, in one embodiment, the Class B instruction templates of the specialized vector-friendly instruction format 700 operate on packed or scalar single / double-precision floating-point data and packed or scalar integer data. A scalar operation is an operation performed on the lowest-order data element position in a zmm / ymm / xmm register; depending on the embodiment, the higher-order data element positions either remain the same as before the instruction or are zeroed.
[0182] Write Mask Registers 815 - In the illustrated embodiment, there are 8 write mask registers (k0 through k7), and each write mask register is 64 bits in size. In alternative embodiments, the write mask registers 815 are 16 bits in size. As previously described, in one embodiment, the vector mask register k0 cannot be used as a write mask; when the encoding that normally indicates k0 is used as a write mask, it selects the hard-wired write mask 0xFFFF, effectively disabling the write mask operation of the instruction.
[0183] General Purpose Registers 825 - In the illustrated embodiment, there are sixteen 64 - bit general purpose registers that are used with existing x86 addressing modes to address memory operands. These registers are referenced by the names RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 through R15.
[0184] Scalar Floating - Point Stack Register File (x87 Stack) 845, overlaid with MMX Packed Integer Flat Register File 850 - In the illustrated embodiment, the x87 stack is an eight - element stack for performing scalar floating - point operations on 32 / 64 / 80 - bit floating - point data using the x87 instruction set extension; and the MMX registers are used to perform operations on 64 - bit packed integer data and to save operands for some operations performed between MMX and XMM registers.
[0185] Other embodiments may use wider or narrower registers. Additionally, alternative embodiments may use more, fewer, or different register files and registers.
[0186] Example Core Architectures, Processors, and Computer Architectures
[0187] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, implementations of such cores can include: 1) general - purpose in - order cores intended for general - purpose computing; 2) high - performance general - purpose out - of - order cores intended for general - purpose computing; 3) specialized cores intended primarily for graphics and / or scientific (throughput) computing. Implementations of different processors can include: 1) a CPU that includes one or more general - purpose in - order cores intended for general - purpose computing and / or one or more general - purpose out - of - order cores intended for general - purpose computing; and 2) a coprocessor that includes one or more specialized cores intended primarily for graphics and / or scientific throughput. Such different processors result in different computer system architectures, which can include: 1) a coprocessor on a chip separate from the CPU; 2) a coprocessor in the same package as the CPU but on a separate die; 3) a coprocessor on the same die as the CPU (in which case, such a coprocessor is sometimes referred to as specialized logic or as a specialized core, such specialized logic being, for example, integrated graphics and / or scientific (throughput) logic); and 4) a system - on - a - chip that can include the described CPU (sometimes referred to as (one or more) application cores or (one or more) application processors), the coprocessor described above, and additional functionality on the same die. An example core architecture is then described, followed by example processors and computer architectures.
[0188] Example Core Architecture
[0189] In - Order and Out - of - Order Core Block Diagrams
[0190] Figure 9ais a block diagram of both an example in-order pipeline and an example register-renamed out-of-order issue / execution pipeline as illustrated. Figure 9b is a block diagram of both an example in-order architecture core to be included in a processor and an out-of-order issue / execution architecture core with example register renaming. Figure 9a – Figure 9b The solid boxes in [[ ]] show the in-order pipeline and in-order core, while the optionally added dashed boxes show the register-renamed out-of-order issue / execution pipeline and core. Given that the in-order aspects are a subset of the out-of-order aspects, the out-of-order aspects will be described.
[0191] In [[ ]] Figure 9a the processor pipeline 900 includes a fetch stage 902, a length decode stage 904, a decode stage 906, an allocation stage 908, a rename stage 910, a schedule (also known as dispatch or issue) stage 912, a register read / memory read stage 914, an execution stage 916, a write-back / memory write stage 918, an exception handling stage 922, and a commit stage 924.
[0192] Figure 9b shows a processor core 990 including a front-end unit 930 coupled to an execution engine unit 950, and both the execution engine unit and the front-end unit are coupled to a memory unit 970. The core 990 can be a reduced instruction set computing (RISC) core, a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, the core 990 can be a specialized core, such as, for example, a network or communication core, a compression engine, a coprocessor core, a general-purpose computing graphics processing unit (GPGPU) core, a graphics core, and so on.
[0193] The front-end unit 930 includes a branch prediction unit 932 that is coupled to an instruction cache unit 934, which is coupled to an instruction translation lookaside buffer (TLB) 936, which is coupled to an instruction fetch unit 938, which is coupled to a decode unit 940. The decode unit 940 (or decoder) can decode the instructions and generate as output one or more micro-operations, micro-code entry points, micro-instructions, other instructions, or other control signals that are decoded from, or otherwise reflect, or are derived from the original instructions. The decode unit 940 can be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, look-up tables, hardware implementations, programmable logic arrays (PLAs), micro-code read-only memories (ROMs), etc. In one embodiment, the core 990 includes a micro-code ROM or other medium (e.g., within the decode unit 940, or otherwise within the front-end unit 930) that stores micro-code for certain macro-instructions. The decode unit 940 is coupled to a rename / allocator unit 952 in the execution engine unit 950.
[0194] The execution engine unit 950 includes a rename / allocator unit 952 that is coupled to a retirement unit 954 and a collection 956 of one or more scheduler units. The scheduler unit 956 represents any number of different schedulers, including reservation stations, a central instruction window, and the like. The scheduler unit 956 is coupled to a physical register file unit 958. Each physical register file unit in the (multiple) physical register file units 958 represents one or more physical register files, where different physical register files store one or more different data types, such as scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points, status (e.g., an instruction pointer that is an address of a next instruction to be executed), and the like. In one embodiment, the (multiple) physical register file units 958 include a vector register unit, a write mask register unit, and a scalar register unit. These register units may provide architectural vector registers, vector mask registers, and general-purpose registers. The (multiple) physical register file units 958 are overlapped by the retirement unit 954 to illustrate various ways in which register renaming and out-of-order execution can be implemented (e.g., using the (multiple) reorder buffers and the (multiple) retirement register files; using the (multiple) future files, the (multiple) history buffers, and the (multiple) retirement register files; using register maps and register pools; and so on). The retirement unit 954 and the (multiple) physical register file units 958 are coupled to the (multiple) execution clusters 960. The (multiple) execution clusters 960 include a collection 962 of one or more execution units and a collection 964 of one or more memory access units. The execution units 962 can perform various operations (e.g., shift, add, subtract, multiply) and can perform on various data types (e.g., scalar floating point, packed integer, packed floating point, vector integer, vector floating point). Although some embodiments may include multiple execution units dedicated to a particular function or a set of functions, other embodiments may include only one execution unit or multiple execution units that all perform all functions. The (multiple) scheduler units 956, the (multiple) physical register file units 958, and the (multiple) execution clusters 960 are shown as potentially having multiple because certain embodiments create separate pipelines for certain types of data / operations (e.g., a scalar integer pipeline, a scalar floating point / packed integer / packed floating point / vector integer / vector floating point pipeline, and / or a memory access pipeline each having its own scheduler unit, the (multiple) physical register file units, and / or execution cluster - and in the case of a separate memory access pipeline, certain embodiments are implemented in which only the execution cluster of that pipeline has the (multiple) memory access units 964). It should also be understood that in the case of using separate pipelines, one or more of these pipelines may be out-of-order issue / execution, and the remaining pipelines may be in-order.
[0195] A set of memory access units 964 is coupled to a memory unit 970, which includes a data TLB unit 972, which is coupled to a data cache unit 974, which is coupled to a level 2 (L2) cache unit 976. In one embodiment, the memory access unit 964 may include a load unit, a store address unit, and a store data unit, each of which is coupled to the data TLB unit 972 in the memory unit 970. An instruction cache unit 934 is further coupled to the level 2 (L2) cache unit 976 in the memory unit 970. The L2 cache unit 976 is coupled to one or more other levels of cache and ultimately to main memory.
[0196] As an example, a register-renamed, out-of-order issue / execution core architecture may implement pipeline 900 as follows: 1) Instruction fetch 938 performs a fetch stage 902 and a length decoding stage 904; 2) A decode unit 940 performs a decode stage 906; 3) A rename / allocator unit 952 performs an allocation stage 908 and a rename stage 910; 4) One or more scheduler units 956 perform a schedule stage 912; 5) One or more physical register file units 958 and the memory unit 970 perform a register read / memory read stage 914; An execution cluster 960 performs an execution stage 916; 6) The memory unit 970 and one or more physical register file units 958 perform a write-back / memory write stage 918; 7) Each unit may be involved in an exception handling stage 922; and 8) A retirement unit 954 and one or more physical register file units 958 perform a commit stage 924.
[0197] The core 990 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions added with newer versions); the MIPS instruction set of MIPS Technologies, Inc., in Sunnyvale, California; the ARM instruction set of ARM Holdings, Inc., in Sunnyvale, California (with optional additional extensions such as NEON)), including the instructions described herein. In one embodiment, the core 990 includes logic for supporting SIMD (Single Instruction, Multiple Data) instruction set extensions (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using SIMD data.
[0198] It should be understood that the core may support multithreading (executing a set of two or more parallel operations or threads), and this multithreading may be accomplished in various ways, including time division multithreading, simultaneous multithreading (where a single physical core provides a logical core for each of the threads for which the physical core is simultaneously multithreading), or a combination thereof (e.g., time division fetching and decoding and thereafter such as Intel Simultaneous multithreading in hyperthreading technology).
[0199] Although register renaming is described in the context of out-of-order execution, it should be understood that register renaming can be used in an in-order architecture. Although the illustrated embodiments of the processor also include separate instruction and data cache units 934 / 974 and a shared L2 cache unit 976, alternative embodiments may have a single internal cache for both instructions and data, such as, for example, a level 1 (L1) internal cache or multiple levels of internal caches. In some embodiments, the system may include a combination of an internal cache and an external cache outside the core and / or processor. Alternatively, all caches may be outside the core and / or processor.
[0200] Example in-order core architecture
[0201] Figures 10a - 10b A block diagram illustrating a more specific example of an in-order core architecture, which would be one of several logic blocks (including other cores of the same type and / or different types) in a chip. Depending on the application, the logic block communicates with some fixed function logic, a memory IO interface, and other necessary IO logic via a high bandwidth interconnect network (e.g., a ring network).
[0202] Figure 10a A block diagram of a single processor core according to one or more embodiments, its connection to the on-die interconnect network 1002, and a local subset 1004 of its second level (L2) cache. In one embodiment, the instruction decoder 1000 supports the x86 instruction set with a compact data instruction set extension. The L1 cache 1006 allows low latency access to the cache memory into the scalar and vector units. Although in one embodiment (for simplicity of design), the scalar unit 1008 and the vector unit 1010 use separate register sets (scalar registers 1012 and vector registers 1014, respectively), and the data transferred between these registers is written to memory and then read back from the level 1 (L1) cache 1006, other embodiments may use different methods (such as using a single register set or including a communication path that allows data to be transferred between the two register banks without being written and read back).
[0203] The local subset 1004 of the L2 cache is part of a global L2 cache that is partitioned into multiple separate local subsets, one for each processor core. Each processor core has a direct access path to its own local subset 1004 of the L2 cache. Data read by a processor core is stored in its L2 cache subset 1004 and can be quickly accessed in parallel with other processor cores accessing their own local L2 cache subsets. Data written by a processor core is stored in that processor core's own L2 cache subset 1004 and flushed from other subsets if necessary. A ring network ensures data sharing consistency. The ring network is bidirectional to allow agents such as processor cores, L2 caches, and other logic blocks to communicate with each other within the chip. Each ring data path is 6012 bits wide in each direction.
[0204] Figure 10b is of an embodiment according to the present specification Figure 10a exploded view of a portion of the processor core in Figure 10b includes a portion of the L1 data cache 1006A that includes the L1 cache 1004, and more details regarding the vector unit 1010 and vector registers 1014. Specifically, the vector unit 1010 is a 16-wide vector processing unit (VPU) (see 16-wide ALU 1028) that executes one or more of integer, single-precision floating-point, and double-precision floating-point instructions. The VPU supports mixing of register inputs using a mixing unit 1020, numerical conversion using numerical conversion units 1022A-B, and copying of memory inputs using a copy unit 1024. A write mask register 1026 allows assertion of the resulting vector writes.
[0205] Figure 11 is a block diagram of a processor 1100 that can have more than one core, can have an integrated memory controller, and can have an integrated graphics device according to an embodiment of the present specification. The solid box in FIG. 6 illustrates a processor 1100 having a single core 1102A, a system agent 1110, and a group of one or more bus controller units 1116, while the optionally added dashed box shows an alternative processor 1100 having multiple cores 1102A-N, a set 1114 of one or more integrated memory controller units within the system agent unit 1110, and dedicated logic 1108.
[0206] Accordingly, different implementations of the processor 1100 can include: 1) a CPU, where the dedicated logic 1108 is integrated graphics and / or scientific (throughput) logic (which can include one or more cores), and the cores 1102A-N are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, a combination of both); 2) a coprocessor, where the cores 1102A-N are a large number of dedicated cores designed primarily for graphics and / or scientific throughput; and 3) a coprocessor, where the cores 1102A-N are a large number of general-purpose in-order cores. Thus, the processor 1100 can be a general-purpose processor, a coprocessor, or a special-purpose processor, such as, for example, a network or communication processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput integrated many-core (MIC) coprocessor (including 30 or more cores), an embedded processor, and so on. The processor can be implemented on one or more chips. The processor 1100 can be part of one or more substrates, and / or can be implemented on one or more substrates using any of a variety of process technologies, such as, for example, BiCMOS, CMOS, or NMOS.
[0207] The memory hierarchy includes one or more levels of cache within the cores, or a collection 1106 of one or more shared cache units, and external memory (not shown) coupled to a collection 1114 of integrated memory controller units. The collection 1106 of shared cache units can include one or more intermediate-level caches, last-level cache (LLC), and / or a combination of the above, with intermediate-level caches such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache. Although in one embodiment, the ring-based interconnect unit 1112 interconnects the integrated graphics logic 1108, the collection 1106 of shared cache units, and the system agent unit 1110 / (multiple) integrated memory controller units 1114, alternative embodiments can use any number of well-known techniques to interconnect such units. In one embodiment, coherence is maintained between one or more cache units 1106 and the cores 1102A-N.
[0208] In some embodiments, one or more of the cores 1102A-N are capable of implementing multithreading. The system agent 1110 includes those components that coordinate and operate the cores 1102A-N. The system agent unit 1110 can include, for example, a power control unit (PCU) and a display unit. The PCU can be the logic and components required to regulate the power states of the cores 1102A-N and the integrated graphics logic 1108, or can include such logic and components. The display unit is used to drive one or more externally connected displays.
[0209] The core 1102A-N can be homogeneous or heterogeneous in terms of the architecture instruction set; that is, two or more of the cores 1102A-N may be capable of executing the same instruction set, while other cores may be capable of executing only a subset of the instruction set or a different instruction set.
[0210] Example computer architecture
[0211] Figures 12 - 15 is a block diagram of an example computer architecture. Other system designs and configurations of laptop devices, desktop computers, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular telephones, portable media players, handheld devices, and various other electronic devices known in the art are also suitable. Generally, a wide variety of systems or electronic devices that can incorporate a processor and / or other execution logic as disclosed herein are generally suitable.
[0212] Now refer to Figure 12 shown is a block diagram of a system 1200 according to one embodiment. The system 1200 may include one or more processors 1210, 1215, which are coupled to a controller hub 1220. In one embodiment, the controller hub 1220 includes a graphics memory controller hub (GMCH) 1290 and an input / output hub (IOH) 1250 (which may be on separate chips); the GMCH 1290 includes a memory and a graphics controller, to which a memory 1240 and a coprocessor 1245 are coupled; the IOH 1250 couples input / output (IO) devices 1260 to the GMCH 1290. Alternatively, one or both of the memory and the graphics controller are integrated within the processor (as described herein), the memory 1240 and the coprocessor 1245 are directly coupled to the processor 1210, and the controller hub 1220 and the IOH 1250 are in a single chip.
[0213] The optional nature of the additional processor 1215 is shown in dashed lines in FIG. 7. Each of the processors 1210, 1215 may include one or more of the processing cores described herein and may be a certain version of the processor 1100.
[0214] The memory 1240 may be, for example, dynamic random access memory (DRAM), phase change memory (PCM), or a combination of the two. For at least one embodiment, the controller hub 1220 communicates with the (multiple) processors 1210, 1215 via a multi-branch bus such as a front-side bus (FSB), a point-to-point interface such as an ultra-path interconnect (UPI), or a similar connection 1295.
[0215] In one embodiment, the coprocessor 1245 is a special-purpose processor, such as, for example, a high-throughput MIC processor, a network or communications processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, and the like. In one embodiment, the controller hub 1220 may include an integrated graphics accelerator.
[0216] There may be various differences in a series of quality metrics, including architecture, microarchitecture, thermal, power consumption characteristics, etc., between the physical resources 1210, 1215.
[0217] In one embodiment, the processor 1210 executes instructions that control general types of data processing operations. Coprocessor instructions may be embedded within these instructions. The processor 1210 identifies these coprocessor instructions as being of a type that should be executed by the attached coprocessor 1245. Accordingly, the processor 1210 issues these coprocessor instructions (or control signals representing coprocessor instructions) on a coprocessor bus or other interconnect to the coprocessor 1245. The (multiple) coprocessor 1245 receives and executes the received coprocessor instructions.
[0218] Now referring Figure 13 , shown is a block diagram of a first more specific example system 1300. As Figure 13 shown, the multiprocessor system 1300 is a point-to-point interconnect system and includes a first processor 1370 and a second processor 1380 coupled via a point-to-point interconnect 1350. Each of the processors 1370 and 1380 may be a certain version of the processor 1100. In one embodiment, the processors 1370 and 1380 are the processor 1210 and the processor 1215, respectively, and the coprocessor 1338 is the coprocessor 1245. In another embodiment, the processors 1370 and 1380 are the processor 1210 and the coprocessor 1245, respectively.
[0219] The processors 1370 and 1380 are shown as including integrated memory controller (IMC) units 1372 and 1382, respectively. The processor 1370 also includes point-to-point (P-P) interfaces 1376 and 1378 as part of its bus controller unit; similarly, the second processor 1380 includes P-P interfaces 1386 and 1388. The processors 1370, 1380 may exchange information via the P-P interface 1350 using the point-to-point (P-P) interface circuits 1378, 1388. As Figure 13 shown, the IMCs 1372 and 1382 couple the processors to the respective memories, namely memories 1332 and 1334, which may be portions of the main memory locally attached to the respective processors.
[0220] Processors 1370, 1380 may each exchange information with chipset 1390 via respective P-P interfaces 1352, 1354 using point-to-point interface circuits 1376, 1394, 1386, 1398. Chipset 1390 may optionally exchange information with coprocessor 1338 via high performance interface 1339. In one embodiment, coprocessor 1338 is a special purpose processor such as, for example, a high throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, and the like.
[0221] A shared cache (not shown) may be included in either processor or external to both processors but connected to these processors via a P-P interconnect such that if a processor is placed in a low power mode, the local cache information of either or both processors may be stored in the shared cache.
[0222] Chipset 1390 may be coupled to first bus 1316 via interface 1396. In one embodiment, first bus 1316 may be a Peripheral Component Interconnect (PCI) bus or a bus such as a PCI Express bus or another third generation I / O interconnect bus, as a non-limiting example.
[0223] As Figure 13 shown, various I / O devices 1314 may be coupled to first bus 1316 along with bus bridge 1318 which couples first bus 1316 to second bus 1320. In one embodiment, one or more additional processors 1315 such as a coprocessor, a high throughput MIC processor, a GPGPU, an accelerator (such as, for example, a graphics accelerator or a Digital Signal Processing (DSP) unit), a Field Programmable Gate Array or any other processor are coupled to first bus 1316. In one embodiment, second bus 1320 may be a Low Pin Count (LPC) bus. In one embodiment, various devices may be coupled to second bus 1320 including, for example, a keyboard and / or mouse 1322, a communication device 1327, and a storage unit 1328 which may include, for example, a disk drive or other mass storage device that may include instructions or code and data 1330. Additionally, audio I / O 1324 may be coupled to second bus 1320. Note that other architectures are possible. For example, instead of Figure 13 the point-to-point architecture, the system may implement a multi-branch bus or other such architecture.
[0224] Now referring Figure 14 to, shown is a block diagram of a second more specific example system 1400. Figure 13 and Figure 14 are denoted with the same reference numerals and have been from Figure 14is omitted in Figure 13 certain aspects of Figure 14 in order to avoid obscuring
[0225] Figure 14 shows that processors 1370, 1380 may respectively include integrated memory and IO control logic (“CL”) 1372 and 1382. Thus, CL 1372, 1382 include integrated memory controller units and include IO control logic. Figure 14 shows that not only memories 1332, 1334 are coupled to CL 1372, 1382, but also IO devices 1414 are coupled to control logics 1372, 1382. Conventional IO device 1415 is coupled to chipset 1390.
[0226] Now refer to Figure 15 , shown is a block diagram of SoC 1500 according to an embodiment. Similar elements in FIG. 10 have the same reference numerals. Additionally, the dashed boxes are optional features on more advanced SoCs. In FIG. 10, the (multiple) interconnect units 1502 are coupled to: an application processor 1510, which includes a set of one or more cores 1102A-N and the (multiple) shared cache units 1106; a system agent unit 1110; the (multiple) bus controller units 1116; the (multiple) integrated memory controller units 1114; a set of one or more coprocessors 1520, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random access memory (SRAM) unit 1530; a direct memory access (DMA) unit 1532; and a display unit 1540 for coupling to one or more external displays. In one embodiment, the (multiple) coprocessors 1520 include dedicated processors, such as, for example, a network or communication processor, a compression engine, a GPGPU, a high throughput MIC processor, an embedded processor, and so on.
[0227] Embodiments of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of such implementations. Some embodiments may be implemented as a computer program or program code executed on a programmable system that includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0228] The program code (such as, Figure 8The code 1330 shown in is applied to the input instructions to perform the various functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0229] The program code can be implemented in a high-level procedural programming language or an object-oriented programming language in order to communicate with the processing system. If desired, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described herein are not limited to any specific programming language scope. In any case, the language can be a compiled language or an interpreted language.
[0230] One or more aspects of at least one embodiment can be implemented by representative instructions stored on a machine-readable medium that represent various logic in a processor, which when read by the machine cause the machine to fabricate logic for performing the techniques described herein. Such representations, referred to as “IP cores,” can be stored on a tangible machine-readable medium and supplied to various customers or manufacturing facilities to be loaded into the manufacturing machines that actually make the logic or processor.
[0231] Such machine-readable storage media can include, but are not limited to, non-transitory, tangible arrangements of articles manufactured or formed by a machine or device that include storage media such as a hard disk; any other type of disk including floppy disks, optical disks, compact disk read only memory (CD-ROM), rewritable compact disk (CD-RW), and magneto-optical disks; semiconductor devices such as read only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM) and static random access memory (SRAM), erasable programmable read only memory (EPROM), flash memory, electrically erasable programmable read only memory (EEPROM); phase change memory (PCM); magnetic or optical cards; or any other type of medium suitable for storing electronic instructions.
[0232] Accordingly, some embodiments also include a non-transitory tangible machine-readable medium that contains instructions or contains design data such as a hardware description language (HDL) that defines the structures, circuits, devices, processors, and / or system features described herein. Such embodiments can also be referred to as program products.
[0233] Emulation (including binary translation, code morphing, etc.)
[0234] In some cases, an instruction converter may be used to convert instructions from a source instruction set to a target instruction set. For example, the instruction converter may transform (e.g., using static binary translation or dynamic binary translation including dynamic compilation), morph, emulate, or otherwise convert instructions into one or more other instructions to be processed by a core. The instruction converter may be implemented in software, hardware, firmware, or a combination thereof. The instruction converter may be on the processor, off the processor, or partly on the processor and partly off the processor.
[0235] Figure 16 is a block diagram contrasting the use of a software instruction converter to convert binary instructions in a source instruction set to binary instructions in a target instruction set. In the illustrated embodiment, the instruction converter is a software instruction converter, but alternatively, the instruction converter may be implemented in software, firmware, hardware, or various combinations thereof. Figure 16 shows that an x86 compiler 1604 may be used to compile a program utilizing a high-level language 1602 to generate x86 binary code 1606 that can be natively executed by a processor 1616 having at least one x86 instruction set core. The processor 1616 having at least one x86 instruction set core represents any processor that can perform substantially the same functions as an Intel processor having at least one x86 instruction set core by compatibly executing or otherwise processing: 1) an essential portion of the instruction set of the Intel instruction set core, or 2) a version of the object code of an application or other software targeted to run on an Intel processor having at least one x86 instruction set core, so as to obtain substantially the same results as an Intel processor having at least one x86 instruction set core. The x86 compiler 1604 represents a compiler operable to generate x86 binary code 1606 (e.g., object code) that can be executed on a processor 1616 having at least one x86 instruction set core with or without additional linking processing. Similarly, Figure 16It is shown that an alternative instruction set compiler 1608 can be used to compile a program of a high-level language 1602 to generate alternative instruction set binary code 1610 that can be natively executed by a processor 1614 that does not have at least one x86 instruction set core (e.g., a processor having a core that executes the MIPS instruction set of MIPS Technologies, Inc. of Sunnyvale, California and / or the ARM instruction set of ARM Holdings plc of Sunnyvale, California). An instruction converter 1612 is used to convert x86 binary code 1606 into code that can be natively executed by a processor 1614 that does not have an x86 instruction set core. The converted code is not likely to be the same as the alternative instruction set binary code 1610 because it is difficult to manufacture an instruction converter that can do so; however, the converted code will perform general operations and is composed of instructions from an alternative instruction set. Thus, the instruction converter 1612 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device that does not have an x86 instruction set processor or core to execute x86 binary code 1606 through emulation, simulation, or any other process.
[0236] The foregoing outlines the features of one or more embodiments of the subject matter disclosed herein. These embodiments are provided so that a person of ordinary skill in the art (PHOSITA) can better understand various aspects of the present disclosure. Certain well-known terms and underlying technologies and / or standards may be referenced without detailed description. It is expected that the PHOSITA will have background knowledge or information in or access to these technologies and standards sufficient to practice the teachings of this specification.
[0237] The PHOSITA will understand that they can readily use the present disclosure as a basis for designing or modifying other processes, structures, or variations to perform the same purposes of the various embodiments described herein and / or achieve the same advantages of the various embodiments described herein. The PHOSITA will also recognize that such equivalent constructs do not depart from the spirit and scope of the present disclosure, and that they can make various changes, substitutions, and alterations herein without departing from the spirit and scope of the present disclosure.
[0238] In the foregoing description, the description of some or all aspects of some embodiments is more detailed than is strictly required to practice the appended claims. These details are provided only by way of non-limiting examples for purposes of providing context and illustration of the disclosed embodiments. Such details should not be construed as being required, nor should they be “read into” the claims as limitations. The phrase may refer to “an embodiment” or “embodiments”. These phrases, as well as any other reference to an embodiment, should be understood broadly as referring to any combination of one or more embodiments. In addition, several features disclosed in a particular “embodiment” may also be distributed across multiple embodiments. For example, if features 1 and 2 are disclosed in an “embodiment”, embodiment A may have feature 1 but lack feature 2, while embodiment B may have feature 2 but lack feature 1.
[0239] This specification may provide descriptions in block diagram form, where certain features are disclosed in separate blocks. These are to be understood broadly as disclosing how the various features interact operatively, but are not intended to imply that those features must be embodied in separate hardware or software. Further, where more than one feature is disclosed in a single block, those features need not necessarily be embodied in the same hardware and / or software. For example, in some cases, computer “memory” may be distributed or mapped among a multi-level cache or local memory, main memory, battery-backed volatile memory, and various forms of persistent memory such as hard disks, storage servers, optical disks, tape drives, or similar persistent memory. In some embodiments, some components may be omitted or combined. In general, the arrangements depicted in the drawings may be more logical in their representation, while the physical architecture may include various arrangements, combinations, and / or mixtures of these elements. An infinite number of possible design configurations may be used to achieve the operational objectives outlined herein. Accordingly, the associated infrastructure has an infinite number of alternative arrangements, design choices, device possibilities, hardware configurations, software implementations, and equipment options.
[0240] This document may refer to a computer-readable medium, which may be a tangible and non-transitory computer-readable medium. As used throughout this specification and the entire claims, "computer-readable medium" should be understood to include one or more computer-readable media of the same or different types. By way of non-limiting example, computer-readable media may include optical disk drives (e.g., CD / DVD / Blue-ray), hard disk drives, solid state drives, flash memory, or other non-volatile media. Computer-readable media may also include media such as read-only memory (ROM); FPGAs or ASICs configured to execute desired instructions, with the stored instructions for programming the FPGA or ASIC to execute the desired instructions; intellectual property (IP) blocks integratable into other circuits in hardware; or instructions directly encoded in the hardware or microcode of a processor (such as a microprocessor, digital signal processor (DSP), microcontroller); or instructions encoded in any other suitable component, device, element, or object as appropriate and based on specific needs. The non-transitory storage media herein are expressly intended to include any non-transitory dedicated or programmable hardware configured to provide the disclosed operations or cause a processor to execute the disclosed operations.
[0241] Throughout the specification and claims, various elements may be "communicatively", "electrically", "mechanically", or otherwise "coupled" to each other. Such couplings may be direct point-to-point couplings, or may include intermediate devices. For example, two devices may be communicatively coupled to each other via a controller that facilitates communication. Devices may be electrically coupled to each other via intermediate devices such as signal boosters, voltage dividers, or buffers. Mechanically coupled devices may be indirectly mechanically coupled.
[0242] Any "module" or "engine" disclosed herein may refer to or include software, software stacks, hardware, firmware, and / or a combination of software, circuitry configured to perform the functions of the engine or module, or any of the computer-readable media disclosed above. Where appropriate, such modules or engines may be provided on or in conjunction with a hardware platform, which may include hardware computing resources such as processors, memory, storage, interconnects, networks and network interfaces, accelerators, or other suitable hardware. Such hardware platforms may be provided as a single monolithic device (e.g., in a PC form factor), or have some or part of the functions distributed (e.g., in a "composite node" in a high-end data center, where computing, memory, storage, and other resources may be dynamically allocated and do not need to be local to each other).
[0243] This document may disclose flowcharts, signal flowcharts, or other diagrams showing operations performed in a particular order. Unless explicitly stated otherwise, or unless required in a particular context, the order should be understood as being only a non-limiting example. Additionally, in cases where one operation is shown as following another operation, other intervening operations may occur that may or may not be relevant. Some operations may also be performed simultaneously or in parallel. In cases where an operation is referred to as being "based on" or "in accordance with" another item or operation, this should be understood to imply that the operation is at least partially based on or at least partially in accordance with the other item or operation. This should not be construed to imply that the operation is only based on or exclusively based on the item or operation, or only in accordance with or exclusively in accordance with the item or operation.
[0244] All or part of any hardware element disclosed herein can be readily provided in a system-on-a-chip (SoC) (including a central processing unit (CPU) package). An SoC represents an integrated circuit (IC) that integrates the components of a computer or other electronic system onto a single chip. Thus, for example, a client device or a server device can be provided, in whole or in part, in an SoC. An SoC can include digital, analog, mixed-signal, and radio frequency functionality, all of which can be provided on a single chip substrate. Other embodiments can include a multi-chip module (MCM), where multiple chips are located within a single electronic package and configured to interact closely with each other through this electronic package.
[0245] In a general sense, any appropriately configured circuit or processor can execute any type of instruction associated with data to implement the operations detailed herein. Any processor disclosed herein can transform an element or article (e.g., data) from one state or thing to another state or thing. Additionally, based on specific needs and implementations, the information tracked, sent, received, or stored in a processor can be provided in any database, register, table, cache, queue, control list, or storage structure, all of which can be referenced within any suitable time frame. Any of the memory or storage elements disclosed herein should be construed to be appropriately encompassed within the broad terms "memory" and "storage".
[0246] The computer program logic for implementing all or part of the functions described herein is embodied in various forms, including but not limited to, source code form, computer-executable form, machine instructions or microcode, programmable hardware, and various intermediate forms (e.g., forms generated by assemblers, compilers, linkers, or locators). In an example, the source code includes a series of computer program instructions implemented in various programming languages or in a hardware description language, various programming languages such as, object code, assembly language, or high-level languages (such as, OpenCL, FORTRAN, C, C++, JAVA, or HTML) for use with various operating systems or operating environments, and hardware description languages such as, Spice, Verilog, and VHDL. The source code can define and use various data structures and communication messages. The source code can be in computer-executable form (e.g., via an interpreter), or the source code can be converted (e.g., via a converter, assembler, or compiler) into computer-executable form, or into an intermediate form (such as, bytecode). In a suitable case, any of the foregoing can be used to build or describe a suitable discrete circuit or integrated circuit, whether sequential, combinational, state machine, or other form.
[0247] In one example embodiment, any number of the circuits in the figures can be implemented on a board of an associated electronic device. The board can be a general-purpose circuit board that can secure various components of the internal electronic system of the electronic device and can further provide connectors for other peripheral devices. Any suitable processor and memory can be appropriately coupled to the board based on specific configuration requirements, processing needs, and computing designs. Note that for the numerous examples provided herein, interactions can be described with two, three, four, or more electrical components. However, this is done only for clarity and example purposes. It should be understood that the system can also be combined or reconfigured in any suitable way. Together with similar design alternatives, any one of the components, modules, and elements shown in the figures can be combined in various possible configurations, all of which are within the broad scope of this specification.
[0248] Numerous other changes, substitutions, variations, alterations, and modifications will be apparent to those skilled in the art, and this disclosure is intended to cover all such changes, substitutions, variations, alterations, and modifications as falling within the scope of the appended claims.
[0249] To assist the United States Patent and Trademark Office (USPTO) and, further, any reader of any patent issued on this application in interpreting the appended claims, the Applicant wishes to note that the Applicant: (a) does not wish any of the appended claims to invoke that paragraph in its application date because of paragraph (6) of 35 U.S.C. § 112 (prior to AIA) or paragraph (f) of the same section (after AIA), unless the words "means for" or "step for" are specifically used in a particular claim; and (b) does not wish any statement in this application file to limit the disclosure in any way not otherwise expressly reflected in the appended claims.
[0250] Example Implementations
[0251] An example of a microprocessor is disclosed that includes: a processing core; and a Total Memory Encryption (TME) engine that is configured to provide TME for a first Trust Domain (TD) and is further configured to: allocate a physical memory block to the first TD and assign a first cryptographic key to the first TD; map a Host Physical Address (HPA) space to a Guest Physical Address (GPA) space of the TD within an Extended Page Table (EPT); create a Memory Ownership Table (MOT) entry for a memory page within the physical memory block, where the MOT table includes a GPA reverse mapping; encrypt the MOT entry using the first cryptographic key; and append MOT entry verification data to the MOT entry, where the MOT entry verification data enables detection of an attack on the MOT entry.
[0252] A further example of a microprocessor is disclosed, where the processor is configured to supply the MOT within a memory range controlled by a physical memory range register in response to one or more instructions.
[0253] A further example of a microprocessor is disclosed, where the TME engine is a multi-key TME engine, where the first cryptographic key provides a first key domain, and where the TME engine is further configured to allocate a second TD with a second key domain.
[0254] A further example of a microprocessor is disclosed, where the MOT further includes a TD Control Structure (TDCS) pointer field.
[0255] A further example of a microprocessor is disclosed, where the entry verification data includes a version number field.
[0256] A further example of a microprocessor is disclosed, where the entry verification data includes an integrity field.
[0257] A further example of a microprocessor is disclosed, where the integrity field includes a cryptographic hash of the MOT entry signed by a first encryption key.
[0258] Examples of a microprocessor are further disclosed, wherein the MOT entry is a 128-bit hash.
[0259] Examples of a microprocessor are further disclosed, wherein the MOT entry is divided into 128-bit rows.
[0260] Examples of a microprocessor are further disclosed, wherein the TME is configured to encrypt the memory at the cache line granularity.
[0261] Examples of a microprocessor are further disclosed, wherein the MOT is configured to divide cache operations into 128-bit aligned blocks.
[0262] Examples of a microprocessor are further disclosed, wherein the processor further includes a page miss handler (PMH), the PMH being configured to walk through a memory page upon a page miss, to determine that an integrity check based on entry verification data has failed, and to invalidate the memory page.
[0263] Examples of a microprocessor are further disclosed, wherein the PMH is further used to send a TD exit signal for a first TD.
[0264] Examples of a computing device including a memory and a microprocessor are also disclosed.
[0265] Examples of a computing device are further disclosed, the computing device further including a virtual machine monitor (VMM), wherein the TME engine is configured to isolate a first TD from the VMM.
[0266] Examples of one or more tangible non-transitory media are also disclosed, the media having instructions stored thereon for providing total memory encryption (TME) for a trust domain (TD), including instructions for: allocating a physical memory block to a first TD and allocating a first cryptographic key to the first TD; mapping a host physical address (HPA) space to a guest physical address (GPA) space of the TD within an extended page table (EPT); creating a memory ownership table (MOT) entry for a memory page within the physical memory block, wherein the MOT table includes a GPA reverse mapping; encrypting the MOT entry using the first cryptographic key; and appending MOT entry verification data to the MOT entry, wherein the MOT entry verification data enables detection of an attack on the MOT entry.
[0267] Examples of one or more tangible non-transitory media are further disclosed, wherein the first cryptographic key provides a first key domain, and wherein the instructions are further for allocating a second TD having a second key domain.
[0268] Examples of one or more tangible non-transitory media are further disclosed, wherein the MOT further includes a TD control structure (TDCS) pointer field.
[0269] Examples of one or more tangible non-transitory media are further disclosed, wherein the entry verification data includes an integrity field.
[0270] Examples of one or more tangible non-transitory media are further disclosed, wherein the entry verification data includes a version number field.
[0271] Examples of one or more tangible non-transitory media are further disclosed, wherein the integrity field includes a cryptographic hash of the MOT entry signed by a first cryptographic key.
[0272] Examples of one or more tangible non-transitory media are further disclosed, wherein the MOT entry is a 128-bit hash.
[0273] Examples of one or more tangible non-transitory media are further disclosed, wherein the MOT entry is divided into 128-bit rows.
[0274] Examples of one or more tangible non-transitory media are further disclosed, further including encrypting the memory at cache line granularity.
[0275] Examples of one or more tangible non-transitory media are further disclosed, further including dividing cache operations into 128-bit aligned blocks.
[0276] Examples of one or more tangible non-transitory media are further disclosed, further including: walking through a memory page on a page miss, determining that an integrity check based on the entry verification data has failed, and invalidating the memory page.
[0277] Examples of one or more tangible non-transitory media are further disclosed, further including sending a TD exit signal for a first TD.
[0278] A computer-implemented method for providing total memory encryption (TME) for a trust domain (TD) is also disclosed, the method including: allocating a physical memory block to a first TD and allocating a first cryptographic key to the first TD; mapping a host physical address (HPA) space to a client physical address (GPA) space of the TD within an extended page table (EPT); creating a memory ownership table (MOT) entry for a memory page within the physical memory block, wherein the MOT table includes a GPA reverse mapping; encrypting the MOT entry using the first cryptographic key; and appending MOT entry verification data to the MOT entry, wherein the MOT entry verification data enables detection of an attack on the MOT entry.
[0279] A method is further disclosed, wherein a first cryptographic key provides a first key domain, further comprising allocating a second TD having a second key domain.
[0280] A method is further disclosed, wherein the MOT further comprises a TD control structure (TDCS) pointer field.
[0281] A method is further disclosed, wherein the entry verification data includes an integrity field.
[0282] A method is further disclosed, wherein the entry verification data includes a version number field.
[0283] A method is further disclosed, wherein the integrity field includes a cryptographic hash of the MOT entry signed by a first encryption key.
[0284] A method is further disclosed, wherein the MOT entry is a 128-bit hash.
[0285] A method is further disclosed, wherein the MOT entry is divided into 128-bit rows.
[0286] A method is further disclosed, further comprising encrypting the memory at cache line granularity.
[0287] A method is further disclosed, further comprising dividing cache operations into 128-bit aligned blocks.
[0288] A method is further disclosed, further comprising: walking through a memory page upon a page miss, determining that an integrity check based on the entry verification data has failed, and invalidating the memory page.
[0289] A method is further disclosed, further comprising sending a signal for TD exit for a first TD.
[0290] A device is further disclosed, comprising means for performing the method of one or more examples of this specification.
[0291] A device is further disclosed, wherein the means for performing the method comprises a processor and a memory.
[0292] A device is further disclosed, wherein the memory comprises machine-readable instructions that, when executed, cause the device to perform the method of one or more examples of this specification.
[0293] A device is further disclosed, wherein the device is a computing system.
[0294] Further disclosed is at least one computer-readable medium comprising instructions that, when executed, implement the method as claimed in one or more examples of this specification or implement the apparatus as claimed in one or more examples of this specification.
[0295] A computing device includes: a hardware platform; and a Total Memory Encryption (TME) device for providing TME for a first Trust Domain (TD), and further for: allocating a memory block to the first TD and allocating a first cryptographic key to the first TD; mapping a first physical address space to a second physical address space of the TD within an Extended Page Table (EPT); creating a Memory Ownership Table (MOT) entry for a memory page within the memory block, wherein the MOT table includes a reverse mapping from the second physical address space to the first physical address space; encrypting the MOT entry using the first cryptographic key; and appending MOT entry verification data to the MOT entry, wherein the MOT entry verification data enables detection of an attack on the MOT entry.
[0296] Examples are further described, wherein the hardware platform includes a processing device for supplying the MOT within a memory range controlled by a physical memory range register in response to one or more instructions.
[0297] Examples are further described, wherein the TME device includes a multi-key TME engine, wherein the first cryptographic key provides a first key domain, and wherein the TME device is further for allocating a second TD having a second key domain.
[0298] Examples are further described, wherein the MOT further includes a TD Control Structure (TDCS) pointer field.
[0299] Examples are further described, wherein the entry verification data includes an integrity field.
[0300] Examples are further described, wherein the entry verification data includes a version number field.
[0301] Examples are further described, wherein the integrity field includes a cryptographic hash of the MOT entry signed by a first encryption key.
[0302] Examples are further described, wherein the MOT entry is a 128-bit hash.
[0303] Examples are further described, wherein the MOT entry is divided into 128-bit rows.
[0304] Examples are further described, wherein the TME is configured to encrypt the memory at a cache line granularity.
[0305] Examples are further described, where the MOT is configured to divide cache operations into 128-bit aligned chunks.
[0306] Examples are further described, where the processor device further includes a page miss handler (PMH) that is configured to walk through a memory page upon a page miss, to determine that an integrity check based on entry verification data has failed, and to invalidate the memory page.
[0307] Examples are further described, where the PMH is further used to send a signal for TD exit for a first TD.
[0308] Examples of a computing device including a memory and a hardware platform of any of the foregoing examples are further described.
[0309] Examples are further described, further including a virtual machine monitor (VMM), where the TME device is configured to isolate a first TD from the VMM.
Claims
1. A microprocessor, comprising: A processing core; And A total memory encryption TME engine, the TME engine being configured to provide TME for a first trust domain TD and further configured to: Allocate a physical memory block to the first trust domain TD and allocate a first cryptographic key to the first trust domain TD; Map a host physical address HPA space to a guest physical address GPA space of the first trust domain TD within an extended page table EPT; Create a memory ownership table MOT entry for a memory page within the physical memory block, wherein the MOT table includes a GPA reverse mapping; Encrypt the MOT entry using the first cryptographic key; and Append MOT entry verification data to the MOT entry, wherein the MOT entry verification data enables detection of an attack on the MOT entry.
2. The microprocessor according to claim 1, characterized in that, The processor is configured to supply the MOT within a memory range controlled by a physical memory range register in response to one or more instructions.
3. The microprocessor according to claim 1, characterized in that, The TME engine is a multi-key TME engine, wherein the first cryptographic key provides a first key domain, and wherein the TME engine is further configured to allocate a second TD having a second key domain.
4. The microprocessor according to claim 1, wherein The MOT further includes a TD control structure TDCS pointer field.
5. The microprocessor according to claim 1, characterized in that, The entry verification data includes a version number field.
6. The microprocessor according to claim 1, wherein The entry verification data includes an integrity field.
7. The microprocessor according to claim 6, characterized in that, The integrity field includes a cryptographic hash of the MOT entry signed by an encryption key.
8. The microprocessor according to claim 7, wherein The MOT entry is a 128-bit hash.
9. The microprocessor according to claim 7, wherein, The MOT entry is divided into 128-bit rows.
10. The microprocessor according to claim 1, wherein, The TME is configured to encrypt memory at cache line granularity.
11. The microprocessor according to claim 10, wherein, The MOT is configured to divide cache operations into 128-bit aligned blocks.
12. The microprocessor according to any one of claims 1-11, characterized in that, The processor further includes a page miss handler PMH, the PMH being configured to walk through a memory page on a page miss, to determine that an integrity check based on the entry verification data has failed, and to invalidate the memory page.
13. The microprocessor according to claim 12, wherein, The PMH is further configured to signal a TD exit for the first trust domain TD.
14. A computing device comprising a memory and a microprocessor according to any one of claims 1-13.
15. The computing device according to claim 14, further comprising a virtual machine monitor VMM, wherein, The TME engine is configured to isolate the first trust domain TD from the VMM.
16. One or more tangible non-transitory media having instructions stored thereon for providing total memory encryption TME for a first trust domain TD, the instructions including instructions for: Allocate a physical memory block to the first trust domain TD and allocate a first cryptographic key to the first trust domain TD; Map a host physical address HPA space to a guest physical address GPA space of the first trust domain TD within an extended page table EPT; Create a Memory Ownership Table (MOT) entry for the memory page within the physical memory block, wherein, The MOT table includes a GPA reverse mapping; Encrypt the MOT entry using the first cryptographic key; and Append MOT entry verification data to the MOT entry, wherein the MOT entry verification data enables detection of attacks on the MOT entry.
17. One or more tangible non-transitory media as described in claim 16, wherein The first cryptographic key provides a first key domain, and wherein the instructions are further for allocating a second TD having a second key domain.
18. One or more tangible non-transitory media as recited in claim 16, wherein, The MOT further includes a TD control structure TDCS pointer field.
19. One or more tangible non-transitory media as recited in claim 16, wherein The entry verification data includes an integrity field.
20. One or more tangible non-transitory media as claimed in claim 16, characterized in that, The entry verification data includes a version number field.
21. One or more tangible non-transitory media as described in claim 19, characterized in that, The integrity field includes a cryptographic hash of the MOT entry signed by an encryption key.
22. The one or more tangible non-transitory media according to claim 21, wherein, The MOT entry is a 128-bit hash.
23. One or more tangible non-transitory media as described in claim 21, characterized in that, The MOT entry is divided into 128-bit rows.
24. One or more tangible non-transitory media as described in claim 16, characterized in that, The instructions are for encrypting the memory at cache line granularity.
25. One or more tangible non-transitory media as claimed in claim 16, characterized in that, The instructions are for dividing cache operations into 128-bit aligned chunks.
Citation Information
Patent Citations
Remote access to hosted virtual machines by enterprise users
CN102420846A
Digital rights management method
CN102882677A