Processors, methods, systems, and instructions to save and restore protected execution environment context
By introducing selective saving circuits, lazy recovery circuits and lazy access control inspection circuits into the processor, the processor can optimize when saving and restoring the context, solving the performance and power consumption problems in the prior art and achieving more efficient context management.
Patent Information
- Application Number
- CN202411890035.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-01
AI Technical Summary
Existing processors have performance and power consumption issues when saving and restoring contexts, especially when switching threads frequently or executing protected execution environments.
By introducing selective saving circuits, lazy recovery circuits, and lazy access control inspection circuits, the processor can selectively save and restore the context, reducing unnecessary power and performance consumption. These circuits determine which contexts need to be saved or restored through context metadata, avoiding frequent storage and recovery of unmodified contexts.
This method effectively reduces the power and performance consumption of the processor when saving and restoring the context, especially when frequently switching threads or executing a protected execution environment, improving the overall performance and efficiency of the system.
Smart Images

Figure CN120234100A_ABST
Abstract
Description
Technical Field
[0001] Embodiments generally relate to processors. Specifically, the embodiments described herein generally relate to saving and restoring the context of a processor. Background Art
[0002] During use, a processor generates and maintains an execution state or context while executing threads. The context may include data or values stored in the architectural registers of the processor. When switching between threads, it is generally necessary to save and restore such context. For example, when switching out of an outgoing thread, the context of the outgoing thread can be saved from the architectural registers of the processor to system memory. Similarly, when switching into an incoming thread, the context of the incoming thread can be restored from system memory into the architectural registers of the processor. Brief Description of the Drawings
[0003] Examples in accordance with the present disclosure will be described with reference to the accompanying drawings, in which:
[0004] Figure 1 is a block diagram of a computer system in which embodiments of the invention may be implemented.
[0005] Figure 2 is a block diagram of a first exemplary embodiment in which context element metadata is stored together with corresponding context elements.
[0006] Figure 3 is a block diagram of a second exemplary embodiment in which registers are used to store context element metadata from corresponding context elements.
[0007] Figure 4 is a block diagram of a first exemplary embodiment in which context element metadata is defined at register granularity.
[0008] Figure 5 is a block diagram of a second exemplary embodiment in which context element metadata is defined at register set granularity.
[0009] Figure 6 is a flowchart of an embodiment of a method of entering a protected execution environment.
[0010] Figure 7 is a block diagram of an embodiment of a device that operates an embodiment for executing a control primitive to exit a protected execution environment or asynchronously exit a protected execution environment due to an exception condition.
[0011] Figure 8 is a block diagram of an embodiment of a processor that operates an embodiment for executing an instruction to exit a protected execution environment.
[0012] Figure 9A block diagram of an example apparatus that operates to execute a command for exiting a protected execution environment.
[0013] Figure 10 A flowchart of an example method for exiting a protected execution environment.
[0014] Figure 11 A block diagram of an example processor that operates to execute a context access instruction.
[0015] Figure 12 A flowchart of an example method for lazily restoring context elements and committing an operation to read a context element.
[0016] Figure 13 A flowchart of an example method for committing an operation to write a context element.
[0017] Figure 14 A flowchart of an example method for lazily restoring context elements and committing an operation to read and then write a context element.
[0018] Figure 15 An illustrated example computing system.
[0019] Figure 16 A block diagram illustrating an example processor and / or system-on-a-chip (SoC) that may have one or more cores and an integrated memory controller.
[0020] FIG. 17(A) is a block diagram illustrating an example in-order pipeline and example register renaming, out-of-order issue / execution pipeline both according to an example.
[0021] FIG. 17(B) is a block diagram illustrating an example in-order architecture core and example register renaming, out-of-order issue / execution architecture core both to be included in a processor according to an example.
[0022] Figure 18 An illustration of an example of (one or more) execution unit circuitry.
[0023] Figure 19 A block diagram of a register architecture according to some examples.
[0024] Figure 20 An illustration of an example instruction format.
[0025] Figure 21 An illustration of an example addressing information field.
[0026] Figure 22 An illustration of an example of a first prefix.
[0027] Figure 23(A)-Figure 23(D) Illustration of how to use Figure 22 Examples of the R, X, and B fields of the first prefix in
[0028] Figure 24(A)-Figure 24(B) Illustration of an example of the second prefix
[0029] Figure 25 Illustration of an example of the third prefix
[0030] Figure 26 is a block diagram illustrating, according to an example, the conversion of binary instructions in a source instruction set architecture into binary instructions in a target instruction set architecture using a software instruction converter. Detailed implementation manners
[0031] The present disclosure relates to methods, devices, systems, instructions, and non-transitory computer-readable storage media for saving and restoring contexts of protected execution environments. In the following description, numerous specific details are set forth (e.g., specific methods, operations, instructions, processor configurations, microarchitecture details, etc.). However, embodiments may be practiced without these specific details. In other instances, well-known circuits, structures, and techniques have not been shown in detail to avoid obscuring the understanding of this specification.
[0032] Figure 1 is a block diagram of a computer system 100 in which embodiments of the present invention may be implemented. In various embodiments, the computer system may represent a server, a workstation, a desktop computer, a laptop computer, a notebook computer, a tablet computer, a smart phone, a set-top box, a network device (e.g., a router, a switch, etc.), or various other types of computer systems known in the art.
[0033] The computer system includes a processor 101 and a system memory 116. The processor and the system memory are coupled to each other (e.g., via one or more interconnects, memory controllers, chipset components, etc.). The computer system is shown and described to better illustrate certain concepts, but it is to be appreciated that other embodiments relate only to the processor and not to the system memory.
[0034] The system memory 116 may include one or more types of memory. Suitable types of memory include, but are not limited to: random access memory (RAM), such as dynamic random access memory (DRAM); non-volatile memory, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), other types of read-only memory (ROM), and flash memory; persistent memory; other types of memory known in the art; and combinations of the foregoing.
[0035] In some embodiments, the processor may be a general-purpose processor (e.g., the type of general-purpose microprocessor or central processing unit (CPU) used in desktop computers, laptop computers, servers, and other computer systems). In other embodiments, the processor may be a special-purpose processor. Examples of suitable special-purpose processors include, but are not limited to, co-processors, graphics processors (e.g., general-purpose GPUs), security processors, machine learning processors, artificial intelligence processors, network processors, and controllers (e.g., microcontrollers).
[0036] When in use, the protected software 117 and the untrusted privileged system software 118 may be stored in the system memory. The protected software may broadly represent software to be protected from the influence of untrusted software (e.g., untrusted privileged system software). The untrusted privileged system software may broadly represent privileged system software that is not trusted by the protected software. As an example, the untrusted privileged system software may be untrusted because there is a possibility that the untrusted privileged system software becomes corrupted and then behaves improperly (such as, for example, stealing secret information (e.g., passwords, cryptographic keys, confidential data, etc.), falsely authenticating itself as the protected software, or otherwise mimicking the protected software, etc.). In some embodiments, the untrusted privileged system software may be outside the trusted computing base (TCB) of the protected software.
[0037] Typically, untrusted privileged system software may include at least one operating system 119 (e.g., a standard operating system (OS), a real-time operating system, a highly streamlined operating environment with limited conventional OS functionality). In some cases, the untrusted privileged system software may also include at least one virtual-machine monitor (VMM) 120. A VMM is sometimes referred to as a hypervisor. The VMM may present or expose an abstraction of one or more virtual machines (VMs) to other software (e.g., software referred to as "guest" software). The VMM may emulate or otherwise provide a bare-metal interface to the VMs. The VMM may assist in managing the VMs (e.g., managing resource allocation for the VMs). In some embodiments, a VM may include a guest OS and protected software. In some cases, the untrusted privileged system software may optionally include multiple VMMs (e.g., nested VMMs). A VMM may be used with some types but not all possible types of protected software.
[0038] The processor may support or provide a protected execution environment 103. The protected execution environment may help to allow protected software 117 to execute in a manner that is protected from untrusted software (e.g., untrusted privileged system software 118). In embodiments, the protected execution environment may be a trusted execution environment (TEE), an isolated execution environment (IEE), a hardware-isolated virtual machine (VM), a secure VM, a protected VM execution space, a protected container, etc. The protected execution environment may use various means to prevent unauthorized access and / or modification of code and data while in use (e.g., even from untrusted privileged system software). For example, the protected execution environment may provide one or more of the following: data confidentiality (e.g., where untrusted privileged system software and other unauthorized entities are not permitted to view the data while it is in use within the protected execution environment), data integrity (e.g., where untrusted privileged system software and other unauthorized entities are not permitted to add, remove, or change the data while it is in use within the protected execution environment), and code integrity (e.g., where untrusted privileged system software and other unauthorized entities are not permitted to add, remove, or change the code executing within the protected execution environment). In some embodiments, the protected execution environment may optionally provide one or more of data replay protection, memory remapping protection, etc. Depending on the specific type of protected execution environment, the protection may be provided at the virtual machine level, the per-application level, or the computing function level.
[0039] To further illustrate certain concepts, several specific examples of possible protected execution environments, protected software, and untrusted privileged system software will be described. In some embodiments, the protected execution environment and protected software may be Software Guard Extensions ( Software GuardExtensions, A protected execution environment and protected software in a secure enclave of SGX (SGX). A secure enclave may represent a protected container. A secure enclave may represent code and data in a protected memory area in the address space of a program, where only code in the protected memory area can access the code and data in the protected memory area. Code outside the protected memory area (e.g., untrusted privileged system software) cannot access the code and data in the protected memory area. A secure enclave may be used with a VMM or may not be used with a VMM. When residing outside a processor and only the processor can know the encryption key, the code and data of the secure enclave may be encrypted and integrity protected by the cryptographic unit of the processor. This may help protect the secure enclave even in the presence of untrusted privileged system software.
[0040] In other embodiments, the protected execution environment and the protected software may be Trust Domain Extensions ( Trust Domain Extensions, A protected execution environment and protected software in a TDX Trust Domain (TD). A TD may represent a hardware-isolated VM. A TD may be hardware-isolated from a VMM and other non-TD software. A TD may use a protection management module (referred to as a TDX module) that may run in a secure arbitration mode (SEAM). At a high level, a TDX module may act as a middleware between the protected software and the VMM to help manage the security or protection of the protected software from the VMM.
[0041] In other embodiments, the protected execution environment and protected software can be a protected execution environment and protected software of an ARM realm. A realm can represent a protected execution environment or a protected virtual machine executed in a realm security state. A realm can use a protection management module called a Realm Management Monitor (RMM). At a high level, the RMM module can act as middleware between the protected software and the VMM to help manage the security or protection of the protected software from the VMM.
[0042] In other embodiments, the protected execution environment and the protected software can be the protected execution environment and the protected software of an AMD Secure Encrypted Virtualization (SEV) VM. When the content of the SEV VM is stored in the system memory, the content can be encrypted using the cryptographic key of the SEV VM that is secret from the VMM.
[0043] In still other embodiments, the protected execution environment and the protected software can be the protected execution environment and the protected software of an AMD SEV Secure Nested Paging (SEV-SNP) VM. Similar to SEV, when the content of the SEV-SNP VM is stored in the system memory, the content can be encrypted using the cryptographic key of the SEV-SNP VM that is secret from the VMM. Additionally, access to the pages of the SEV-SNP VM can be restricted based on ownership, page type, and other attributes maintained in a Reverse Map Table (RMP).
[0044] In still other embodiments, the protected execution environment and the protected software can be the protected execution environment and the protected software of Nvidia's Nvidia confidential computing. Isolation can be provided at the virtual machine level or at the multi-user GPU instance level.
[0045] Refer again to Figure 1 , the processor includes circuitry or other logic 113 that supports the protected execution environment. The circuitry / logic 113 can support any of the protected execution environments and / or various types of protected software discussed above (e.g., secure enclaves, TDs, SEV VMs, SEV-SNP VMs, etc.). For example, the logic / circuitry can include: cryptographic logic / circuitry for encrypting the code and data of the protected container before it is transferred to the system memory and decrypting it after the encrypted code and data are loaded into the processor, logic / circuitry for tagging or labeling the decrypted code and data when it resides in the internal structure of the processor and preventing unauthorized entities from accessing the decrypted code and data, logic / circuitry for restricting access to the code and data of the protected software when it is in the system memory (e.g., a memory management unit (MMU)), and so on.
[0046] Refer again to Figure 1, the processor includes at least one logical processor 102. A logical processor may also be referred to as a processor element. Examples of suitable logical processors include, but are not limited to, cores, hardware threads, thread units and thread slots, and other logical processors or processor elements having dedicated contexts (e.g., architectural registers, program counters, or instruction pointers, etc.). The term core is typically used to refer to the logic located on an integrated circuit that can maintain an independent context and where the context is associated with dedicated execution and certain other resources. In contrast, the term hardware thread is typically used to refer to the logic located on an integrated circuit that can maintain an independent context and where the context shares access to execution and certain other resources. When some execution and / or other resources are shared for two or more contexts and other execution and / or other resources are dedicated to a certain context, the boundary between such uses of the terms core and hardware thread may tend to be less significant. However, cores, hardware threads, thread units and thread slots, and other logical processors or processor elements are generally treated as separate logical processors or processor elements by software. Software threads, processes, or workloads can be scheduled on each of cores, hardware threads, thread units and thread slots, and other logical processors or processor elements, and are independently associated with each of hardware threads, thread units and thread slots, and other logical processors or processor elements.
[0047] The logical processor includes data processing circuitry 104 such as an instruction processing pipeline (e.g., a decoding unit, an execution unit, etc.). The data processing circuitry can execute software (including protected software 117) and perform data processing on a context 106. The context may include execution states and / or data stored in the architectural registers and / or other architectural storage devices of the logical processor. These architectural registers and / or other architectural storage devices are shown in the figure as context storage 105. Examples of suitable context storage include, but are not limited to, a set of general-purpose registers, one or more sets of vector registers having one or more vector widths, a set of mask registers, a set of tile registers or other tile storage devices for storing matrices or other two-dimensional data, other types of architectural storage, and various combinations of the above. A specific example of suitable context storage includes Figure 19 the registers shown in. The context may represent data or values in such context storage. In some embodiments, the processor may optionally include one or more coprocessors and / or accelerators 114 (e.g., a shared matrix processing accelerator, a shared vector execution unit, a shared migration engine, etc.), and the context may also optionally include the context 115 of one or more coprocessors and / or accelerators.
[0048] The logical processor includes a context save and restore unit 108. The processor may periodically exchange at least some of the context 106 between the context storage device 105 and the system memory 116 (e.g., a context save area 121 in the system memory). For example, when the logical processor enters the protected execution environment 103, the logical processor may restore or load the context 122 from the system memory (e.g., from the context save area 121) into the context storage device 105 as the context 106. Similarly, the logical processor may exit the protected execution environment either synchronously (e.g., to make a system call) or asynchronously (e.g., due to a timer interrupt). When exiting the protected execution environment, the logical processor may save or store the context 106 from the context storage device 105 in the system memory 116 as the context 122. Such context save and restore may occur when switching threads, at interrupts, when untrusted privileged system software needs to handle certain events, and for other reasons.
[0049] Saving and restoring context saves power and takes time, which may tend to limit performance. This may be especially true when there is a large amount of context and / or when the context is saved and restored frequently (e.g., when the system is busy and asynchronous exits are frequent, and / or when the protected execution environment frequently calls untrusted / unprotected code). Moreover, many current processors have a large amount of context, such as, for example, as a result of many registers, wide registers (e.g., wide vector registers), on-chip registers, or other on-chip storage devices, etc.
[0050] In some embodiments, the context save and restore unit may include a selective save circuit or other logic 109 that operates to selectively save and / or store a subset of the context 106 from the processor's context storage device 105 to the system memory 116. In some embodiments, the selective save circuit / logic may selectively save / store a first subset of the context, which may be modified or otherwise written after entry into the protected execution environment and during execution within the protected execution environment, from the processor's context storage device to the context in the system memory. However, the selective save circuit / logic may determine not to save / store a different second subset of the context, which has not been modified or otherwise written after entry into the protected execution environment and during execution within the protected execution environment, from the processor's context storage device to the context in the system memory, and may not save / store the different second subset of the context from the processor's context storage device to the context in the system memory. For example, the context in a first subset of a set of vector registers that has been written during execution within the protected execution container may be selectively saved / stored to the system memory, but the context in a different second subset of the set of vector registers that has not been written during execution within the protected execution environment may not be saved / stored to the system memory. For example, during a context save operation, the context save logic may bypass portions of the context storage device where no context change has occurred since the most recent context save operation, and may selectively save the context from portions of the context storage device where a context change has occurred since the most recent context save operation. Advantageously, saving / storing only a subset of the context to the system memory may tend to help reduce latency and / or improve performance and / or reduce power consumption.
[0051] In some embodiments, the context save and restore unit may include a selective sanitization circuit or other logic 111 that operates to selectively "sanitize" a subset of the context storage device 105. Although within the protected execution environment, the logic / circuit 113 may help protect the confidentiality of the context 106. However, in some embodiments, after exiting the protected execution environment, such protection may not be available, and / or the risk of an untrusted entity reading the context from the context storage device may be higher. In some embodiments, when exiting the protected execution environment, the selective sanitization circuit / logic 111 may operate to selectively sanitize a first subset of the context storage device 105 that has been read and / or written after entering the protected execution environment 103 and during execution within the protected execution environment 103. However, when exiting the protected execution environment, the selective sanitization circuit / logic may determine not to sanitize a different second subset of the context storage device 105 that has not been read or written after entering the protected execution environment 103 and during execution within the protected execution environment 103, and may not sanitize the different second subset of the context storage device 105. The sanitization of the different second subset of the context storage device may be omitted or avoided. For example, a first subset of a set of vector registers that has been read and / or written after entering the protected execution environment and during execution within the protected execution environment may be selectively sanitized, but a different second subset of the set of vector registers that has not been read or written after entering the protected execution environment and during execution within the protected execution environment may not be sanitized. The subset of the context storage device that has been written represents the subset of the context storage device that contains confidential information and is thus the subset of the context storage device that may be selectively sanitized. Examples of suitable ways to sanitize the context storage device include, but are not limited to: overwriting the context 106 (e.g., with all 0s, all 1s, a predetermined meaningless value, a predetermined bit pattern, or some other non-confidential value), erasing the context in the context storage device, resetting the context storage device to a predetermined value, default value, or reset value, encrypting, scrambling, or otherwise obfuscating the context in the context storage device, destroying the context in the context storage device in a way that obfuscates the context, wiping the context storage device clean, or otherwise changing or altering the context. Advantageously, selectively sanitizing only a subset of the context and / or the context storage device may tend to help reduce latency and / or improve performance and / or reduce power consumption.
[0052] In some embodiments, the context save and restore unit may include a lazy restore circuit or other logic 110 that operates to lazily restore and / or load at least some of the context 122 from the system memory 116 (e.g., from the context save region 121) into the context storage device 105 as context 106. In some embodiments, the lazy restore circuit / logic may lazily and / or as needed and / or dynamically and / or when needed and / or selectively load such context during the operation or execution of the protected software within the protected execution environment 103 and throughout the operation or execution. In some embodiments, a first subset of the context 122 that is read, written, or otherwise needed after entry into the protected execution environment and during execution within the protected execution environment may be selectively restored / loaded from the system memory (e.g., from the context save region) into the context storage device of the processor. However, a different second subset of the context 122 that is not read, written, or otherwise needed after entry into the protected execution environment and during execution within the protected execution environment may not be restored / loaded from the system memory (e.g., from the context save region) into the context storage device of the processor. For example, portions of the context 122 may be lazily and / or as needed and / or dynamically and / or when needed and / or selectively restored / loaded when the instructions of the protected software are executed and access the corresponding portions of the context storage device 106 (e.g., registers) (e.g., read from and / or written to the corresponding portions). Another possible approach is to eagerly restore / load all context in one go immediately upon entry into the protected execution environment. However, such eager restoration of all context may restore some context that is never needed, which may unnecessarily consume power and take additional time needlessly, potentially tending to degrade performance.
[0053] In some embodiments, the context save and restore unit may include a lazy access control check circuit or other logic 112 that operates to lazily perform access control checks for pages or other portions of system memory 116 used by the protected execution environment 103. Certain types of protected execution environments may eagerly or preemptively perform certain page access control checks when a thread enters the protected execution environment. For example, a protected execution environment may use certain pages to securely store processor context under certain conditions (e.g., during an asynchronous exit). As an example, in SGX, State Save Area (SSA) pages may be used to store context during an asynchronous exit, and conventionally, enclave page cache map (EPCM) access control checks may be eagerly or preemptively performed on these SSA pages. Conventionally, such access control checks are eagerly or preemptively performed to ensure that there is a valid location to store data under certain conditions when exiting the protected execution environment (e.g., during an asynchronous exit). If instead the checks are not performed until the context is written, there is a possibility that the checks may fail, meaning that there is no secure location to store the modified context, potentially resulting in data loss. The operations performed during the access control check may vary depending on the type of protected execution environment. Commonly, the access control check may perform one or more checks to check whether a page is available and can be used or to ensure that a page is available and can be used (e.g., the page exists, the page has the correct page type, the page has appropriate ownership, the page has the correct access permissions (e.g., read and / or write and / or execute permissions), or any combination of the above). Conventional execution environments generally do not need to preemptively perform such page access control checks because they can trust privileged system software (e.g., the OS and / or VMM) to store the processor context at a valid location, but the same does not hold true for certain protected execution environments.
[0054] In some embodiments, lazy access control checking circuitry / logic 112 may be operable to perform access control checks on pages or other portions of system memory 116 used by protected execution environment 103 lazily and / or as needed and / or dynamically and / or when needed and / or selectively during and throughout operation or execution of protected software 117 within protected execution environment 103. In some embodiments, access control checks may optionally be performed on a first subset of pages that are actually needed under certain conditions (e.g., asynchronous exit) to store corresponding context that has been read, written, or modified after entry into the protected execution environment and during execution within the protected execution environment. However, access control checks may optionally be omitted, skipped, bypassed, or otherwise passed over for a different second subset of pages that are not actually needed to store corresponding context under these conditions (e.g., asynchronous exit) because the context has not been read, written, or modified after entry into the protected execution environment and during execution within the protected execution environment. For example, access control checks may optionally be delayed or deferred and performed lazily and / or on demand and / or dynamically and / or when needed and / or selectively as instructions of the protected software are executed and access context corresponding to the first subset of pages (e.g., context that may be stored in those pages under certain conditions) (e.g., reading from and / or writing to the context corresponding to the first subset of pages).
[0055] If the access control check fails at this point in time, the processor can abandon the read, write, or modify operation, take an exit from the protected execution environment, and signal the error without losing context. Advantageously, avoiding some of the access control checks that are not needed can tend to help reduce latency and / or improve performance and / or reduce power consumption. Note that this may introduce page access control checks for instructions that interact with the context, even in cases where the instructions do not originally access memory (e.g., register-to-register move instructions, instructions that add two registers and store the sum to a destination register, etc.). Moreover, this may cause such instructions (e.g., those that do not originally access memory) to signal page faults at irregular intervals.
[0056] In some embodiments, to help support the selective save circuitry / logic 109, the selective purge circuitry / logic 111, the lazy restore circuitry / logic 110, and the lazy access control check circuitry / logic 112, the logical processor may operate to maintain context metadata 107. The logical processor may have circuitry or other logic for monitoring and detecting accesses to the context storage 105 (e.g., reads from and writes to the context storage 105), and recording or otherwise storing indications of such accesses in the context metadata 107. For example, the context metadata may include a collection of context element metadata for the corresponding context elements of the context. The context element metadata may broadly represent one or more bits or values that describe or indicate that a context element has been accessed by at least one retired instruction since entering the protected execution environment. In some embodiments, the context element metadata may include a first set of one or more bits or values for indicating whether a context element has been read from, and a second set of one or more bits or values for indicating whether a context element has been written to.
[0057] In some embodiments, the context metadata 107 may include a set of read indication bits (e.g., an array of read indication bits, referred to herein as R[]). Each read indication bit (e.g., R[i]) may correspond to a different context element in a set of context elements. Each read indication bit (e.g., R[i]) may either have a first bit value (e.g., cleared to binary 0) indicating that the corresponding context element has not been read from by a retired operation since entering the protected execution environment, or a different second bit value (e.g., set to binary 1) indicating that the corresponding context element has been read from by a retired operation since entering the protected execution environment. Similarly, the context metadata may include a set of write indication bits (e.g., an array of write indication bits, referred to herein as W[]). Each write indication bit (e.g., W[i]) may correspond to a different context element in the same set of context elements. Each write indication bit (e.g., W[i]) may either have a first bit value (e.g., cleared to binary 0) indicating that the corresponding context element has not been written to by a retired operation since entering the protected execution environment, or a different second bit value (e.g., set to binary 1) indicating that the corresponding context element has been written to by a retired operation since entering the protected execution environment. As a specific example, if the logical processor includes 32 vector registers and 8 matrix tiles, the size of R[] may be 40 bits and the size of W[] may be 40 bits, where each bit corresponds to a single register / tile. In some embodiments, each logical processor may have a different corresponding set of read indication bits and a different corresponding set of write indication bits.
[0058] The context element metadata 107 may be stored in different places in different embodiments. Figure 2 FIG. is a block diagram of a first exemplary embodiment in which context element metadata 227 is stored together with the corresponding context element 226. The first context element 226-1, the second context element 226-2, and (optionally) other context elements are shown. The context elements may represent elements of a context (e.g., a register or a set of registers, a slice or a set of slices, etc.). The first context element metadata 227-1 corresponds to the first context element and is stored together with the first context element. The second context element metadata 227-2 corresponds to the second context element and is stored together with the second context element. As an example, the context element metadata may be stored in one or more bits similar to those used to store error correction codes for registers, contamination indicators for registers, etc. There may optionally be other context element metadata for other context elements. In some embodiments, the first context element metadata may include a first read indication bit (e.g., R[1]) in a set of read indication bits (e.g., R[]) and a first write indication bit (e.g., W[1]) in a set of write indication bits (e.g., W[]), and the second context element metadata may include a second read indication bit (e.g., R[2]) in a set of read indication bits (e.g., R[]) and a second write indication bit (e.g., W[2]) in a set of write indication bits (e.g., W[]).
[0059] Figure 3Block diagram of a second exemplary embodiment in which register 328 is used to store context element metadata 327 from corresponding context elements 326. The first context element 326-1, the second context element 326-2, and (optionally) other context elements are shown. Context elements can represent elements of a context (e.g., a register or a set of registers, a slice or a set of slices, etc.). The first context element metadata 327-1 corresponds to the first context element. The second context element metadata 327-2 corresponds to the second context element. In this embodiment, the first and second context element metadata are stored in register 328, which can be, for example, a model specific register (MSR) or other control and / or configuration register. The register can typically be physically separate from the register or other context storage device corresponding to the first and second context elements. The register can also optionally store other context element metadata for other context elements. In other embodiments, instead of register 328, the context element metadata can optionally be stored in an on-die hardware table or other hardware structure, a table or other data structure in a sufficiently protected memory, or other location. In some embodiments, the first context element metadata can include a first read indication bit (e.g., R[1]) in a set of read indication bits (e.g., R[]) and a first write indication bit (e.g., W[1]) in a set of write indication bits (e.g., W[]), and the second context element metadata can include a second read indication bit (e.g., R[2]) in a set of read indication bits (e.g., R[]) and a second write indication bit (e.g., W[2]) in a set of write indication bits (e.g., W[]).
[0060] Context element metadata can be defined at different context granularity levels in different embodiments. Figure 4Block diagram of a first exemplary embodiment in which context element metadata 427 is defined at the register granularity. Register set 406 includes a first register 426-1 representing a first context element, a second register 426-2 representing a second context element, and (optionally) other registers representing other context elements. As an example, the register set may represent a vector register set of an architecture, and the first and second registers may represent the first and second vector registers in the vector register set. First context element metadata 427-1 corresponds to the first register. Second context element metadata 427-2 corresponds to the second register. As previously described, the context element metadata may be stored with these registers in a separate register (e.g., a control and / or configuration register), a hardware structure on another die, or other locations. In some embodiments, the first context element metadata may include a first read indication bit (e.g., R[1]) in a set of read indication bits (e.g., R[]) and a first write indication bit (e.g., W[1]) in a set of write indication bits (e.g., W[]), and the second context element metadata may include a second read indication bit (e.g., R[2]) in a set of read indication bits (e.g., R[]) and a second write indication bit (e.g., W[2]) in a set of write indication bits (e.g., W[]). In other embodiments, the context element metadata may optionally be defined at the granularity of register portions.
[0061] Figure 5Block diagram of a second exemplary embodiment in which context element metadata 527 is defined at the register set granularity. A first associated register set (e.g., a vector register set) represents a first context element 526-1, a second associated register set (e.g., a general-purpose register set) represents a second context element 526-2, and optionally, other register sets may represent other context elements. The register set may either include a complete set or a subset containing two or more registers. The first context element metadata 527-1 corresponds to the first register set. The second context element metadata 527-2 corresponds to the second register set. As previously described, the context element metadata may be stored with these registers in a separate register (e.g., a control and / or configuration register), another die structure, or other locations. In some embodiments, the first context element metadata may include a first read indication bit (e.g., R[1]) in a set of read indication bits (e.g., R[]) and a first write indication bit (e.g., W[1]) in a set of write indication bits (e.g., W[]), and the second context element metadata may include a second read indication bit (e.g., R[2]) in a set of read indication bits (e.g., R[]) and a second write indication bit (e.g., W[2]) in a set of write indication bits (e.g., W[]).
[0062] Figure 6 Flow block diagram of an embodiment of a method 630 for entering a protected execution environment. In some embodiments, the method may correspond to and / or be executed in response to a control primitive for entering a protected execution environment. In various embodiments, the method may be executed by a processor, digital logic device, or integrated circuit. As an example, processor 101 or logical processor 102 may execute the method. The components, features, and specific optional details described herein for processor 101 or logical processor 102 may also optionally apply to the method. Alternatively, method 630 may be executed by a similar or different processor or logical processor. Moreover, processor 101 and logical processor 102 may execute the same, similar, or different methods as method 630.
[0063] The method includes: at block 631, receiving a control primitive for entering a protected execution environment. Various different types of control primitives are suitable. An example of a suitable control primitive is a machine language instruction and / or an instruction in the instruction set of a processor. The instruction may have an opcode or operation code that at least partially specifies an operation to be performed (e.g., an operation for entering a protected execution environment). An example of a suitable instruction for entering a protected execution environment is the ERESUME instruction in SGX that has been modified to perform the method. Another example of a suitable instruction for entering a protected execution environment is the SEAMCALL[TDH.VP.ENTER] instruction in TDX that has been modified to perform the method. These are merely a few illustrative examples.
[0064] Another example of a suitable control primitive is an Application Binary Interface (ABI) command, an Application Programming Interface (API) command, or other commands (e.g., command codes, command identifiers, etc.) that can be stored in one or more control registers (e.g., one or more memory-mapped input and / or output (MMIO) registers). The command may at least partially specify an operation to be performed. In some cases, one or more additional data structures and / or storage locations explicitly specified or otherwise indicated (e.g., implicitly indicated) by the command may optionally be used to provide additional information for specifying the operation to be performed, for providing output values, and for receiving output values (e.g., a first destination storage location).
[0065] The method includes: at block 632, omitting / delaying the restoration of at least some context from system memory (e.g., from a context save area in system memory) to a context storage device of the processor (e.g., to allow the context to be restored lazily and / or on demand). As discussed above, the context may be restored / loaded lazily and / or on demand and / or dynamically and / or when needed and / or selectively during the operation or execution of protected software within the protected execution environment and throughout the operation or execution. For example, portions of the context (e.g., data in registers) may be restored / loaded lazily and / or on demand and / or dynamically and / or when they are needed and / or selectively when the instructions of the protected software are executed and access the corresponding portions of the context storage device of the processor (e.g., read from and / or write to the corresponding portions).
[0066] In some embodiments, omitting and / or delaying the restoration of context can optionally be applied to all contexts of the processor (e.g., regardless of the number of such contexts, the degree of likelihood that such contexts will be modified during execution within the protected execution environment, etc.). In other embodiments, omitting and / or delaying the restoration of context can optionally be applied only to a subset of all contexts of the processor. For example, omitting and / or delaying the restoration of context can optionally be applied to a subset of contexts having a relatively large amount of data (e.g., vector registers and scratchpad memory), but not to another subset of contexts having a relatively small amount of data (e.g., instruction pointer, general-purpose registers). As another example, omitting and / or delaying the restoration of context can optionally be applied to a subset of contexts that are more likely to not be used (e.g., the widest set of vector registers, scratchpad memory, accelerator contexts, etc.), but not to another subset of contexts that will typically always be used (e.g., instruction pointer, general-purpose registers, etc.). Moreover, embodiments can use heuristics to determine which contexts are frequently used by a logical processor within the protected execution environment, and optionally eagerly restore at least some of the frequently used contexts.
[0067] The method includes: at block 633, for a context whose restoration has been delayed, changing or updating context element metadata to indicate that operations that have not been retired or otherwise committed since entering the protected execution environment read from or write to the corresponding context element. For example, this can include: clearing or otherwise changing or updating a set of read indication bits (e.g., R[]) such that they indicate that operations that have not been retired since entering the protected execution environment read from the corresponding context element; and clearing or otherwise changing or updating a set of write indication bits (e.g., W[]) such that they indicate that operations that have not been retired since entering the protected execution environment write to the corresponding context element. Subsequently, the read and write indication bits can be updated when reading from and writing to their corresponding context elements in order to track context elements that have been accessed during execution within the protected execution environment.
[0068] The method may optionally include: at block 634, omitting and / or delaying access control checks for pages or other portions of system memory to be used by the protected execution environment. As discussed above, in some embodiments, these access control checks may optionally be performed lazily and / or as needed and / or dynamically and / or when needed and / or selectively during and throughout the operation or execution of protected software within the protected execution environment. For example, the access control checks may optionally be delayed or postponed and performed lazily and / or as needed and / or dynamically and / or when needed and / or selectively when the instructions of the protected software are executed and access context corresponding to a first subset of pages (e.g., context that may be stored in those pages under certain conditions) (e.g., reading from and / or writing to context corresponding to a first subset of pages). For example, in some embodiments, the access control checks may optionally be delayed or postponed until the corresponding context is first read, written, or modified after entry into the protected execution environment. Delaying the access control checks at block 634 is optional. Moreover, some types of protected execution environments may not have to implement such access control checks and thus may not need to delay such access control checks. The method includes: at block 635, entering a protected execution container. Changing the operating mode to a mode for the protected execution environment (e.g., this may include changing one or more bits in one or more control and / or status accumulators).
[0069] Figure 6 is merely an example method. Upon initial entry into the protected execution environment, some types of protected execution environments initialize certain context elements to a predetermined context or value. Such initialization also takes time and consumes power. In some embodiments, as Figure 6 the method may also include selective initialization of a subset of context elements to allow the initialization to be delayed and performed lazily and / or as needed and / or only when needed. At least some such initialization may optionally be omitted entirely in cases where the corresponding context elements are never actually accessed and / or in cases where writing to a corresponding context element occurs before reading from the corresponding context element.
[0070] Figure 7 is a block diagram of an embodiment of a device 736 that operates to execute a control primitive to exit the protected execution environment or asynchronously exits the protected execution environment due to an exception condition 798. In various embodiments, the device may be a processor, an integrated circuit, a system-on-chip (SoC), or a computer system.
[0071] The device may be coupled to receive a control primitive 738. Suitable types of control primitives include, but are not limited to, those described above forFigure 6 Those being discussed (e.g., instructions in a processor's instruction set, ABI commands, API commands, another type of command, etc.). A specific example of a suitable instruction is the ERESUME instruction in SGX that has been modified to have selective context save and selective context store device sanitization aspects and uses context metadata as further described below. In other embodiments, instead of exiting due to a control primitive, an asynchronous exit can be initiated due to an exceptional condition 798 (e.g., an interrupt, an exception, etc.).
[0072] The device includes a context store 705. Examples of suitable context stores include, but are not limited to, a set of general-purpose registers, one or more sets of vector registers with different vector widths, mask registers, other types of architectural registers, a set of tile registers or other tile storage devices for storing matrices or other two-dimensional data, the context stores of one or more coprocessors or accelerators, various other types of architectural storage devices, and various combinations of the foregoing.
[0073] The context store is used to store contexts. In the illustrated example, contexts are grouped into four types or four bins. Each of the four context types or four context bins has a different one of four combinations of context metadata 707. Specifically, in the illustrated example embodiment, the context metadata includes a set of read indication bits R[] for indicating whether the content has been read (marked as "yes") or has not been read (marked as "no") during execution after entering the protected execution environment. Similarly, the illustrated example context metadata also includes a set of write indication bits W[] for indicating whether the content has been written (marked as "yes") or has not been written (marked as "no") during execution after entering the protected execution environment. The first context 706-1 is a subset of contexts whose corresponding metadata bits indicate that it has not been read and has not been written. The second context 706-2 is a subset of contexts whose corresponding metadata bits indicate that it has been read but has not been written. The third context 706-3 is a subset of contexts whose corresponding metadata bits indicate that it has not been read but has been written. The fourth context 706-4 is a subset of contexts whose corresponding metadata bits indicate that it has been read and has been written. The context metadata can be stored anywhere as previously described (e.g., stored with the corresponding context, stored in control and / or configuration registers, stored in a dedicated structure of the processor, stored in a table in memory, etc.).
[0074] The device also includes an execution unit 737. The execution unit can broadly represent circuitry or other logic that operates to perform operations corresponding to control primitive 738. In some embodiments, the execution unit can operate to selectively save / store, from the context storage device 705, contexts that have been written during execution after entering and while within the protected execution environment to the system memory (e.g., to a context save area in the system memory), without saving / storing contexts that have not been written during execution after entering and while within the protected execution environment from the context storage device to the system memory. For example, in the illustrated embodiment, the execution unit can save (i.e., operation 740) the third context 706-3 and the fourth context 706-4 to the system memory (e.g., to a context save area in the system memory), but not save the first context 706-1 or the second context 706-2 to the system memory. In some embodiments, the execution unit can optionally include a selective context save unit 709 for selectively saving such contexts that have been written. The selective context save unit can include hardware, firmware, software, or a combination thereof (e.g., at least some circuitry potentially combined with some firmware).
[0075] In some embodiments, the execution unit may optionally operate to selectively sanitize contexts and / or context storage devices that have been read and / or written during execution after entering the protected execution environment and while within the protected execution environment, without sanitizing contexts and / or context storage devices that have not been read and have not been written during execution after entering the protected execution environment and while within the protected execution environment. For example, in the illustrated embodiment, the execution unit may optionally sanitize (i.e., operation 743) the second context 706-2, the third context 706-3, and the fourth context 706-4. Sanitizing these contexts may also represent sanitizing their corresponding context storage devices. Examples of suitable ways to sanitize include, but are not limited to: overwriting the context and / or context storage device (e.g., with all 0s, all 1s, a predetermined meaningless value, or some other benign value), resetting the context storage device to a predetermined value, default value, or reset value, encrypting, scrambling, or otherwise obfuscating the context in the context storage device, destroying the context in the context storage device in a way that obfuscates the context, or otherwise changing or altering the context. However, for the illustrated example embodiment, since the first context 706-1 has not been read and has not been written, the execution unit may not sanitize the first context 706-1. That is, sanitization of at least a portion of the context that does not contain confidential information may optionally be omitted. In some embodiments, the execution unit may optionally include a selective context and / or context storage device sanitization unit 711 for selectively sanitizing contexts and / or context storage devices. The selective context and / or context storage device sanitization unit may include hardware, firmware, software, or a combination thereof (e.g., at least some circuitry potentially combined with some firmware).
[0076] The execution unit may also operate to exit the protected execution environment. This may be done in different ways in different embodiments depending on the type of protected execution environment. In some embodiments, this may include providing an exit control to change the operating mode to a mode for untrusted software (e.g., in some cases, this may be done by writing one or more bits in one or more control and / or configuration registers). In some embodiments, metadata may optionally be recorded (e.g., to describe the reason for the exit). In some embodiments, the execution unit may optionally include a protected execution environment exit unit 739 for exiting the protected execution environment. The protected execution environment exit unit may include hardware, firmware, software, or a combination thereof (e.g., at least some circuitry potentially combined with some firmware).
[0077] Figure 8FIG. 0 is a block diagram of an embodiment of a processor 801 that implements an embodiment of an instruction 838 for exiting a protected execution environment. General purpose processors and special purpose processors of the type previously described are suitable. The processor may have any architecture of a variety of complex instruction set computing (CISC) architectures, reduced instruction set computing (RISC) architectures, very long instruction word (VLIW) architectures, or hybrid architectures. In some embodiments, the processor may include at least one integrated circuit or semiconductor die (e.g., disposed on at least one integrated circuit or semiconductor die). In some embodiments, the processor may include at least some hardware (e.g., transistors, capacitors, circuits, non-volatile memory that stores circuit-level instructions / control signals).
[0078] The processor includes a decode unit 845 (e.g., including decode circuitry). The decode unit may be coupled to receive the instruction 838. The instruction may represent a macro instruction, a machine code instruction, or other instructions in the instruction set of the processor. The instruction may have an opcode or operation code that at least partially specifies an operation to be performed (e.g., for exiting a protected execution environment, for selectively storing context, optionally for selectively purging, etc.). The instruction may have various formats or encodings, such as, for example, those further described below (e.g., for Figure 20-Figure 25 ).
[0079] The decode unit may be operative to decode the instruction into one or more lower-level control signals, operations, or decoded instructions 846 (e.g., one or more microinstructions, micro-operations, microcode entry points, etc.). Various instruction decoding mechanisms may be used to implement the decode unit, including but not limited to microcode read only memory (ROM), lookup tables, hardware implementations, programmable logic arrays (PLAs), other mechanisms suitable for implementing an instruction decode unit, and combinations of the above. In some embodiments, the decode unit may include at least some hardware (e.g., transistors, integrated circuits, on-die read only memory, or other non-volatile memory that stores microcode or other hardware-level instructions, or any combination of the above). In some embodiments, the decode unit may be included on a die, integrated circuit, or semiconductor substrate.
[0080] The execution unit 837 (e.g., including execution circuitry) is coupled to the decoding unit 845 (e.g., to receive one or more lower-level control signals, operations, or decoded instructions 846). In some embodiments, the execution unit may be on the die or integrated circuit together with the decoding unit. The execution unit may be operative to perform operations corresponding to the instructions. For example, one or more lower-level control signals, operations, or decoded instructions may control the execution unit to perform operations corresponding to the instructions. The operations may be the same as or similar to those Figure 7 previously described. To avoid obscuring the present specification, those operations will not be repeated. As previously described, the processor may also have architectural registers and / or other context storage means (not shown).
[0081] Figure 9 FIG. is a block diagram of an embodiment of a device 936 that is operative to execute a command 938 for exiting a protected execution environment. In embodiments, the device may be a processor, integrated circuit, system-on-chip (SoC), or computer system.
[0082] The device includes one or more control registers 947 that are coupled to receive and store the command 938. In some embodiments, the control registers may optionally be memory-mapped input / output (MMIO) registers, mailbox registers, etc. The command may be an ABI command, API command, or another type of command. The command may include a command code, command identifier, etc. The command code or command identifier may at least partially specify the operation to be performed (e.g., for exiting a protected execution environment, for selectively storing context, optionally for selectively purging, etc.). In some cases, one or more registers and / or data structures and / or storage locations explicitly specified or otherwise indicated (e.g., implicitly indicated) by the command may optionally be used to provide additional information for specifying the operation to be performed, for providing output values, and for receiving output values and / or status.
[0083] The execution unit 937 (e.g., including execution circuitry) is coupled to the (one or more) control registers. The execution unit may also be coupled to access source and destination storage locations and / or data structures. As shown, in some embodiments, the execution unit may optionally be part of a security processor (e.g., a security coprocessor), but this is not required. The execution unit may include hardware, software, firmware, or a combination thereof. The execution unit may be operative to perform operations corresponding to the command. The operations may be the same as or similar to those Figure 7 previously described. To avoid obscuring the present specification, those operations will not be repeated.
[0084] Figure 10 It is a flowchart of an embodiment of method 1050 for exiting a protected execution environment. In some embodiments, the method may correspond to and / or be executed in response to a control primitive for exiting a protected execution environment. In other embodiments, the method may correspond to and / or be executed in response to an exception condition (e.g., an exception, an interrupt, etc.). In various embodiments, the method may be executed by a processor, digital logic device, or integrated circuit. As an example, device 736 and / or execution unit 737 may execute the method. The components, features, and specific optional details described herein for device 736 and / or execution unit 737 may also optionally apply to the method. Alternatively, method 1050 may be executed by a similar or different device, logical processor, or execution unit. Moreover, device 736 and execution unit 737 may execute the same, similar, or different methods as method 1050.
[0085] The method includes: at block 1051, receiving a control primitive or an exception condition for causing an exit from a protected execution environment. Suitable types of control primitives include, but are not limited to, those control primitives discussed (e.g., an instruction in a processor's instruction set, an ABI command, an API command, another type of command).
[0086] The method includes: at block 1052, selectively saving and / or storing the context that has been written during execution within the protected execution environment from the processor's context storage device to system memory (e.g., saving and / or storing to a context save area in system memory). In some embodiments, a first subset of the context that has been modified or otherwise written during execution within the protected execution environment may be selectively saved and / or stored from the processor's context storage device to system memory (e.g., saving and / or storing to a context save area in system memory). However, a determination may be made that a different second subset of the context that has not been modified or otherwise written during execution within the protected execution environment (e.g., at least some context) may not be saved and / or stored from the processor's context storage device to system memory. In some embodiments, the protected execution environment may have additional or separate control over what context is selectively saved (e.g., control over skipping saving some context). For example, SGX and TDX may cause the saving of some context to be skipped.
[0087] The method optionally includes, at block 1053, selectively purifying contexts and / or context storage devices that have been read and / or written during execution within the protected execution environment. In some embodiments, the first subset of contexts and / or the first subset of context storage devices that have been read and / or written during execution within the protected execution environment may be selectively purified when exiting the protected execution environment. However, it may be determined that different second subsets of contexts and / or different second subsets of context storage devices that have not been read or written during execution within the protected execution environment may not be purified when exiting the protected execution environment. That is, purification of at least portions of the contexts that do not contain confidential information may optionally be omitted. Examples of suitable ways to perform purification include, but are not limited to, overwriting the contexts and / or context storage devices (e.g., with all 0s, all 1s, a predetermined meaningless value, or some other benign value), resetting the context storage devices to a predetermined value, default value, or reset value, encrypting, scrambling, or otherwise obfuscating the contexts in the context storage devices, destroying the contexts in the context storage devices in a way that obfuscates the contexts, or otherwise changing or altering the contexts.
[0088] The method includes, at block 1054, exiting the protected execution container. This may be done as previously described.
[0089] Figure 11 FIG. 7 is a block diagram of an embodiment of a processor 1101 that is operative to execute context access instructions 1160. In some embodiments, the processor may be a general-purpose processor (e.g., the type of general-purpose microprocessor or central processing unit (CPU) used in desktop computers, laptop computers, servers, and other computer systems). Alternatively, the processor may be a special-purpose processor. Examples of suitable special-purpose processors include, but are not limited to, coprocessors, graphics processors (e.g., general-purpose GPUs), security processors, machine learning processors, artificial intelligence processors, network processors, and controllers (e.g., microcontrollers). The processor may have any architecture among various complex instruction set computing (CISC) architectures, reduced instruction set computing (RISC) architectures, very long instruction word (VLIW) architectures, hybrid architectures, and other types of architectures. In some embodiments, the processor may include at least one integrated circuit or semiconductor die (e.g., disposed on at least one integrated circuit or semiconductor die), and may include at least some hardware (e.g., transistors, circuits, etc.).
[0090] The processor may be coupled to receive a context access instruction 1160 (e.g., from system memory). The context access instruction may represent a macro instruction, a machine code instruction, or other instruction in the instruction set of the processor. The context access instruction broadly represents an instruction that, when executed, accesses (e.g., reads, writes, or reads and writes) at least some context (e.g., one or more architectural registers, on-chip storage, accelerator context, etc.). In some embodiments, the context access instruction may explicitly specify or otherwise indicate (e.g., implicitly indicate) one or more source operands and / or one or more destination operands (e.g., via one or more fields or sets of bits). The number and type of operands may vary depending on the type of context access instruction. As an illustrative example, the context access instruction may be an instruction for moving context from a source vector register to a destination vector register, and the context access instruction may specify or otherwise indicate the source vector register and the destination vector register. As another illustrative example, the context access instruction may be an instruction for performing an arithmetic operation on first and second source vector registers and storing the resulting vector in a destination vector register, and the context access instruction may specify or otherwise indicate the first and second source vector registers and the destination vector register. In some cases, the instruction may have source and / or destination operand specification fields to specify the register, memory location, or other storage location for the operand. In other cases, the register or other context may be implicit for the instruction and need not be specified via a field. Combinations of these approaches may also be used. The instruction may have one or more fields for an opcode that at least partially or fully specifies the operation to be performed. The instruction may have various formats or encodings, such as, for example, those further described below (e.g., for Figure 20-Figure 25 ).
[0091] Referring again to Figure 11, the processor includes a decoding unit 1145 (e.g., a decoding circuit). The decoding unit can be coupled to receive context access instructions. The decoding unit can be operative to decode the context access instructions into one or more lower-level control signals, operations, or decoded instructions (e.g., one or more microinstructions, micro-operations, microcode entry points, etc.). In some embodiments, the decoding unit can include: at least one input structure (e.g., a port, an interconnect, or an interface), coupled to receive context access instructions; instruction recognition and decoding logic, coupled to the at least one input structure to recognize the context access instructions and decode the context access instructions into one or more lower-level control signals, operations, or decoded instructions; and at least one output structure (e.g., a port, an interconnect, or an interface), coupled to the instruction recognition and decoding logic to output the one or more lower-level control signals, operations, or decoded instructions. Various instruction decoding mechanisms can be used to implement the decoding unit and / or its instruction recognition and decoding logic, and the various instruction decoding mechanisms include, but are not limited to, microcode read-only memory (ROM), look-up tables, hardware implementations, programmable logic arrays (PLAs), and other mechanisms suitable for implementing an instruction decoding unit, as well as combinations of the foregoing. In some embodiments, the decoding unit can include at least some hardware (e.g., transistors, integrated circuits, on-die read-only memory, or other non-volatile memory that stores microcode or other hardware-level instructions, or any combination of the foregoing). In some embodiments, the decoding unit can be included on a die, an integrated circuit, or a semiconductor substrate.
[0092] The processor further includes a backend unit and / or circuit 1161, and the backend unit and / or circuit 1161 includes several components as shown. In some examples, the register renaming, register allocation, and / or scheduling circuit 1162 can provide functions for one or more of the following: 1) renaming logical operand values to physical operand values (e.g., a register alias table in some examples); 2) allocating status bits and flags to the decoded instructions; and 3) scheduling the decoded instructions from an instruction pool for execution by an execution circuit (e.g., using reservation stations in some examples).
[0093] The processor further includes a context storage device 1105. The context storage device is used to store context elements. The context elements have corresponding context metadata 1107. The context metadata can be stored in various places described previously.
[0094] Execution unit 1137 (e.g., execution circuitry) is coupled to decoding unit 1145 (e.g., to receive one or more lower-level control signals, operations, or decoded instructions). In some embodiments, the execution unit may be on the die or integrated circuit together with the decoding unit. The execution unit may be operative to perform operations corresponding to context access instructions 1160. For example, one or more lower-level control signals, operations, or decoded instructions may be executed by a control unit to control the execution unit to perform operations corresponding to context access instructions. The operations may include accessing one or more context elements in context storage device 1105 specified by a particular context access instruction. For example, depending on the particular type of context access instruction, this may include: reading a context in a register, writing a context to a register, or both reading and writing a context from / to a register, or a combination of the above. In some embodiments, depending on the particular type of instruction, the operations may also optionally include other operations (e.g., adding, multiplying, or other arithmetic operations, shifting, rotating, or other logical operations, etc.). The scope of the present invention is not limited to any known type of context access instruction or these various possible types of operations.
[0095] The execution unit and / or processor may include specific or particular logic operative to execute context access instructions (e.g., transistors, integrated circuits, or other hardware potentially in combination with firmware (e.g., instructions stored in non-volatile memory) and / or software). In some cases, the execution unit may include an arithmetic unit, an arithmetic logic unit, or digital circuitry for performing arithmetic and logical operations, etc. In some embodiments, the execution unit may include one or more input structures (e.g., ports, interconnections, or interfaces) coupled to receive one or more source operands, circuitry or logic coupled to the one or more input structures to receive and process the one or more source operands and generate a result operand; and at least one output structure (e.g., port, interconnection, or interface) coupled to the circuitry or logic to output the result operand. In one example, the execution unit includes the (one or more) execution clusters 1760 shown in FIG. 17(B).
[0096] Referring again to Figure 11 , the retirement or other commit unit 1163 may be operative to commit context access instructions. Committing a context access instruction may include committing the result of the context access instruction to the architectural state.
[0097] In some embodiments, the commit unit may include an optional context metadata check unit 1164, an optional context metadata update unit 1165, and an optional lazy context element restoration unit 1110. Each of these units may be implemented using hardware, firmware, software, or a combination thereof (e.g., at least some circuitry potentially combined with some firmware). In some embodiments, the commit unit and / or the optional context metadata check unit may be operative to check corresponding context metadata 1107 for one or more context elements to be accessed. In some embodiments, the commit unit and / or the optional context metadata update unit may be operative to update corresponding context metadata for one or more context elements to be accessed. In some embodiments, the commit unit and / or the optional lazy context element restoration unit may be operative to lazily restore (i.e., operation 1166) the one or more context elements from system memory to the context storage device when the corresponding context metadata indicates that the one or more context elements have not been restored. In some embodiments, the commit unit and / or the optional lazy context element restoration unit may be operative to lazily perform one or more access control checks on one or more corresponding pages to be used for the one or more context elements when the corresponding context metadata indicates that the one or more context elements have not been restored. In some embodiments, depending on the particular type of context access instruction, the commit unit and its optional subunits may be operative to perform Figure 12-Figure 14 any one or more of the operations. Alternatively, such subunits and their operations need not be performed during the commit phase, but may be performed during the execution phase or another phase (e.g., if a mechanism is included for rolling back changes made for non-committed execution).
[0098] Figure 12-Figure 14 is a flowchart of embodiments of the method. In some embodiments, each of these methods may be performed as part of the execution of a context access instruction (e.g., an instruction that causes a register or other context to be read and / or written when executed), and / or in response to a context access instruction (e.g., an instruction that causes a register or other context to be read and / or written when executed). In embodiments, the method may be performed by a processor, digital logic, or an integrated circuit. In some embodiments, the method may be performed by Figure 1 processor 101 and / or Figure 11 processor 1101. The components, features, and particular optional details described herein for processor 101 and / or processor 1101 may also optionally apply to these methods. Alternatively, these methods may be performed by similar or different processors. Moreover, processor 101 and / or processor 1101 may perform the same, similar, or different methods as these methods.
[0099] Figure 12 It is a flowchart of an embodiment of method 1268 that lazily restores context elements and commits the operation of reading context elements. At block 1269, the operation of reading context elements is received at a retirement unit or other commit unit. The commit unit may be for performing a commit that submits the context to an architecturally visible context.
[0100] At block 1270, a determination is made as to whether the context metadata corresponding to the context element indicates that the context element has been read or written. For example, in some embodiments, this may include checking the corresponding read indication bit (e.g., R[i]) in a set of read indication bits (e.g., R[]), to see if the corresponding read indication bit has been cleared or otherwise indicates that the corresponding context element has not been read after entering the protected execution environment. Additionally, in some embodiments, this may include checking the corresponding write indication bit (e.g., W[i]) in a set of write indication bits (e.g., W[]), to see if the corresponding write indication bit has been cleared or otherwise indicates that the corresponding context element has not been written after entering the protected execution environment.
[0101] If the context metadata indicates that the context element has been read or written (i.e., the determination at block 1270 is "yes"), the method may proceed to block 1276. Since the context element has been read or written, the context element does not need to be restored from system memory and will generally have the correct and expected value. At block 1276, the operation may be retired or otherwise committed.
[0102] Conversely, if the context metadata indicates that the context element has not been read and has not been written (i.e., the determination at block 1270 is "no"), the method may proceed to block 1271. Since the context element has not been read and has not been written, the context element will typically have an incorrect value and should be removed from the pipeline. At block 1271, the pipeline of the operation may be cleared. For example, the pipeline may be blocked, nuked, etc.
[0103] At block 1272, an access control check may optionally be performed on the corresponding page or other portion of system memory that is used for the context element (e.g., for storing the context element in the case of an asynchronous exit and / or certain other conditions). As described above for Figure 6As discussed with respect to the box 634, the access control check for the corresponding page or other portion of the system memory used for the context element may optionally have been initially omitted or deferred. As discussed above, in some embodiments, the access control check may optionally be performed lazily and / or as needed and / or dynamically and / or when needed and / or selectively during the operation or execution of the protected software within the protected execution environment. In some embodiments, the operation of reading the context element may be a triggering event for initiating the access control check or causing the access control check to be performed. Since the operation of reading the context element arrives at box 1272 after determining that the context element has not been read and has not been written, the operation of reading the context element may represent the first reading of the context element after entering the protected execution environment. Performing the access control check at this point may mean performing the access control check in a timely manner before there is a possibility that the context may be lost in the case of an asynchronous exit and / or some other situation due to the lack of a suitable place to store the context element. The access control check may vary depending on the type of protected execution environment. Commonly, the access control check may perform one or more checks to check whether the page is available and can be used or to ensure that the page is available and can be used (e.g., the page exists, the page has the correct page type, the page has the appropriate ownership, the page has the correct access permissions (e.g., read and / or write and / or execute permissions), or any combination of the above). It should be appreciated that performing the access control check is optional and not mandatory. On the one hand, in other embodiments, the access control check need not be deferred, but may instead be performed as part of the Figure 6 method. On the other hand, some types of protected execution environments may not natively feature such access control checks.
[0104] At box 1273, the context metadata may be updated or otherwise changed to indicate that the context element has been read. For example, in some embodiments, this may include setting, updating, or otherwise changing the corresponding read indication bit in the set of read indication bits (e.g., R[]) (e.g., R[i]) such that it indicates that the corresponding context element has been read after entering the protected execution environment and during the execution within the protected execution environment.
[0105] At box 1274, the context element may be restored or otherwise loaded from the system memory. For example, a value may be loaded from the context save area in the memory into the corresponding register of the processor or other context storage location. At box 1275, the operation may be restarted. The restarted operation may now read the restored context and continue to proceed through box 1269, box 1270, and box 1276, where at box 1276, the restarted operation is committed.
[0106] Figure 13FIG. 1399 is a flow diagram of an embodiment of method 1399 for operations that commit write context elements. At block 1377, an operation to write a context element is received at a retirement unit or other commit unit. A commit unit may be for performing a commit that submits a context to an architecturally visible context.
[0107] At block 1378, an access control check may optionally be performed on a corresponding page or other portion of system memory that is used for the context element (e.g., for storing the context element in the case of an asynchronous exit and / or certain other conditions). As discussed above for Figure 6 block 634, the access control check on the corresponding page or other portion of system memory that is used for the context element may optionally have been initially omitted or deferred. As discussed above, in some embodiments, the access control check may optionally be performed lazily and / or as needed and / or dynamically and / or when needed and / or selectively during the operation or execution of protected software within a protected execution environment. In some embodiments, the operation of writing a context element may be a triggering event for initiating or causing an access control check to be performed. Since block 1378 is reached after determining that the context element has not been read and has not been written, the operation of writing the context element may represent the first write to the context element after entering the protected execution environment. Performing the access control check at this point may represent performing the access control check in a timely manner before there is a possibility that the context may be lost in the case of an asynchronous exit and / or certain other conditions due to not having a suitable place to store the context element. It is to be appreciated that performing the access control check is optional and not required. On the one hand, in other embodiments, the access control check need not be deferred, but may instead be performed as part of Figure 6 the method. On the other hand, some types of protected execution environments may not natively feature such access control checks.
[0108] At block 1379, the corresponding context metadata may be updated or otherwise changed to indicate that the context element has been written. For example, in some embodiments, this may include setting, updating, or otherwise changing a corresponding write indication bit in a set of write indication bits (e.g., W[]) (e.g., W[i]) such that it indicates that the corresponding context element has been written after entry into the protected execution environment and during execution within the protected execution environment. At block 1380, the operation of writing the context element may be retired or otherwise committed.
[0109] Figure 14It is a flowchart of an embodiment of method 1481 that lazily restores context elements and commits operations that read and then write context elements. At block 1482, an operation that reads and then writes a context element is received at a retirement unit or other commit unit.
[0110] At block 1483, a determination is made as to whether the context metadata corresponding to the context element indicates that the context element has been read or written. For example, in some embodiments, this may include checking a corresponding read indication bit (e.g., R[i]) in a set of read indication bits (e.g., R[]) to see if the corresponding read indication bit has been cleared or otherwise indicates that the corresponding context element has not been read after entering the protected execution environment. Additionally, in some embodiments, this may include checking a corresponding write indication bit (e.g., W[i]) in a set of write indication bits (e.g., W[]) to see if the corresponding write indication bit has been cleared or otherwise indicates that the corresponding context element has not been written after entering the protected execution environment.
[0111] If the context metadata indicates that the context element has been read or written (i.e., the determination at block 1483 is "yes"), the method may proceed to block 1489. Since the context element has been read or written, the context element does not need to be restored from system memory and will generally have the correct and expected value. At block 1489, the operation may be retired or otherwise committed.
[0112] Conversely, if the context metadata indicates that the context element has not been read and has not been written (i.e., the determination at block 1483 is "no"), the method may proceed to block 1484. Since the context element has not been read and has not been written, the context element will typically have an incorrect value and should be removed from the pipeline. At block 1484, the pipeline of the operation may be cleared. For example, the pipeline may be blocked, blown up, etc.
[0113] At block 1485, an access control check may optionally be performed on the corresponding page or other portion of system memory that is used for the context element (e.g., for storing the context element in the case of an asynchronous exit and / or certain other conditions). As described above for Figure 6As discussed with respect to the box 634, the access control check for the corresponding page or other portion of the system memory used for the context element may optionally have been initially omitted or delayed. As discussed above, in some embodiments, the access control check may optionally be performed lazily and / or as needed and / or dynamically and / or when needed and / or selectively during the operation or execution of the protected software within the protected execution environment. In some embodiments, the operation of reading the context element may be a triggering event for initiating or causing the access control check to be performed. Since the operation of reading the context element arrives at box 1485 after determining that the context element has not been read and has not been written, the operation of reading the context element may represent the first reading of the context element after entering the protected execution environment. Performing the access control check at this point may mean performing the access control check in a timely manner before there is a possibility of context loss in the case of an asynchronous exit and / or certain other conditions due to the lack of a suitable place to store the context element. It is to be appreciated that performing the access control check is optional and not mandatory. On the one hand, in other embodiments, the access control check need not be delayed, but may instead be performed as part of the Figure 6 method. On the other hand, some types of protected execution environments may not natively feature such access control checks.
[0114] At box 1486, the context metadata may be updated or otherwise changed to indicate that the context element has been read and written. For example, in some embodiments, this may include setting, updating, or otherwise changing the corresponding read indication bit (e.g., R[i]) in the set of read indication bits (e.g., R[]) such that it indicates that the corresponding context element has been read after entry into the protected execution environment. Also, in some embodiments, this may include setting, updating, or otherwise changing the corresponding write indication bit (e.g., W[i]) in the set of write indication bits (e.g., W[]) such that it indicates that the corresponding context element has been written after entry into the protected execution environment and during the execution within the protected execution environment.
[0115] At box 1487, the context element may be restored or otherwise loaded from the system memory. For example, a value may be loaded from the context save area in the memory into the corresponding register or other context storage location of the processor. At box 1488, the operation may be restarted. The restarted operation may now read the restored context and continue to proceed through box 1482, box 1483, and box 1489, where at box 1489, the restarted operation may be committed.
[0116] Figure 12-Figure 14This is an example embodiment of a method that can be executed. The method has been described in a basic form in the flowchart, but operations can optionally be added to and / or removed from these methods. For example, optional operations can be removed. Additionally, although the flowchart depicts a particular order of operations according to an embodiment, this order is exemplary. Alternative embodiments can perform operations in a different order, combine certain operations, overlap certain operations, etc. For example, the order of blocks 1273 and 1274 (or 1486 and 1487) can be swapped, the order of blocks 1272 and 1273 (or 1485 and 1486) can be swapped, or a combination of the above can be done.
[0117] In some embodiments, the techniques disclosed herein can optionally be applied to all contexts of a processor (e.g., regardless of the number of such contexts, the degree to which such contexts are likely to be modified during execution within a protected execution environment, etc.). In other embodiments, the techniques disclosed herein can optionally be selectively applied only to a subset of all contexts of a processor. For example, the techniques disclosed herein can optionally be selectively applied to a subset of contexts having a relatively large amount of data (e.g., vector registers and on-chip storage), but not to another subset of contexts having a relatively small amount of data (e.g., instruction pointer, general-purpose registers). As another example, the techniques disclosed herein can optionally be selectively applied to a subset of contexts that are more likely to be unused (e.g., the widest set of vector registers, on-chip storage, accelerator contexts, etc.), but not to another subset of contexts that are typically always in use (e.g., instruction pointer, general-purpose registers, etc.). Moreover, embodiments can use heuristics to determine which contexts are frequently used by a logical processor within a protected execution environment. When one or more context elements are determined to be frequently used, embodiments can choose to eagerly restore those context elements rather than lazily restore them, which would also require setting the corresponding bits in R[] during entry into the protected execution environment. Also, in some embodiments, the protected execution environment can have additional control over which contexts are selectively saved (e.g., control over skipping saving some contexts). For example, SGX and TDX can cause saving of some contexts to be skipped.
[0118] To further illustrate certain concepts, a particular detailed example embodiment of how Intel SGX can be modified to include the features of the embodiments disclosed herein will now be described. It is to be appreciated that this is merely a particular detailed example embodiment. Even when implementing the features disclosed herein into SGX, the scope of the invention is not so limited. Many variations are possible and many variations are contemplated.
[0119] Intel SGX is an instruction set architecture (ISA) extension that enables the creation and attestation of secure enclaves, which can represent regions of user code and data that are protected from privileged software. A logical processor can enter an enclave by calling the EENTER instruction for synchronous entry or the ERESUME instruction for asynchronous entry (e.g., for resuming operation after an interruption). A logical processor can exit an enclave by calling the EEXIT instruction for synchronous exit, or a logical processor can experience an asynchronous enclave exit (AEX) when it is interrupted or encounters an exception. All enabled processor contexts can be restored and saved via ERESUME and AEX operations, respectively. The Intel SGX architecture defines a State Save Area (SSA), which is used by each enclave software thread to save / restore the processor context when asynchronously exiting / entering an enclave (e.g., during AEX or ERESUME, respectively). The SSA is an example of a context save area.
[0120] The EENTER and ERESUME operations can perform an enclave page cache map (EPCM) check on all State Save Area (SSA) pages that can be used to securely store the processor context. Intel's first-generation Advanced Matrix Extensions (AMX) support two palettes: in palette 0, the logical processor does not read / write to the AMX tiles, and these AMX tiles will also not be saved / restored via XSAVE / XRSTOR; in palette 1, the logical processor can fully utilize all tiles, and all tiles will be saved / restored. Even if AMX is enabled in palette 0, EENTER and ERESUME may not assume that an enclave thread starting in palette 0 will not switch to palette 1, and if the thread does switch to palette 1, AEX will generally save all tiles. Therefore, EENTER and ERESUME can perform an additional EPCM check on the SSA pages that would be used to save the tiles, thereby incurring overhead even when the tiles will not be used. Additionally, ERESUME and AEX can also save and restore all tiles again, even when only a subset of these tiles (or just these tiles) are in use. Such access control checks on the SSA pages can be omitted using the methods disclosed herein.
[0121] In some embodiments, the processor may be modified to include a read indication bit set R[] and a write indication bit set W[]. Each bit may correspond to a register, a slice, or other context element. The ERESUME instruction may be modified to clear R[] and W[]. The EENTER instruction does not restore the processor state from the SSA, so it may not clear R[] and W[]. Alternatively, the EENTER instruction may set all bits in R[] and W[] to prevent the processor context from being lazily restored after the logical processor enters via EENTER, and to force all extended processor contexts to be saved when the logical processor exits via a subsequent AEX. The ERESUME and EENTER instructions may also be modified to omit the EPCM check for SSA pages (e.g., the SSA page for storing AMX state) whose contents will only be lazily read and written. Generally, the EEXIT instruction does not save the extended processor state to the SSA, so the EEXIT instruction may not need to be modified. The AEX may be modified to selectively save only those context elements when W[] indicates that a write has been made to the context element, and to selectively purge those context elements when W[] indicates that a write has been made to the context element or R[] indicates that a read has been made from the context element. For example, the operational behavior of the AEX may be modified to store the XSAVE state into the XSAVE region of the current SSA frame using the physical address determined and cached at enclave entry with CR_XSAVE_PAGE_i. For each XSAVE state i defined by (SECS.ATTRIBUTES.XFRM[i] = 1, the destination address cached in CR_XSAVE_PAGE_n)
[0122] In some embodiments, changes may also be made to the behavior of the commit pipeline stage. When a logical processor is executing within an enclave and a commit operation reads some context, the processor may perform the same as for Figure 12Operations similar to those shown. If the corresponding bits in R[] and W[] are both cleared, the processor may stall the pipeline and load the context element from the SSA frame. For example, it may trigger a microcode assist to locate and load the context element, in which the microcode may use existing control registers (e.g., CR_XSAVE_PAGE_n) to identify the base address of the context (XSAVE) region in the current SSA frame of the logical processor. For example, the microcode may then derive an index into the XSAVE region where the saved state of the context element is located and then load that value into the context element of the logical processor. Depending on the extended features defined by the processor, etc., the index derivation function may vary with embodiments. When a logical processor executes in an enclave and commits an operation to write to a context element, the processor may set the corresponding bit in W[], similar to that described for Figure 13 When a logical processor executes in an enclave and commits an operation to modify a context element, the processor may perform operations similar to those shown for Figure 14 For example, the microcode assist may load the context element when the corresponding bits in R[] and W[] are not set.
[0123] In some embodiments, in one or more scenarios where an SSA page is to be read or written lazily and no EPCM check has been performed on the page since the logical processor entered the enclave, the microcode assist may perform the corresponding EPCM check. This new behavior may have architectural side effects. Specifically, instructions that would not otherwise access memory may trigger an error in the event that one of these EPCM checks fails. Since the OS (e.g., the SGX driver) may load the missing SSA page without introspecting on the instruction that triggered the error, no significant OS or platform software changes should be required.
[0124] In some embodiments, Intel SGX may also include another architectural extension. Different from some other TEEs, SGX allows the SSA to be both software-readable and software-writable. A possible scenario is as follows. A logical processor enters the enclave asynchronously via ERESUME. The modified ERESUME behavior clears R[] and W[], and does not restore the processor's context from the SSA. The logical processor overwrites SSA.XMM1 (e.g., the state save field for the XMM1 register) with a new value V' different from the original value V. The logical processor executes an instruction that moves data from XMM1 to XMM2. When the instruction is committed, the processor observes that both R[XMM1] and W[XMM1] are cleared, so it stalls the pipeline and loads XMM1 from the SSA, and then re-executes the vector move operation that triggered the stall. After this operation, both XMM1 and XMM2 contain the new value V'. However, since the SSA was modified after ERESUME, the Intel SGX architecture expects both XMM1 and XMM2 to contain V. In some embodiments, SGX may be modified or extended to allow the XSAVE state to be maintained in a read-only page. For example, Intel's Control-flow Enforcement Technology (CET) can be used with SGX. A new read-only SSA, called CET SSA, can be used.
[0125] To further illustrate certain concepts, a specific detailed example embodiment of how Intel TDX can be modified to include the features of the embodiments disclosed herein will now be described. It is to be appreciated that this is merely a specific detailed example embodiment. Even when implementing the features disclosed herein into TDX, the scope of the present invention is not so limited. Many variations are possible and many variations are contemplated.
[0126] Intel TDX is an ISA extension that enables the creation and attestation of trust domains (TDs). A TD can represent a virtual machine (VM) that contains user code and data and is protected from an untrusted Virtual Machine Monitor (VMM) and other VMs. A logical processor can enter a TD through an intermediate Intel TDX module that acts as a trusted VMM. The Intel TDX module executes in a specialized TDX Root Secure-Arbitration Mode (SEAM) of operation. When an untrusted VMM wants a logical processor to enter a TD (e.g., start / resume execution of the TD's virtual CPU, i.e., VCPU), the untrusted VMM can call the SEAMCALL[TDH.VP.ENTER] instruction to enter the Intel TDX module. This instruction can restore the TD VCPU context and determine whether the TD should be entered synchronously or asynchronously (e.g., to resume execution after an interruption or exception). Similarly, a logical processor can exit a TD through the Intel TDX module. A logical processor can perform a synchronous exit by calling the TDCALL[TDG.VP.VMCALL] instruction, or an asynchronous exit can occur if the logical processor is interrupted or an exception occurs. In both cases, the logical processor can transfer control from the VCPU of the TD to the Intel TDX module. The Intel TDX module can save the context of the TD's VCPU and call the SEAMRET instruction to transfer control to the untrusted VMM.
[0127] The Trust Domain Virtual Processor State (TDVPS) is used by each VCPU as a context save area to save / restore the VCPU processor context when exiting / entering a TD. The TDVPS structure is similar to the virtual machine control structure (VMCS) structure in the Intel virtual machine extension (VMX) architecture and includes the VCPU architecture context. The VCPU processor context can be saved and restored not only during asynchronous exit / entry of the trust domain but also during synchronous exit / entry.
[0128] In some embodiments, the processor may be modified to include a read indication bit set R[] and a write indication bit set W[]. Each bit may correspond to a register, slice, or other context element. The TDX module may be modified to clear R[] and W[] when the SEAMCALL[TDH.VP.ENTER] instruction is invoked. Additionally, when returning from a VMCALL, the TDX ABI may disclose an option to mask some registers to prevent their restoration from the TDVPS structure. For each bit set in such a mask, SEAMCALL[TDH.VP.ENTER] may set the corresponding bits in R[] and W[] to prevent such state from being overwritten, which could otherwise violate the expected architectural behavior of TDX. VM exits that occur in the TDX non-root SEAM mode (e.g., when control is passed from TD to the Intel TDX module) may be modified to selectively save context elements only if the corresponding bit in W[] is set, and to selectively sanitize context elements when the corresponding bit in W[] or R[] is set.
[0129] In some embodiments, changes may also be made to the behavior of the commit pipeline stage. The TDVPS structure has a field called XBUFF that contains XSAVE data for the current VP. The TDX microarchitecture may be extended to have additional control registers to cache the physical address of such an XBUFF, which can then be used by the microcode to locate the saved values of the accessed context elements (e.g., similar to the role of CR_XSAVE_PAGE_n in the SGX description above). Alternatively, the retirement stage may trigger a VMEXIT after flushing the pipeline. The VMEXIT may transition into the SEAM, which can then locate the context element save state and load the context element save state from the XBUFF into the context elements of the logical processor.
[0130] Example computer architecture.
[0131] The following describes an example computer architecture in detail. Other system designs and configurations known in the art for laptops, desktop computers, handheld personal computers (PCs), personal digital assistants, engineering workstations, servers, blade servers, network devices, network hubs, switches, routers, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and various other electronic devices are also suitable. In general, various systems or electronic devices capable of incorporating the processors and / or other execution logic disclosed herein are generally suitable.
[0132] Figure 15 An example computing system is illustrated. The multiprocessor system 1500 is an interface-based system and includes multiple processors or cores, including a first processor 1570 and a second processor 1580 coupled via an interface 1550 (e.g., a point-to-point (P-P) interconnect, a fabric, and / or a bus). In some examples, the first processor 1570 and the second processor 1580 are homogeneous. In some examples, the first processor 1570 and the second processor 1580 are heterogeneous. Although the example system 1500 is shown as having two processors, the system can have three or more processors, or can be a single-processor system. In some embodiments, the computing system is a system-on-chip (SoC).
[0133] The processors 1570 and 1580 are shown as including integrated memory controller (IMC) circuits 1572 and 1582, respectively. The processor 1570 also includes interface circuits 1576 and 1578; similarly, the second processor 1580 includes interface circuits 1586 and 1588. The processors 1570, 1580 can exchange information via the interface circuits 1578, 1588 through the interface 1550. The IMCs 1572 and 1582 couple the processors 1570, 1580 to their respective memories, namely memory 1532 and memory 1534, which can be part of the main memories locally attached to the respective processors.
[0134] The processors 1570, 1580 can each utilize interface circuits 1576, 1594, 1586, 1598 to exchange information with a network interface (NW I / F) 1590 via respective interfaces 1552, 1554. The network interface 1590 (such as one or more of an interconnect, a bus, and / or a fabric, and in some examples, a chipset) can optionally exchange information with the coprocessor 1538 via the interface circuit 1592. In some examples, the coprocessor 1538 is a dedicated processor, such as a high-throughput processor, a network or communication processor, a compression engine, a graphics processor, a general purpose graphics processing unit (GPGPU), a neural-network processing unit (NPU), an embedded processor, and so on.
[0135] A shared cache (not shown) can be included in either of the processors 1570, 1580, or outside both processors but connected to these processors via an interface (such as a P-P interconnect), such that: if a processor is placed in a low-power mode, the local cache information of either or both processors can also be stored in the shared cache.
[0136] The network interface 1590 can be coupled to a first interface 1516 via the interface circuit 1596. In some examples, the first interface 1516 can be an interface such as a Peripheral Component Interconnect (PCI) interconnect, a PCI Express interconnect, or another I / O interconnect. In some examples, the first interface 1516 is coupled to a power control unit (PCU) 1517, and the PCU 1517 can include circuitry, software, and / or firmware to perform power management operations regarding the processors 1570, 1580, and / or the coprocessor 1538. The PCU 1517 provides control information to a voltage regulator (not shown) such that the voltage regulator generates an appropriate regulated voltage. The PCU 1517 also provides control information to control the generated operating voltage. In various examples, the PCU 1517 can include various power management logic units (circuits) to perform hardware-based power management. Such power management can be completely controlled by the processor (e.g., controlled by various processor hardware and can be triggered by workload and / or power constraints, thermal constraints, or other processor constraints), and / or the power management can be performed in response to an external source (e.g., a platform or a power management source or system software).
[0137] The PCU 1517 is shown as existing as separate logic from the processor 1570 and / or the processor 1580. In other cases, the PCU 1517 may execute on one or more cores (not shown) given in the core of the processor 1570 or 1580. In some cases, the PCU 1517 may be implemented as a microcontroller (dedicated or general-purpose) or other control logic, which is configured to execute its own dedicated power management code (sometimes called P-code). In still other examples, the power management operations to be performed by the PCU 1517 may be implemented outside the processor, for example, by a separate power management integrated circuit (PMIC) or another component outside the processor. In still other examples, the power management operations to be performed by the PCU 1517 may be implemented within the BIOS or other system software.
[0138] A variety of I / O devices 1514 and the bus bridge 1518 may be coupled to the first interface 1516, and the bus bridge 1518 couples the first interface 1516 to the second interface 1520. In some examples, one or more additional processors 1515 are coupled to the first interface 1516, such as a coprocessor, a high-throughput many integrated core (MIC) processor, a GPGPU, an accelerator (such as a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array (FPGA), or any other processor. In some examples, the second interface 1520 may be a low pin count (LPC) interface. A variety of devices may be coupled to the second interface 1520, and these devices include, for example, a keyboard and / or a mouse 1522, a communication device 1527, and a storage circuit 1528. The storage circuit 1528 may be one or more non-transitory machine-readable storage media described below, such as a disk drive or other mass storage device, which may include instructions / code and data 1530 in some examples and may implement the storage device `ISAB03. Additionally, the audio I / O 1524 may be coupled to the second interface 1520. Note that other architectures are possible in addition to the point-to-point architecture described above. For example, a system such as the multiprocessor system 1500 may implement a multi-drop interface or other such architecture instead of the point-to-point architecture.
[0139] Example core architectures, processors, and computer architectures.
[0140] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, the implementations of these cores can include: 1) general-purpose in-order cores, for general computing purposes; 2) high-performance general-purpose out-of-order cores, for general computing purposes; 3) specialized cores, mainly for graphics and / or scientific (throughput) computing purposes. The implementations of different processors can include: 1) a CPU, including one or more general-purpose in-order cores for general computing purposes and / or one or more general-purpose out-of-order cores for general computing purposes; and 2) a coprocessor, including one or more specialized cores mainly for graphics and / or scientific (throughput) computing purposes. These different processors result in different computer system architectures, which can include: 1) the coprocessor and the CPU on separate chips; 2) the coprocessor and the CPU on separate dies within the same package; 3) the coprocessor and the CPU on the same die (in this case, such a coprocessor is sometimes referred to as specialized logic, such as integrated graphics and / or scientific (throughput) logic, or as a specialized core); and 4) a system-on-chip (SoC), which can include the described CPU (sometimes referred to as one or more application cores or one or more application processors), the above-mentioned coprocessor, and additional functions on the same die. Example core architectures are described next, followed by a description of example processors and computer architectures.
[0141] Figure 16 FIG. illustrates a block diagram of an example processor and / or SoC 1600, which can have one or more cores and have an integrated memory controller. The processor 1600 illustrated by the solid-line block diagram has a single core 1602(A), a system agent unit circuit 1610, and a set of one or more interface controller unit circuits 1616, while the optionally added dashed-line block diagram illustrates an alternative processor 1600 as having multiple cores 1602(A)-(N), a set of one or more integrated memory control unit circuits 1614 in the system agent unit circuit 1610, specialized logic 1608, and a set of one or more interface controller unit circuits 1616. Note that the processor 1600 can be Figure 15 one of the processors 1570 or 1580 or the coprocessors 1538 or 1515.
[0142] Thus, different implementations of the processor 1600 may include: 1) a CPU, where the dedicated logic 1608 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores, not shown), and the cores 1602(A)-(N) are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of both); 2) a coprocessor, where the cores 1602(A)-(N) are a large number of dedicated cores mainly for graphics and / or scientific (throughput) purposes; and 3) a coprocessor, where the cores 1602(A)-(N) are a large number of general-purpose in-order cores. Thus, the processor 1600 may be a general-purpose processor, a coprocessor, or a special-purpose processor, such as a network or communication processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput integrated many-core (MIC) coprocessor (including 30 or more cores), an embedded processor, and so on. The processor may be implemented on one or more chips. The processor 1600 may be part of one or more substrates and / or may be implemented on one or more substrates using any of a variety of process technologies, such as complementary metal oxide semiconductor (CMOS), bipolar CMOS (BiCMOS), P-type metal oxide semiconductor (PMOS), or N-type metal oxide semiconductor (NMOS).
[0143] The memory hierarchy includes one or more levels of cache unit circuits 1604(A)-(N) within cores 1602(A)-(N), a group of one or more shared cache unit circuits 1606, and an external memory (not shown) coupled to the group of integrated memory controller unit circuits 1614. The group of one or more shared cache unit circuits 1606 may include one or more intermediate-level caches, such as a second level (L2), third level (L3), fourth level (L4), or other levels of cache, such as a last level cache (LLC), and / or combinations thereof. Although in some examples the interface network circuit 1612 (e.g., a ring interconnect) provides an interface to dedicated logic 1608 (e.g., integrated graphics logic), the group of shared cache unit circuits 1606, and the system agent unit circuit 1610, alternative examples use any number of well-known techniques to provide an interface to these units. In some examples, coherence is maintained between one or more of the circuits in the shared cache unit circuits 1606 and the cores 1602(A)-(N). In some examples, the interface controller unit circuit 1616 couples these cores 1602 to one or more other devices 1618, such as one or more I / O devices, storage devices, one or more communication devices (e.g., wireless networks, wired networks, etc.), and so on.
[0144] In some examples, one or more of the cores 1602(A)-(N) have multithreading capabilities. The system agent unit circuit 1610 includes those components that coordinate and operate the cores 1602(A)-(N). The system agent unit circuit 1610 may include, for example, a power control unit (PCU) circuit and / or a display unit circuit (not shown). The PCU may be (or may include) the logic and components required to regulate the power states of the cores 1602(A)-(N) and / or the dedicated logic 1608 (e.g., integrated graphics logic). The display unit circuit is used to drive one or more externally connected displays.
[0145] The cores 1602(A)-(N) may be homogeneous in terms of the instruction set architecture (ISA). Alternatively, the cores 1602(A)-(N) may be heterogeneous in terms of the ISA; that is, a subset of the cores 1602(A)-(N) may be capable of executing one ISA, while other cores may be capable of executing only a subset of that ISA or may be capable of executing another ISA.
[0146] Example core architecture - in-order and out-of-order core block diagrams.
[0147] The block diagram of FIG. 17(A) illustrates an example in-order pipeline and an example register renaming, out-of-order issue / execution pipeline both according to some examples. The block diagram of FIG. 17(B) illustrates an example in-order architecture core and an example register renaming, out-of-order issue / execution architecture core both to be included in a processor according to examples. The solid boxes in FIGS. 17(A)-(B) illustrate the in-order pipeline and in-order core, while the optionally added dashed boxes illustrate the register renaming, out-of-order issue / execution pipeline and core. Considering that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.
[0148] In FIG. 17(A), the processor pipeline 1700 includes a fetch stage 1702, an optional length decoding stage 1704, a decode stage 1706, an optional allocation (Alloc) stage 1708, an optional rename stage 1710, a schedule (also referred to as dispatch or issue) stage 1712, an optional register read / memory read stage 1714, an execution stage 1716, a write-back / memory write stage 1718, an optional exception handling stage 1722, and an optional commit stage 1724. One or more operations may be performed in each of these processor pipeline stages. For example, during the fetch stage 1702, one or more instructions are fetched from an instruction memory, and during the decode stage 1706, the fetched one or more instructions may be decoded, an address using a forwarding register port (e.g., a load store unit (LSU) address) may be generated, and branch forwarding (e.g., immediate offset or link register (LR)) may be performed. In one example, the decode stage 1706 and the register read / memory read stage 1714 may be combined into one pipeline stage. In one example, during the execution stage 1716, the decoded instructions may be executed, LSU address / data pipelining to an Advanced Microcontroller Bus (AMB) interface may be performed, multiplication and addition operations may be performed, arithmetic operations with branch results may be performed, and so on.
[0149] As an example, the exemplary register renaming, out-of-order issue / execution architecture core of FIG. 17(B) may implement pipeline 1700 in the following manner: 1) Instruction fetch circuitry 1738 performs fetch and length decoding stages 1702 and 1704; 2) Decoding circuitry 1740 performs decoding stage 1706; 3) Rename / allocator unit circuitry 1752 performs allocation stage 1708 and rename stage 1710; 4) (One or more) scheduler circuitry 1756 performs scheduling stage 1712; 5) (One or more) physical register file circuitry 1758 and memory unit circuitry 1770 perform register read / memory read stage 1714; (One or more) execution clusters 1760 perform execution stage 1716; 6) Memory unit circuitry 1770 and (one or more) physical register file circuitry 1758 perform writeback / memory write stage 1718; 7) Various circuitry may be involved in exception handling stage 1722; and 8) Retirement unit circuitry 1754 and (one or more) physical register file circuitry 1758 perform commit stage 1724.
[0150] FIG. 17(B) shows that the processor core 1790 includes a front-end unit circuitry 1730 coupled to an execution engine unit circuitry 1750, and both are coupled to a memory unit circuitry 1770. The core 1790 may be a reduced instruction set architecture computing (RISC) core, a complex instruction set architecture computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, the core 1790 may be a dedicated core, such as a network or communication core, a compression engine, a coprocessor core, a general purpose computing graphics processing unit (GPGPU) core, a graphics core, and so on.
[0151] The front-end unit circuit 1730 may include a branch prediction circuit 1732, which is coupled to an instruction cache circuit 1734, which is coupled to an instruction translation lookaside buffer (TLB) 1736, which is coupled to an instruction fetch circuit 1738, which is coupled to a decoding circuit 1740. In one example, the instruction cache circuit 1734 is included in the memory unit circuit 1770 rather than in the front-end circuit 1730. The decoding circuit 1740 (or decoder) may decode the instruction and generate one or more micro-operations, microcode entry points, micro-instructions, other instructions, or other control signals as outputs, which are decoded from the original instruction, or otherwise reflect the original instruction, or are derived from the original instruction. The decoding circuit 1740 may also include an address generation unit (AGU, not shown) circuit. In one example, the AGU uses a forwarded register port to generate an LSU address and may further perform branch forwarding (e.g., immediate offset branch forwarding, LR register branch forwarding, etc.). Various different mechanisms may be utilized to implement the decoding circuit 1740. Examples of suitable mechanisms include, but are not limited to, look-up tables, hardware implementations, programmable logic arrays (PLAs), microcode read only memories (ROMs), etc. In one example, the core 1790 includes a microcode ROM (not shown) or other medium that stores microcode for certain macro instructions (e.g., in the decoding circuit 1740 or otherwise within the front-end circuit 1730). In one example, the decoding circuit 1740 includes a micro-operation (micro-op) or operation cache (not shown) to save / cache decoded operations, micro-tokens, or micro-operations generated during the decoding or other stages of the processor pipeline 1700. The decoding circuit 1740 may be coupled to a rename / allocator unit circuit 1752 in the execution engine circuit 1750.
[0152] The execution engine circuit 1750 includes a rename / allocator unit circuit 1752, which is coupled to a retirement unit circuit 1754 and a set of one or more scheduler circuits 1756. The scheduler circuits 1756 represent any number of different schedulers, including reservation stations, a central instruction window, and so on. In some examples, the (one or more) scheduler circuits 1756 may include an arithmetic logic unit (ALU) scheduler / scheduling circuit, an ALU queue, an address generation unit (AGU) scheduler / scheduling circuit, an AGU queue, and so on. The (one or more) scheduler circuits 1756 are coupled to the (one or more) physical register file circuits 1758. Each of the (one or more) physical register file circuits 1758 represents one or more physical register files, and different physical register files among these physical register files store one or more different data types, such as scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points, status (e.g., an instruction pointer, i.e., the address of the next instruction to be executed), and so on. In one example, the (one or more) physical register file circuits 1758 include a vector register unit circuit, a write mask register unit circuit, and a scalar register unit circuit. These register units may provide architected vector registers, vector mask registers, general-purpose registers, and so on. The (one or more) physical register file circuits 1758 are coupled to the retirement unit circuit 1754 (also referred to as a retirement queue) to illustrate various ways that can be used to implement register renaming and out-of-order execution (e.g., using the (one or more) reorder buffer(s) (ROB) and the (one or more) retirement register files; using the (one or more) future heaps, the (one or more) history buffers, and the (one or more) retirement register files; using a register map and a pool of registers; and so on). The retirement unit circuit 1754 and the (one or more) physical register file circuits 1758 are coupled to the (one or more) execution clusters 1760. The (one or more) execution clusters 1760 include a set of one or more execution unit circuits 1762 and a set of one or more memory access circuits 1764. The (one or more) execution unit circuits 1762 may perform various arithmetic, logical, floating-point, or other types of operations (e.g., shift, addition, subtraction, multiplication) on various types of data (e.g., scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points). Although some examples may include several execution units or execution unit circuits dedicated to a specific function or set of functions, other examples may include only one execution unit circuit or multiple execution units / execution unit circuits that perform all functions.(One or more) scheduler circuits 1756, (one or more) physical register file circuits 1758, and (one or more) execution clusters 1760 are shown as potentially being plural because some examples create separate pipelines for certain types of data / operations (e.g., scalar integer pipelines, scalar floating point / tight integer / tight floating point / vector integer / vector floating point pipelines, and / or memory access pipelines, each having its own scheduler circuit, (one or more) physical register file circuits, and / or execution cluster - and in the case of a separate memory access pipeline, in some examples only the execution cluster of that pipeline has (one or more) memory access unit circuits 1764). It should also be understood that in cases where separate pipelines are used, one or more of these pipelines can be out-of-order issue / execution while the rest are in-order.
[0153] In some examples, the execution engine unit circuit 1750 may perform load / store unit (LSU) address / data pipelining to an advanced microcontroller bus (AMB) interface (not shown), as well as address phase and writeback, data phase load, store, and branch.
[0154] A set of memory access circuits 1764 is coupled to a memory unit circuit 1770, which includes a data TLB circuit 1772, which is coupled to a data cache circuit 1774, which is coupled to a level 2 (L2) cache circuit 1776. In one example, the memory access circuits 1764 may include a load unit circuit, a store address unit circuit, and a store data unit circuit, each of which is coupled to the data TLB circuit 1772 in the memory unit circuit 1770. An instruction cache circuit 1734 is further coupled to the level 2 (L2) cache circuit 1776 in the memory unit circuit 1770. In one example, the instruction cache 1734 and the data cache 1774 are combined into a single instruction and data cache (not shown) in the L2 cache circuit 1776, a level 3 (L3) cache circuit (not shown), and / or main memory. The L2 cache circuit 1776 is coupled to one or more other levels of cache and ultimately to main memory.
[0155] The core 1790 may support one or more instruction sets (e.g., x86 instruction set architecture (optionally with some extensions added with updated versions); MIPS instruction set architecture; ARM instruction set architecture (optionally with optional additional extensions, e.g., NEON)), which includes the (one or more) instructions described herein. In one example, the core 1790 includes logic to support a packed data instruction set architecture extension (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using packed data.
[0156] (One or more) example execution unit circuits.
[0157] Figure 18 An example of (one or more) execution unit circuits is illustrated, such as the (one or more) execution unit circuits 1762 of FIG. 17(B). As shown, the (one or more) execution unit circuits 1762 may include one or more ALU circuits 1801, an optional vector / single instruction multiple data (SIMD) circuit 1803, a load / store circuit 1805, a branch / jump circuit 1807, and / or a floating-point unit (FPU) circuit 1809. The ALU circuit 1801 performs integer arithmetic and / or Boolean operations. The vector / SIMD circuit 1803 performs vector / SIMD operations on packed data (e.g., SIMD / vector registers). The load / store circuit 1805 executes load and store instructions to load data from memory into registers or store data from registers to memory. The load / store circuit 1805 may also generate addresses. The branch / jump circuit 1807 causes a branch or jump to a certain memory address depending on the instruction. The FPU circuit 1809 performs floating-point arithmetic. The width of the (one or more) execution unit circuits 1762 varies depending on the example and may be in a range, for example, from 16 bits to 1024 bits. In some examples, two or more smaller execution units are logically combined to form a larger execution unit (e.g., two 128-bit execution units are logically combined to form a 256-bit execution unit).
[0158] Example register architecture.
[0159] Figure 19is a block diagram of a register architecture 1900 according to some examples. As shown, the register architecture 1900 includes vector / SIMD registers 1910, whose widths vary from 128 bits to 1024 bits. In some examples, the vector / SIMD registers 1910 are physically 512 bits, and depending on the mapping, only some of the lower bits are used. For example, in some examples, the vector / SIMD registers 1910 are 512-bit ZMM registers: the lower 256 bits are used for YMM registers, and the lower 128 bits are used for XMM registers. Thus, there is register overlap. In some examples, the vector length field selects between a maximum length and one or more other shorter lengths, where each such shorter length is half the length of the previous length. A scalar operation is an operation performed on the lowest-order data element position in a ZMM / YMM / XMM register; the higher-order data element positions are either kept the same as they were before the instruction or are zeroed, depending on the example.
[0160] In some examples, the register architecture 1900 includes write mask / predicate registers 1915. For example, in some examples, there are 8 write mask / predicate registers (sometimes called k0 to k7), each of which is 16 bits, 32 bits, 64 bits, or 128 bits in size. The write mask / predicate registers 1915 can allow merging (e.g., allowing any set of elements in the destination to be protected from update during the execution of any operation) and / or zeroing (e.g., zeroing the vector mask allows any set of elements in the destination to be zeroed during the execution of any operation). In some examples, each data element position in a given write mask / predicate register 1915 corresponds to a data element position in the destination. In other examples, the write mask / predicate registers 1915 are scalable and consist of a set number of enable bits for a given vector element (e.g., 8 enable bits for each 64-bit vector element).
[0161] The register architecture 1900 includes multiple general-purpose registers 1925. These registers can be 16 bits, 32 bits, 64 bits, etc., and are capable of being used for scalar operations. In some examples, these registers are named RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 to R15.
[0162] In some examples, the register architecture 1900 includes a scalar floating-point (FP) register file 1945 that is used to perform scalar floating-point operations on 32 / 64 / 80-bit floating-point data using the x87 instruction set architecture extensions, or as MMX registers to perform operations on 64-bit packed integer data, and to hold operands for some operations performed between MMX and XMM registers.
[0163] One or more flag registers 1940 (e.g., EFLAGS, RFLAGS, etc.) store status and control information for arithmetic, comparison, and system operations. For example, one or more flag registers 1940 can store condition code information such as carry, parity, auxiliary carry, zero, sign, and overflow. In some examples, one or more flag registers 1940 are referred to as program status and control registers.
[0164] Segment registers 1920 contain segment pointers for accessing memory. In some examples, these registers are named CS, DS, SS, ES, FS, and GS.
[0165] Machine-specific registers (MSRs) 1935 control and report on processor performance. Most MSRs 1935 handle system-related functions and are not accessible to applications. Machine check registers 1960 consist of control, status, and error reporting MSRs for detecting and reporting hardware errors.
[0166] One or more instruction pointer registers 1930 store instruction pointer values. One or more control registers 1955 (e.g., CR0 - CR4) determine the operating mode of the processor (e.g., processors 1570, 1580, 1538, 1515, and / or 1600) and the characteristics of the currently executing task. Debug registers 1950 control and allow monitoring of the debug operations of the processor or core.
[0167] Memory (mem) management registers 1965 specify the locations of data structures used in protected mode memory management. These registers can include the global descriptor table register (GDTR), the interrupt descriptor table register (IDTR), the task register, and the local descriptor table register (LDTR).
[0168] Alternative examples may use wider or narrower registers. Additionally, alternative examples may use more, fewer, or different register files and registers. The register architecture 1900 may be used, for example, in the register file / memory ISA B08, or in the physical register file circuitry 1758.
[0169] Instruction Set Architecture.
[0170] An instruction set architecture (ISA) may include one or more instruction formats. A given instruction format may define various fields (e.g., number of bits, position of bits) to specify the operation to be performed (e.g., opcode) and the operand(s) on which the operation is to be performed and / or other data field(s) (e.g., mask), etc. Some instruction formats are further decomposed by the definition of instruction templates (or sub-formats). For example, an instruction template of a given instruction format may be defined as having different subsets of the fields of that instruction format (the fields included are typically in the same order, but at least some have different bit positions since fewer fields are included) and / or as having a given field interpreted in a different way. Thus, each instruction of the ISA is expressed using a given instruction format (and if defined, using a given instruction template in the instruction template of that instruction format), and includes fields for specifying the operation and operands. For example, an example ADD instruction has a specific opcode and an instruction format that includes an opcode field to specify the opcode and an operand field to select the operands (source1 / destination and source2); and the occurrence of this ADD instruction in the instruction stream will have specific contents in the operand fields that select the specific operands. Additionally, although the following description is in the context of the x86 ISA, applying the teachings of the present disclosure to other ISAs is within the knowledge of those skilled in the art.
[0171] Example instruction formats.
[0172] Examples of the (one or more) instructions described herein may be implemented in different formats. Additionally, example systems, architectures, and pipelines are detailed below. Examples of the (one or more) instructions may be executed on these systems, architectures, and pipelines, but are not limited to those detailed.
[0173] Figure 20Illustrates an example of an instruction format. As shown, an instruction can include multiple components, which include but are not limited to one or more fields for the following: one or more prefixes 2001, an opcode 2003, addressing information 2005 (e.g., register identifiers, memory addressing information, etc.), a displacement value 2007, and / or an immediate value 2009. Note that some instructions utilize some or all of the fields of this format, while other instructions may only use the fields of the opcode 2003. In some examples, the illustrated order is the order in which these fields are to be encoded, however it should be understood that in other examples, these fields can be encoded in a different order, combined, etc.
[0174] (One or more) prefix fields 2001 modify the instruction when used. In some examples, one or more prefixes are used for repeat string instructions (e.g., 0xF0, 0xF2, 0xF3, etc.), provide section override (e.g., 0x2E, 0x36, 0x3E, 0x26, 0x64, 0x65, 0x2E, 0x3E, etc.), perform bus lock operations, and / or change operand (e.g., 0x66) and address size (e.g., 0x67). Certain instructions require mandatory prefixes (e.g., 0x66, 0xF2, 0xF3, etc.). Some of these prefixes can be considered "traditional" prefixes. Other prefixes (one or more examples of which are detailed herein) indicate and / or provide further capabilities, such as specifying particular registers, etc. These other prefixes typically follow the "traditional" prefixes.
[0175] The opcode field 2003 is used to at least partially define the operation to be performed upon decoding of the instruction. In some examples, the length of the main opcode encoded in the opcode field 2003 is one, two, or three bytes. In other examples, the main opcode can be of other lengths. An additional 3-bit opcode field is sometimes encoded in another field.
[0176] The addressing information field 2005 is used to address one or more operands of the instruction, such as a location in memory or one or more registers. Figure 21Illustrates an example of the addressing information field 2005. In this illustration, an optional MOD R / M byte 2102 and an optional Scale, Index, Base (SIB) byte 2104 are shown. The MOD R / M byte 2102 and the SIB byte 2104 are used to encode up to two operands of an instruction, and each operand is either a direct register or an effective memory address. Note that these fields are all optional, that is, not all instructions include one or more of these fields. The MOD R / M byte 2102 includes a MOD field 2142, a register (reg) field 2144, and an R / M field 2146.
[0177] The content of the MOD field 2142 differentiates between memory access and non-memory access modes. In some examples, when the MOD field 2142 has a binary value of 11 (11b), register direct addressing mode is used, otherwise register indirect addressing mode is used.
[0178] The register field 2144 can encode a destination register operand or a source register operand, or can also encode an opcode extension and is not used to encode any instruction operand. The content of the register field 2144 directly specifies or specifies the location of the source or destination operand (in a register or in memory) through address generation. In some examples, the register field 2144 is supplemented with additional bits from a prefix (e.g., prefix 2001) to allow for larger addressing.
[0179] The R / M field 2146 can be used to encode an instruction operand that references a memory address, or can be used to encode a destination register operand or a source register operand. Note that in some examples, the R / M field 2146 can be combined with the MOD field 2142 to specify an addressing mode.
[0180] The SIB byte 2104 includes a scale field 2152, an index field 2154, and a base field 2156 for address generation. The scale field 2152 indicates a scaling factor. The index field 2154 specifies the index register to be used. In some examples, the index field 2154 is supplemented with additional bits from a prefix (e.g., prefix 2001) to allow for larger addressing. The base field 2156 specifies the base register to be used. In some examples, the base field 2156 is supplemented with additional bits from a prefix (e.g., prefix 2001) to allow for larger addressing. In practice, the content of the scale field 2152 allows the content of the index field 2154 to be scaled for memory address generation (e.g., for address generation using 2 缩放 *index + base).
[0181] Some addressing forms utilize displacement values to generate memory addresses. For example, memory addresses can be generated according to 2 缩放 *index + base + displacement, index * scale + displacement, r / m + displacement, instruction pointer (RIP / EIP) + displacement, register + displacement, etc. The displacement can be a value such as 1 byte, 2 bytes, 4 bytes, etc. In some examples, the displacement field 2007 provides this value. Additionally, in some examples, the use of the displacement factor is encoded in the MOD field of the addressing information field 2005, which indicates a compressed displacement scheme for which the displacement value is calculated and stored in the displacement field 2007.
[0182] In some examples, the immediate value field 2009 specifies an immediate value for the instruction. The immediate value can be encoded as a 1-byte value, 2-byte value, 4-byte value, etc.
[0183] Figure 22 An example of the first prefix 2001(A) is illustrated. In some examples, the first prefix 2001(A) is an example of a REX prefix. Instructions using this prefix can specify general-purpose registers, 64-bit packed data registers (e.g., single instruction multiple data (SIMD) registers or vector registers), and / or control registers and debug registers (e.g., CR8 - CR15 and DR8 - DR15).
[0184] Instructions using the first prefix 2001(A) can specify up to three registers using a 3-bit field, depending on the format: 1) using the reg field 2144 and the R / M field 2146 of the MOD R / M byte 2102; 2) using the MOD R / M byte 2102 and the SIB byte 2104, including using the reg field 2144 and the base field 2156 and the index field 2154; or 3) using the register field of the opcode.
[0185] In the first prefix 2001(A), bit positions 7:4 are set to 0100. Bit position 3 (W) can be used to determine the operand size but cannot determine the operand width alone. Thus, when W = 0, the operand size is determined by the code segment descriptor (CS.D), and when W = 1, the operand size is 64 bits.
[0186] Note that adding another bit allows addressing of 16 (2 4 ) registers, while the separate MOD R / M reg field 2144 and the R / M field 2146 of the MOD R / M can each address only 8 registers.
[0187] In the first prefix 2001(A), bit position 2 (R) can be an extension of the reg field 2144 of MOD R / M, and can be used to modify the reg field 2144 of MOD R / M when this field encodes a general-purpose register, a 64-bit packed data register (e.g., an SSE register), or a control or debug register. When the MOD R / M byte 2102 specifies other registers or defines an extended opcode, R is ignored.
[0188] Bit position 1 (X) can modify the SIB byte index field 2154.
[0189] Bit position 0 (B) can modify the base address in the R / M field 2146 of MOD R / M or the base address field 2156 of the SIB byte; or it can modify the opcode register field for accessing a general-purpose register (e.g., general-purpose register 1925).
[0190] Figures 23(A)-(D) illustrate examples of how the R, X, and B fields of the first prefix 2001(A) are used. Figure 23(A) illustrates that when the SIB byte 2104 is not used for memory addressing, R and B from the first prefix 2001(A) are used to extend the reg field 2144 and the R / M field 2146 of the MOD R / M byte 2102. Figure 23(B) illustrates that when the SIB byte 2104 is not used, R and B from the first prefix 2001(A) are used to extend the reg field 2144 and the R / M field 2146 of the MOD R / M byte 2102 (register-register addressing). Figure 23(C) illustrates that when the SIB byte 2104 is used for memory addressing, R, X, and B from the first prefix 2001(A) are used to extend the reg field 2144, the index field 2154, and the base address field 2156 of the MOD R / M byte 2102. Figure 23(D) illustrates that when a register is encoded in the opcode 2003, B from the first prefix 2001(A) is used to extend the reg field 2144 of the MOD R / M byte 2102.
[0191] Figures 24(A)-(B) illustrate an example of a second prefix 2001(B). In some examples, the second prefix 2001(B) is an example of a VEX prefix. The second prefix 2001(B) encoding allows an instruction to have more than two operands and allows SIMD vector registers (e.g., vector / SIMD register 1910) to be longer than 64 bits (e.g., 128 bits and 256 bits). The use of the second prefix 2001(B) provides a syntax for three-operand (or more) operations. For example, a previous two-operand instruction performed an operation such as A = A + B, which overwrote the source operand. The use of the second prefix 2001(B) enables the operands to perform non-destructive operations, such as A = B + C.
[0192] In some examples, the second prefix 2001(B) has two forms - a two-byte form and a three-byte form. The two-byte second prefix 2001(B) is mainly used for 128-bit, scalar, and some 256-bit instructions; while the three-byte second prefix 2001(B) provides a compact replacement for 3-byte opcode instructions and the first prefix 2001(A).
[0193] Figure 24(A) illustrates an example of the two-byte form of the second prefix 2001(B). In one example, the format field 2401 (byte 0 2403) contains the value C5H. In one example, byte 1 2405 includes an "R" value in bit [7]. This value is the complement of the "R" value of the first prefix 2001(A). Bit [2] is used to specify the length (L) of the vector (where a value of 0 is scalar or a 128-bit vector, and a value of 1 is a 256-bit vector). Bits [1:0] provide opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). Bits [6:3], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (ones' complement) form and is valid for instructions with 2 or more source operands; 2) encoding the destination register operand, which is specified in ones' complement form, for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a value, such as 1111b.
[0194] Instructions using this prefix can use the R / M field 2146 of the MOD R / M to encode instruction operands that reference a memory address, or to encode a destination register operand or a source register operand.
[0195] Instructions using this prefix can use the reg field 2144 of the MOD R / M to encode a destination register operand or a source register operand, or be treated as an opcode extension and not be used to encode any instruction operand.
[0196] For instruction syntax that supports four operands, vvvv, the R / M field 2146 of the MOD R / M, and the reg field 2144 of the MOD R / M encode three of the four operands. Then bits [7:4] of the immediate value field 2009 are used to encode the third source register operand.
[0197] Figure 24(B) illustrates an example of the three-byte form of the second prefix 2001(B). In one example, the format field 2411 (byte 0 2413) contains the value C4H. Byte 1 2415 includes "R", "X", and "B" in bits [7:5], which are the complements of these values of the first prefix 2001(A). Bits [4:0] of byte 1 2415 (shown as mmmmm) include content for encoding one or more implicit leading opcode bytes as needed. For example, 00001 means a 0FH leading opcode, 00010 means a 0F38H leading opcode, 00011 means a 0F3AH leading opcode, and so on.
[0198] The use of bit [7] of byte 2 2417 is similar to W of the first prefix 2001(A), including helping to determine the size of the promotable operand. Bit [2] is used to specify the length (L) of the vector (where a value of 0 is a scalar or a 128-bit vector, and a value of 1 is a 256-bit vector). Bits [1:0] provide opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). Bits [6:3], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (one's complement) form and is valid for instructions with two or more source operands; 2) encoding the destination register operand, which is specified in one's complement form and is used for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a certain value, such as 1111b.
[0199] Instructions using this prefix can use the R / M field 2146 of the MOD R / M to encode an instruction operand that references a memory address, or encode a destination register operand or a source register operand.
[0200] Instructions using this prefix can use the reg field 2144 of the MOD R / M to encode the destination register operand or the source register operand, or be treated as an opcode extension and not be used to encode any instruction operand.
[0201] For instruction syntax that supports four operands, vvvv, the R / M field 2146 of the MOD R / M, and the reg field 2144 of the MOD R / M encode three of the four operands. Then bits [7:4] of the immediate value field 2009 are used to encode the third source register operand.
[0202] Figure 25 Illustrates an example of the third prefix 2001(C). In some examples, the third prefix 2001(C) is an example of an EVEX prefix. The third prefix 2001(C) is a four-byte prefix.
[0203] The third prefix 2001(C) can encode 32 vector registers (e.g., 128-bit, 256-bit, and 512-bit registers) in 64-bit mode. In some examples, instructions with a write mask / operation mask (see the discussion of registers in the previous figure, e.g., Figure 19 ) or predicates utilize this prefix. The operation mask register allows conditional processing or selective control. Operation mask instructions - whose source / destination operand is the operation mask register and treat the content of the operation mask register as a single value - are encoded using the second prefix 2001(B).
[0204] The third prefix 2001(C) can encode function specific to instruction classes (e.g., packed instructions with "load + operation" semantics can support an embedded broadcast function, floating-point instructions with rounding semantics can support a static rounding function, floating-point instructions with non-rounding arithmetic semantics can support a "suppress all exceptions" function, etc.).
[0205] The first byte of the third prefix 2001(C) is the format field 2511, which has a value of 62H in one example. The subsequent bytes are called the payload bytes 2515 - 2519 and together form a 24-bit value of P[23:0], providing specific capabilities in the form of one or more fields (detailed herein).
[0206] In some examples, P[1:0] of payload byte 2519 is the same as the two lowermost mm bits. In some examples, P[3:2] is reserved. Bit P[4] (R') allows access to the upper 16 vector register set when combined with P[7] and the reg field 2144 of MOD R / M. When SIB type addressing is not needed, P[6] can also provide access to the upper 16 vector registers. P[7:5] consists of R, X, and B, which are operand specifier modifier bits for vector registers, general-purpose registers, and memory addressing, and when combined with the MOD R / M register field 2144 and the R / M field 2146 of MOD R / M, allow access to the next set of 8 registers beyond the lower 8 registers. P[9:8] provides opcode extension equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). P
[10] is a fixed value 1 in some examples. P[14:11], shown as vvvv, can be used for: 1) encoding a first source register operand, which is specified in inverted (one's complement) form and is valid for instructions with 2 or more source operands; 2) encoding a destination register operand, which is specified in one's complement form and is used for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a value, e.g., 1111b.
[0207] P
[15] is similar to W of the first prefix 2001(A) and the second prefix 2011(B), and can be used as an opcode extension bit or an operand size promotion.
[0208] P[18:16] specifies the index of the register in the operation mask (write mask) register (e.g., write mask / predicate register 1915). In one example, a specific value aaa = 000 has special behavior, implying that no operation mask is used for this particular instruction (this can be achieved in multiple ways, including using a hard-wired all-ones operation mask or hardware that bypasses the masking hardware). When combined, the vector mask allows any set of elements in the destination to be protected from update during the execution of any operation (specified by the base and enhanced operations); in another example, the old value of each element of the destination is retained (if the corresponding mask bit has a value of 0). In contrast, when zeroing, the vector mask allows any set of elements in the destination to be zeroed during the execution of any operation (specified by the base and enhanced operations); in one example, the elements of the destination are set to 0 when the corresponding mask bit has a value of 0. A subset of this functionality is the ability to control the vector length of the operation being performed (i.e., the span of the elements being modified, from the first to the last); however, the elements being modified do not have to be contiguous. Thus, the operation mask field allows for partial vector operations, including loads, stores, arithmetic, logic, etc. While in the described example, the content of the operation mask field selects which of several operation mask registers contains the operation mask to be used (so the content of the operation mask field indirectly identifies the masking to be performed), alternatively or additionally, alternative examples allow the content of the mask write field to directly specify the masking to be performed.
[0209] P
[19] can be combined with P[14:11] to encode the second source vector register in a non-destructive source syntax that can utilize P
[19] to access the upper 16 vector registers. P
[20] encodes various functions that vary among different classes of instructions and can affect the meaning of the vector length / rounding control specifier field (P[22:21]). P
[23] indicates support for merge-write masking (e.g., when set to 0) or support for zeroing and merge-write masking (e.g., when set to 1).
[0210] The following table details exemplary examples of the encoding of registers in instructions using the third prefix 2001(C). Table 1: 32-register support in 64-bit mode Table 2: Encoding register specifiers in 32-bit mode Table 3: Operation mask register specifier encoding
[0211] The program code can be applied to the input information to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor, such as a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a microprocessor, or any combination thereof.
[0212] The program code can be implemented in a high-level procedural or object-oriented programming language to communicate with the processing system. If desired, the program code can also be implemented in assembly or machine language. In fact, the mechanisms described herein are not limited in scope to any particular programming language. In any case, the language can be a compiled language or an interpreted language.
[0213] Examples of the mechanisms disclosed herein can be implemented in hardware, software, firmware, or any combination of these implementation approaches. The examples can be implemented as a computer program or program code, executed on a programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0214] One or more aspects of at least one example can be implemented by representative instructions stored on a machine-readable medium, which represent various logics within a processor. When these instructions are read by the machine, they cause the machine to fabricate the logic for performing the techniques described herein. These representations are referred to as “intellectual property (IP) cores” and can be stored on a tangible machine-readable medium and provided to various customers or manufacturing facilities to be loaded into the manufacturing machines that make the logic or processor.
[0215] These machine-readable storage media can include—but are not limited to—non-transitory tangible arrangements of articles manufactured or formed by a machine or device, including storage media such as: hard disks, any other type of disk (including floppy disks, optical disks, compact disk read-only memory (CD-ROM), compact disk rewritable (CD-RW), and magneto-optical disks), semiconductor devices (such as read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM) and static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM), phase change memory (PCM)), magnetic or optical cards, or any other type of medium suitable for storing electronic instructions.
[0216] Accordingly, examples also include non-transitory tangible machine-readable media that contain instructions or contain design data that defines the structural, circuit, device, processor, and / or system features described herein, such as a Hardware Description Language (HDL). Such examples may also be referred to as program products.
[0217] Emulation (including binary translation, code morphing, etc.).
[0218] In some cases, an instruction converter can be used to convert instructions from a source instruction set architecture to a target instruction set architecture. For example, the instruction converter can translate (e.g., using static binary translation, dynamic binary translation including dynamic compilation), morph, emulate, or otherwise convert instructions to one or more other instructions to be processed by the core. The instruction converter can be implemented in software, hardware, firmware, or a combination thereof. The instruction converter can be on the processor, off the processor, or part on the processor and part off the processor.
[0219] Figure 26The block diagram illustrates the use of a software instruction converter according to an example for converting binary instructions in a source ISA to binary instructions in a target ISA. In the illustrated example, the instruction converter is a software instruction converter, but alternatively, the instruction converter may be implemented in software, firmware, hardware, or various combinations thereof. Figure 26 Illustrated is that a program in a high-level language 2602 can be compiled using a first ISA compiler 2604 to generate a first ISA binary code 2606, which can be natively executed by a processor 2616 having at least one first ISA core. The processor 2616 having at least one first ISA core represents any such processor that can perform substantially the same functions as an Intel processor having at least one first ISA core by compatibly executing or otherwise processing (1) a substantial portion of the first ISA or (2) a target code version of an application or other software targeted to run on a processor having at least one first ISA core, so as to achieve substantially the same results as a processor having at least one first ISA core. The first ISA compiler 2604 represents a compiler operable to generate the first ISA binary code 2606 (e.g., target code) that can be executed on the processor 2616 having at least one first ISA core with or without additional linking processing. Similarly, Illustrated is that a program in a high-level language 2602 can be compiled using an alternative ISA compiler 2608 to generate an alternative ISA binary code 2610, which can be natively executed by a processor 2614 that does not have a first ISA core. An instruction converter 2612 is used to convert the first ISA binary code 2606 into code that can be natively executed by the processor 2614 that does not have a first ISA core. This converted code does not necessarily have to be the same as the alternative ISA binary code 2610; however, the converted code will implement the overall operation and be composed of instructions from that alternative ISA. Thus, the instruction converter 2612 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device that does not have a first ISA processor or core to execute the first ISA binary code 2606 through emulation, simulation, or any other process. Figure 26
[0220] The components, features, and details described for any of the processors disclosed herein can optionally be applied to any of the methods disclosed herein. In an embodiment, the method can optionally be performed by and / or utilize such a processor. In an embodiment, any of the processors described herein can optionally be included in any of the systems disclosed herein. Any of the processors disclosed herein can optionally have any of the microarchitectures shown herein. Any of the instructions disclosed herein can optionally be executed by any of the processors disclosed herein. Additionally, in some embodiments, any of the instructions disclosed herein can optionally have any of the features or details of the instruction formats shown herein.
[0221] References to "an example", "example", etc. indicate that the described example may include a particular feature, structure, or characteristic, but not every example necessarily includes that particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same example. Additionally, when a particular feature, structure, or characteristic is described in connection with an example, it is considered within the knowledge of one of ordinary skill in the art to affect such feature, structure, or characteristic in connection with other examples whether or not explicitly described.
[0222] The processor components disclosed herein may be said and / or claimed to be operable, capable of operating, able to, can, be configured, be adapted, or otherwise perform an operation. For example, a decoder may be said and / or claimed to be for decoding instructions, an execution unit may be said and / or claimed to be for storing results, and so on. As used herein, these expressions refer to the characteristics, properties, or attributes of a component when in a powered-off state, and do not imply that the component or the device or apparatus in which the component is included is currently powered on or operating. For clarity, it should be understood that the processors and devices claimed herein are not required to be powered on or operating.
[0223] In the specification and claims, the terms "coupled" and / or "connected" and their derivatives may be used. These terms are not intended as synonyms for each other. Instead, in various embodiments, "connected" may be used to indicate that two or more elements are in direct physical and / or electrical contact with each other. "Coupled" may mean that two or more elements are in direct physical and / or electrical contact with each other. However, "coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other. For example, an execution unit may be coupled to a context unit through one or more intermediary components. In the figures, arrows are used to show connections and couplings.
[0224] Some embodiments include an article (e.g., a computer program product) that includes a machine-readable medium. The medium can include a mechanism for providing (e.g., storing) information in a machine-readable form. The machine-readable medium can provide or store thereon instructions or sequences of instructions that, if executed by a machine and / or when executed by a machine, are operable to cause the machine to perform and / or cause the machine to perform one or more operations, methods, or techniques disclosed herein.
[0225] In some embodiments, the machine-readable medium can include a tangible and / or non-transitory machine-readable storage medium. For example, the non-transitory machine-readable storage medium can include a floppy disk, an optical storage medium, an optical disc, an optical data storage device, a CD-ROM, a magnetic disk, a magneto-optical disk, a read only memory (ROM), a programmable ROM (PROM), an erasable-and-programmable ROM (EPROM), an electrically-erasable-and-programmable ROM (EEPROM), a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a flash memory, a phase change memory, a phase change data storage material, a non-volatile memory, a non-volatile data storage device, a non-transitory memory, or a non-transitory data storage device, and so on. The non-transitory machine-readable storage medium does not consist of a transitory propagated signal. In some embodiments, the storage medium can include a tangible medium that includes solid-state substances or materials such as, for example, semiconductor materials, phase change materials, magnetic solid materials, solid data storage materials, and the like. Alternatively, a non-tangible transitory computer-readable transmission medium can optionally be used, such as, for example, electrical, optical, acoustic, or other forms of propagated signals - such as carrier waves, infrared signals, and digital signals.
[0226] Examples of suitable machines include, but are not limited to, general-purpose processors, special-purpose processors, digital logic circuits, integrated circuits, and the like. Some other examples of suitable machines include computer systems or other electronic devices that include a processor, digital logic circuit, or integrated circuit. Examples of such computer systems or electronic devices include, but are not limited to, desktop computers, laptop computers, notebook computers, tablet computers, netbooks, smart phones, cellular phones, servers, network devices (e.g., routers and switches), mobile Internet devices (MID), media players, smart TVs, Internet machines, set-top boxes, and video game controllers.
[0227] Moreover, in the examples described above, unless otherwise specifically stated, separating languages such as the phrase "at least one of A, B, or C" or "A, B, and / or C" are intended to be understood to mean A, B, or C, or any combination thereof (i.e., A and B, A and C, B and C, and A, B, and C).
[0228] In the above description, specific details have been set forth to provide a thorough understanding of the embodiments. However, other embodiments may be practiced without some of these specific details. Various modifications and changes may be made to the present disclosure without departing from the broader spirit and scope of the present disclosure as set forth in the claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. The scope of the present invention is not intended to be determined by the specific examples provided above, but only by the appended claims. In other instances, well-known circuits, structures, devices, and operations have been shown in block diagram form and / or without detail to avoid obscuring the understanding of the specification.
[0229] Example embodiments
[0230] The following examples relate to further embodiments. The details in the examples may be used anywhere in one or more embodiments.
[0231] Example 1 is a device that includes: context storage means for storing the context of a logical processor; and an execution unit coupled to the context storage means. The execution unit is configured to perform an operation corresponding to a control primitive or perform an operation in response to an exception condition. The operations include: selectively saving a first subset of the context from a first subset of the context storage means that is written after entry into a protected execution environment to a system memory. The operations further include: causing the logical processor to exit the protected execution environment.
[0232] Example 2 includes the device as described in Example 1, wherein, in order to perform an operation corresponding to a control primitive, the execution unit is configured to determine not to write a second subset of the context to the system memory from a second subset of the context storage means that is not written after entry into the protected execution environment.
[0233] Example 3 includes the device as described in any one of Examples 1 or 2, wherein, in order to selectively save the first subset of the context, the execution unit is configured to selectively save only the first subset of the context from only the first subset of the context storage means to a context save area in the system memory.
[0234] Example 4 includes the apparatus according to any one of Examples 1 to 3, wherein, to perform an operation corresponding to a control primitive, the execution unit is configured to selectively sanitize a second subset of a context storage device that is read, written, or both read and written after entry into the protected execution environment.
[0235] Example 5 includes the apparatus according to Example 4, wherein, to selectively sanitize the second subset of the context storage device, the execution unit is configured to perform at least one operation selected from the group consisting of: (1) erasing contexts in the second subset of the context storage device; (2) overwriting contexts in the second subset of the context storage device; (3) obfuscating contexts in the second subset of the context storage device; (4) destroying contexts in the second subset of the context storage device; and (5) resetting the second subset of the context storage device.
[0236] Example 6 includes the apparatus according to any one of Examples 1 to 5, wherein, to perform an operation corresponding to a control primitive, the execution unit is configured to determine not to sanitize a subset of the context storage device that is not read or written after entry into the protected execution environment.
[0237] Example 7 includes the apparatus according to any one of Examples 1 to 6, wherein, to perform an operation corresponding to a control primitive, the execution unit is configured to: determine that a context element has been written after entry into the protected execution environment by determining that context element metadata corresponding to the context element in the first subset of the context storage device indicates that the context element has been written.
[0238] Example 8 includes the apparatus according to any one of Examples 1 to 7, optionally wherein the context element includes a register of a logical processor, and optionally wherein the context element metadata includes a write indication bit in a set of write indication bits, each write indication bit corresponding to a different context element of the logical processor.
[0239] Example 9 includes the apparatus according to any one of Examples 1 to 8, wherein the logical processor includes: a decoding unit configured to decode a context access instruction; and circuitry coupled to the decoding unit and configured to perform an operation corresponding to the context access instruction. The operation includes: determining that a context element in the context storage device has not been read or written after entry into the protected execution environment; and loading a context from a context save area of a system memory into the context element in the context storage device.
[0240] Example 10 includes the apparatus according to any one of Examples 1 to 9, wherein the logical processor includes: a decoding unit configured to decode a context access instruction; and circuitry coupled to the decoding unit and configured to perform an operation corresponding to the context access instruction. The operation includes performing an access control check on a page to be used to store a context to be accessed by the context access instruction.
[0241] Example 11 includes the apparatus according to any one of Examples 1 to 10, wherein the logical processor includes an execution unit configured to perform operations corresponding to a second control primitive, including: changing context metadata corresponding to a context element in a context storage device to indicate that the context element has not been read and has not been written; and causing the logical processor to enter a protected execution environment.
[0242] Example 12 includes the apparatus according to any one of Examples 1 to 11, wherein the control primitive is one of the following: (1) an instruction in an instruction set of a processor, wherein the apparatus further includes a decoding unit coupled to the execution unit and configured to decode the instruction; or (2) a command to be stored in a location selected from a group including a control register and a memory-mapped input / output (MMIO) region.
[0243] Example 13 is a method that includes: storing a context of a logical processor in a context storage device; and performing an operation corresponding to a control primitive or an exception condition. The operation includes: selectively saving a first subset of the context from a first subset of the context storage device that has been written after entering a protected execution environment to a system memory; and causing the logical processor to exit the protected execution environment.
[0244] Example 14 includes the method according to Example 13, wherein the operation includes determining not to save a second subset of the context from a second subset of the context storage device that has not been written after entering a protected execution environment to the system memory.
[0245] Example 15 includes the method according to any one of Examples 13 to 14, wherein the operation includes selectively purging a second subset of the context storage device that has been read, written, or both read and written after entering a protected execution environment.
[0246] Example 16 includes the apparatus as described in Example 15, wherein selectively purifying a second subset of the context storage device includes at least one operation selected from the group consisting of: (1) erasing the context in the second subset of the context storage device; (2) overwriting the context in the second subset of the context storage device; (3) obfuscating the context in the second subset of the context storage device; (15) destroying the context in the second subset of the context storage device; and (5) resetting the second subset of the context storage device.
[0247] Example 17 includes the method as described in any one of Examples 13 to 16, wherein the operation includes: determining not to purify a subset of the context storage device that has not been read or written after entering the protected execution environment.
[0248] Example 18 is a non-transitory machine-readable storage medium that stores instructions that, if executed by a machine, cause the machine to perform operations corresponding to control primitives or exceptional conditions. The operations include: storing a first subset of the context of the logical processor in the system memory; and performing an operation corresponding to a control primitive. The operations include: selectively saving a first subset of the context from a first subset of the context storage device that has been written after entering the protected execution environment to the system memory; and causing the logical processor to exit the protected execution environment.
[0249] Example 19 includes the non-transitory machine-readable storage medium as described in Example 18, wherein the operation includes: determining not to write a second subset of the context from a second subset of the context storage device that has not been written after entering the protected execution environment to the system memory.
[0250] Example 20 includes the non-transitory machine-readable storage medium as described in any one of Examples 18 to 19, wherein the operation includes: selectively purifying a subset of the context storage device that has been read, written, or both read and written after entering the protected execution environment.
[0251] Example 21 is an apparatus including: means for performing the method as described in any one of Examples 13 to 17.
[0252] Example 22 is an apparatus including: circuitry for performing the method as described in any one of Examples 13 to 17.
Claims
1. A device comprising: A context storage device for storing the context of the logical processor; as well as An execution unit, coupled to the context storage device, configured to execute an operation corresponding to a control primitive or to execute an operation in response to an abnormal condition, the operation comprising: selectively saving a first subset of the context to system memory from a first subset of the context storage device written upon entry into the protected execution environment; as well as The logical processor is caused to exit the protected execution environment.
2. The device according to claim 1, wherein: To perform the operation, the execution unit is to determine not to write a second subset of the context to the system memory from a second subset of the context storage that has not been written to after entry into the protected execution environment.
3. The device of claim 1, wherein: To selectively save the first subset of the context, the execution unit is configured to selectively save only the first subset of the context from the first subset of the context storage device to a context save area in the system memory.
4. The device of claim 1, wherein: To perform the operation, the execution unit is operable to selectively purge a second subset of the context storage that is read, written, or both read and written after entry into the protected execution environment.
5. The device of claim 4, wherein: To selectively purge the second subset of the context storage, the execution unit is configured to perform at least one operation selected from the group consisting of: erasing contexts in a second subset of the context storage device; overwriting the contexts in the second subset of the context storage device; obfuscating contexts in a second subset of the context storage device; destroying contexts in a second subset of the context storage device; as well as A second subset of the context storage is reset.
6. The device of claim 4, wherein: To perform the operation, the execution unit is operable to determine not to purge a subset of the context storage that has not been read or written after entry into the protected execution environment.
7. The device of claim 1, wherein: To perform the operation, the execution unit is configured to determine that the context element has been written upon entry into the protected execution environment by determining that context element metadata corresponding to the context element in the first subset of the context storage device indicates that the context element has been written.
8. The device of claim 7, wherein: The context element comprises a register of the logical processor, and wherein the context element metadata comprises a write indication bit of a set of write indication bits, each write indication bit corresponding to a different context element of the logical processor.
9. The device of claim 1, wherein: The logical processor comprises: a decoding unit, the decoding unit being configured to decode a context access instruction; and A circuit is coupled to the decoding unit and configured to perform an operation corresponding to the context access instruction, the operation comprising: determining that a context element in the context storage has not been read or written after entry into the protected execution environment; and A context is loaded from a context save area of the system memory into the context element in the context storage device.
10. The device of claim 1, wherein: The logical processor comprises: a decoding unit, the decoding unit being configured to decode a context access instruction; and Circuitry, coupled to the decode unit, for performing operations corresponding to the context access instruction, the operations including performing an access control check on a page to be used to store a context to be accessed by the context access instruction.
11. The device of claim 1, wherein: The logical processor includes an execution unit, and the execution unit is used to execute an operation corresponding to the second control primitive, including: changing context metadata corresponding to a context element in the context storage to indicate that the context element has not been read and has not been written; and The logical processor is caused to enter the protected execution environment.
12. The apparatus of claim 1, wherein: The control primitive is one of the following: instructions in an instruction set of a processor, wherein the device further comprises a decoding unit coupled to the execution unit, the decoding unit being configured to decode the instruction; or For commands to be stored in a location selected from the group consisting of a control register and a memory mapped input / output MMIO area.
13. A method comprising: storing the context of the logical processor in a context storage device; as well as Perform actions corresponding to control primitives or exception conditions, including: selectively saving a first subset of the context to a system memory from a first subset of the context storage that has been written to after entry into the protected execution environment; as well as The logical processor is caused to exit the protected execution environment.
14. The method of claim 13, wherein: The operations include determining not to save a second subset of the context to the system memory from a second subset of the context storage that has not been written to after entry into the protected execution environment.
15. The method of claim 13, wherein: The operations include selectively purge a second subset of the context storage that is read, written, or both read and written after entry into the protected execution environment.
16. The method of claim 15, wherein: Purging the second subset of the context storage comprises at least one operation selected from the group consisting of: erasing contexts in a second subset of the context storage device; overwriting the contexts in the second subset of the context storage device; obfuscating contexts in a second subset of the context storage device; destroying contexts in a second subset of the context storage device; as well as A second subset of the context storage is reset.
17. The method of claim 15, wherein: The operations include determining not to sanitize a subset of the context storage that has not been read or written after entry into the protected execution environment.
18. The method according to any one of claims 13 to 17, wherein: Selectively saving the first subset of the context includes selectively saving only the first subset of the context from the first subset of the context storage to a context save area in the system memory.
19. The method of any one of claims 13 to 17, further comprising: It is determined that the context element has been written upon entry into the protected execution environment by determining that context element metadata corresponding to the context elements in the first subset of the context storage indicates that the context element has been written.
20. The method of any one of claims 13 to 17, wherein: The context element metadata includes a write indication bit in a write indication bit set, each write indication bit corresponding to a different context element of the logical processor.
21. The method of any one of claims 13 to 17, further comprising: Decode context access instructions; as well as Executing an operation corresponding to the context access instruction includes: determining that a context element in the context storage has not been read or written after entry into the protected execution environment; as well as A context is loaded from a context save area of the system memory into the context element in the context storage device.
22. A non-transitory machine-readable storage medium storing instructions that, if executed by a machine, cause the machine to perform operations corresponding to a control primitive or an exception condition, the operations comprising: selectively saving a first subset of the context of the logical processor to a system memory from a first subset of the context storage device written to after entry into the protected execution environment; as well as The logical processor is caused to exit the protected execution environment.
23. The non-transitory machine-readable storage medium of claim 22, wherein: The operations include determining not to write a second subset of the context to the system memory from a second subset of the context storage that has not been written to following entry into the protected execution environment.
24. The non-transitory machine-readable storage medium of any one of claims 18 to 19, wherein: The operations include selectively purge a subset of the context storage that is read, written, or both read and written after entry into the protected execution environment.
25. A device, comprising: first means for storing a context of a logical processor; as well as a second device coupled to the first device, the second device being configured to perform an operation corresponding to a control primitive or to perform an operation in response to an abnormal condition, the operation comprising: selectively saving a first subset of the context to system memory from a first subset of the first means to which it was written upon entry into the protected execution environment; as well as The logical processor is caused to exit the protected execution environment.