Maintenance operations across fragmented memory regions

JP2025514845A5Pending Publication Date: 2026-04-22ARM LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ARM LTD
Filing Date
2023-04-21
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve memory system performance, especially in complex scenarios of multi-region and multi-execution environments, where there are improper problems of data access and maintenance operations.

Method used

By introducing memory protection circuitry that manages execution environment and multi-region memory, defines encryption points, and uses region and execution environment-specific keys for data encryption and decryption, while limiting maintenance operations on encrypted memory.

Benefits of technology

Data isolation and protection for different regions and execution environments are realized, unnecessary cache maintenance operations are reduced, thereby improving the performance and security of the memory system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An apparatus is provided in which a processing circuit performs processing in one of a fixed number of at least two regions. One of the regions is subdivided into a variable number of execution environments, one of the execution environments being a managed execution environment configured to manage the execution environments. A memory protection circuit defines a point of encryption after at least one unencrypted storage circuit of the memory hierarchy and before at least one encrypted storage circuit of the memory hierarchy. The at least one encrypted storage circuit uses a key input to perform encryption or decryption on data of a memory access request issued from within a current one of the regions. The key input is different for each of the regions and each of the execution environments, and the managed execution environment is configured to prevent issuing maintenance operations to the at least one encrypted storage circuit of the memory hierarchy.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present technology relates to data processing.

[0002] It is desirable to reduce memory system performance improvements where possible in the memory hierarchy. Summary of the Invention

[0003] Viewed from a first exemplary configuration, an apparatus is provided, comprising: a processing circuit configured to execute processing in one of a fixed number of at least two regions, one of the regions being subdivided into a variable number of execution environments, one of the execution environments being a managed execution environment configured to manage the execution environments; and a memory protection circuit defining a point of encryption after at least one unencrypted storage circuit of the memory hierarchy and before at least one encrypted storage circuit of the memory hierarchy, the at least one encrypted storage circuit being configured to use a key input to perform encryption or decryption on data of a memory access request issued from within a current one of the regions, the key input being different for each of the regions and each of the execution environments, and the managed execution environment being configured to prevent issuing maintenance operations to the at least one encrypted storage circuit of the memory hierarchy.

[0004] Viewed from a second exemplary configuration, a method is provided, the method comprising: executing a process in one of a fixed number of at least two regions, one of the regions being subdivided into a variable number of execution environments, one of the execution environments being a managed execution environment configured to manage the execution environments; defining a point of encryption after at least one unencrypted storage circuit of the memory hierarchy and before at least one encrypted storage circuit of the memory hierarchy; preventing issuance of maintenance operations to the at least one encrypted storage circuit of the memory hierarchy; and using a key input to perform encryption or decryption on data of a memory access request issued to a memory address from within a current one of the regions, the key input being different for each of the regions and each of the execution environments, and the managed execution environment being configured to prevent issuance of maintenance operations to the at least one encrypted storage data structure of the memory hierarchy.

[0005] Viewed from a third exemplary configuration, there is provided a computer program for controlling a host data processing apparatus to provide an instruction environment for executing target code, the computer program comprising: processing program logic configured to simulate processing of the target code in one of at least two regions, one of the regions being subdivided into a variable number of execution environments, one of the execution environments being a managed execution environment configured to manage the execution environments; and memory protection program logic configured to define a point of encryption after at least one unencrypted storage data structure of the memory hierarchy and before at least one encrypted storage data structure of the memory hierarchy, the at least one encrypted storage data structure being configured to use a key input to perform encryption or decryption on data of a memory access request issued from within a current one of the regions, the key input being different for each of the regions and each of the execution environments, and the managed execution environment being configured to prevent issuing maintenance operations to the at least one encrypted storage data structure of the memory hierarchy. [Brief description of the drawings]

[0006] The present technology will now be further described, by way of example only, with reference to embodiments thereof illustrated in the accompanying drawings, in which: [Figure 1] 1 illustrates an example according to some embodiments. [Diagram 2] 1 illustrates an embodiment of a separate root domain that manages domain switching. [Diagram 3] 2 illustrates a schematic diagram of another embodiment of a processing system. [Figure 4] The granule protection table is used to illustrate how the system physical address space can be partitioned. [Diagram 5] The operation of the address translation circuit and the PAS filter will be summarized. [Figure 6] 1 illustrates an example of a page table entry. [Figure 7]1 illustrates an embodiment of a MECID consumer working in conjunction with a PAS TAG stripper that functions as a memory protection circuit. [Figure 8] 1 illustrates a flow chart according to some of the above embodiments. [Figure 9] 1 illustrates a simulator implementation that may be used. [Figure 10] Illustrates the location of encryption points and the extent to which deletion and revocation operations extend within the system. [Figure 11] 1 illustrates the relationship between the cache hierarchy, PoE, and PoPA. [Figure 12] 1 shows a flowchart illustrating the cache maintenance behavior in more detail. [Figure 13A] 1 illustrates one embodiment of targeting of cache maintenance operations. [Figure 13B] 13 illustrates another embodiment of targeting of cache maintenance operations. [Figure 14] 1 illustrates a method of data processing according to some embodiments. [Figure 15] 1 illustrates a simulator implementation that may be used. [Figure 16] 1 illustrates an example system in accordance with some embodiments. [Figure 17] 1 illustrates an example of a MECID mismatch. [Figure 18] A poison mode of operation is illustrated in which, in response to a mismatch, the associated cache line is poisoned. [Figure 19] 1 illustrates an example implementation in which an alias mode of operation is indicated. [Figure 20] 1 illustrates an embodiment of a delete mode of operation. [Figure 21] 1 illustrates an embodiment of an erase mode of operation. [Figure 22] 1 illustrates, in flow chart form, an embodiment of how discrepancies are handled in different operating modes. [Diagram 23] 1 illustrates, in flow chart form, the interaction between enabled modes and speculative execution. [Figure 24] 1 illustrates a simulator implementation that may be used.

[0007] Before discussing the embodiments with reference to the accompanying drawings, certain embodiments and associated advantages are described below. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0008] According to one example configuration, an apparatus is provided, the apparatus comprising: a processing circuit configured to execute processing in one of a fixed number of at least two regions, one of the regions being subdivided into a variable number of execution environments, one of the execution environments being a managed execution environment configured to manage the execution environments; and a memory protection circuit defining a point of encryption after at least one unencrypted storage circuit of the memory hierarchy and before at least one encrypted storage circuit of the memory hierarchy, the at least one encrypted storage circuit being configured to use a key input to perform encryption or decryption on data of a memory access request issued from within a current one of the regions, the key input being different for each of the regions and each of the execution environments, and the managed execution environment being configured to prevent issuing maintenance operations to the at least one encrypted storage circuit of the memory hierarchy.

[0009] Processing may occur in several (two or more (e.g., three or more)) regions or worlds. One of these regions / worlds is subdivided into several (e.g., multiple) execution environments, one of these execution environments being a managed execution environment responsible for managing each of the execution environments. The managed execution environment handles, for example, cache maintenance operations. Memory protection circuits are provided that protect the memory. For example, they may manage the isolation of the memory used by each of the regions. The memory protection circuits define points of encryption in the memory hierarchy. The memory hierarchy systems (storage circuits) before the point of encryption store unencrypted data, whereas the memory hierarchy systems (storage circuits) after the point of encryption store encrypted data. The encryption used for these encrypted storage circuits is different for each of the regions and each of the execution environments. That is, unless explicitly requested, software executing in one region or execution environment cannot decrypt data belonging to software in another region or execution environment. This is achieved by using key inputs (e.g., a key, a part of a key, or tunable bits) during the encryption process that are different for each region and / or execution environment. At least some of the cache maintenance operations issued by the managed execution environment are directed to unencrypted storage circuits, not to encrypted storage circuits. This is because data belonging to one area or execution environment is generally not accessible to another area or execution environment (except in certain specially defined circumstances), so there may be no need to delete the data to prevent it from inadvertently becoming accessible to another area or execution environment. Thus, the number of cache maintenance operations can be limited, and thus the impact on performance can be reduced.

[0010] In some embodiments, the management execution environment is configured to issue a maintenance operation to at least one unencrypted storage circuit of the memory hierarchy in response to a memory allocation change made to one of the execution environments. Since the unencrypted storage circuits of the memory hierarchy store data in an unencrypted format, it is important that the maintenance operation specifically targets these storage circuits. Since the data generally cannot be accessed by other execution environments (or regions / worlds), it becomes less critical for a particular maintenance operation to be performed. The memory allocation change may occur, for example, as a result of an execution environment being terminated or as a result of a new execution environment being started.

[0011] In some embodiments, the maintenance operation is an invalidation operation. An invalidation operation marks data in the cache as unavailable (e.g., deleted) so that it must be obtained from elsewhere in the memory hierarchy, such as memory. By invalidating to the point of encryption, the data can no longer be accessed unless a decryption process is performed. Thus, if the key input associated with the data has also been erased or lost, the data is no longer accessible. It is important to ensure that the previous execution environment that used the memory space where the data was stored unencrypted in the unencrypted storage circuitry invalidates the data so that it cannot be accessed by the new execution environment. This is accomplished by using a cache maintenance operation that targets the unencrypted storage circuitry. Because the data associated with the old execution environment is encrypted, there is no need to perform the same maintenance operation to target the encrypted storage circuitry. The new execution environment does not have access to the old keys of the old execution environment and therefore cannot decrypt the data.

[0012] In some embodiments, the maintenance operation is a delete and invalidate operation that causes dirty (modified) data to be written further up the memory hierarchy, for example to memory backed by DRAM, and simultaneously invalidates entries in caches of the memory hierarchy such that future accesses to the data are accomplished by retrieving the data from memory.

[0013] In some embodiments, the maintenance operation is configured to invalidate entries in at least one unencrypted storage circuit associated with one of the execution environments. Thus, the invalidate maintenance operation is directed to those entries in the unencrypted storage circuit (where data is stored in an unencrypted manner) that are associated with or belong to a particular one of the execution environments. Data that belongs to other execution environments remains valid unless / until it is targeted by another invalidate operation. Targeting of entries that belong to a particular execution environment may be achieved by issuing a cache maintenance operation to a particular physical address (or range of addresses) that belongs to the particular execution environment. A managing execution environment that manages the execution environments may determine those physical addresses that belong to the execution environment. By indexing the cache using the physical address, each cache may quickly determine whether the associated address is present in the cache. An alternative to this is for the cache maintenance operation to specify the execution environment for which entries should be invalidated. This requires either a lookup of the cache (which may be time consuming) or indexing of the cache according to the execution environment.

[0014] In some embodiments, the change in allocation is an allocation of memory to one of the execution environments, hi some other embodiments, the change in allocation may be a deallocation or non-allocation of memory to one of the execution environments.

[0015] In some embodiments, the maintenance operation is configured to invalidate entries in at least one encrypted storage circuit associated with an expired one of the execution environments. The invalidation may be performed when memory is reallocated (or deallocated) if / when the execution environment is terminated. By performing the invalidation when the previous execution environment is terminated, sensitive data is not kept in an unencrypted manner, which improves the security of the system.

[0016] In some embodiments, each of the execution environments is associated with a cryptographic environment identifier used to generate the key input, and the maintenance operation is configured to invalidate entries in the memory hierarchy associated with the cryptographic environment identifier. Thus, an expired execution environment may be identified in the memory hierarchy based on a cryptographic environment identifier that is unique to the execution environment that expired. Of course, in some situations, the cryptographic environment identifier may be used by multiple execution environments to enable sharing of data between those multiple execution environments. In these situations, the cryptographic environment identifier may be used in an invalidation operation when all of the execution environments expire, or when a specific one or a specific subset of the execution environments expire. As explained above, an alternative way to achieve invalidation is for the management execution environment (which knows the physical addresses assigned to each execution environment) to issue an invalidation request to the physical address associated with the execution environment whose entry is to be invalidated. This avoids the need to index the cache according to the execution environment (identifier) ​​and to painstakingly search the cache for the relevant entry.

[0017] In some embodiments, the memory address at which the memory access request is issued is a physical memory address in one of a plurality of physical address spaces, each of which is associated with one of at least two regions, such that each of the regions may have its own physical address space.

[0018] In some embodiments, the memory protection circuit defines a point of physical aliasing that is located after at least one non-aliased storage circuit in the memory hierarchy and before at least one aliased storage circuit in the memory hierarchy. The at least one non-aliased storage circuit treats physical addresses from different physical address spaces that correspond to the same memory system resource as if the physical addresses correspond to different memory system resources. Similar to a point of encryption (PoE), there may also be a point of physical aliasing (PoPA). Aliasing is prevented before the point of physical aliasing. This means that two memory accesses sent to the same physical address in different physical address spaces are treated (in the component before PoPA) as requests to different memory addresses. This may result in aliasing where the same data is stored twice in the cache under two entries.

[0019] In some embodiments, the point of physical aliasing is at or after the point of encryption, and thus there are zero or more components of the memory hierarchy that have encryption and no physical aliasing.

[0020] In some embodiments, the point of physical aliasing is at the point of encryption, and thus the point of encryption and the point of physical aliasing occur at the same point in the memory hierarchy.

[0021] In some embodiments, in response to a memory transition request requesting a transfer of memory from an origin physical address space to a destination physical address space, the maintenance operation is configured to invalidate at least some entries in at least one aliased storage circuit. When memory (e.g., at least one page) is transferred between address spaces, the maintenance operation can extend up to, but not beyond, the point of physical aliasing. Thus, it can extend beyond the point of encryption.

[0022] In some embodiments, at least some of the entries are allocated to one of at least two regions associated with the origin physical address space, and thus invalidation may be limited to entries that have been moved out of the origin physical address space.

[0023] The description of the examples below may also be relevant.

[0024] According to one example configuration, an apparatus is provided, the apparatus comprising: a processing circuit configured to perform processing in one of a fixed number of at least two regions, where one of the regions is subdivided into a variable number of execution environments; and a memory protection circuit configured to use a key input to perform encryption or decryption on data of a memory access request issued to a memory address from within a current one of the regions, where the key input is different for each of the regions and each of the execution environments, where the key input for each of the regions is fixed at boot time of the apparatus and the key input for each of the execution environments is dynamic.

[0025] Of the at least two regions, at least one of the regions (there can be several) is subdivided such that several execution environments operate within the region. A memory access request is a read or write request to data stored in a memory in the memory hierarchy. The data may be additionally cached within the memory hierarchy. However, the data is ultimately stored in memory (e.g., DRAM). A key input is used to perform encryption or decryption of data that goes to or from the memory. A key input may be considered as an input or parameter to an encryption or decryption algorithm that is kept secret to protect the confidentiality of the encrypted data. This includes the key itself, as well as parts of the key, tweakable bits, etc. The key input is different for each of the fixed number of regions. Furthermore, within at least one subdivided region, the key input is different for each execution environment. As a result, data encrypted by one execution environment cannot be accessed in another execution environment in the subdivided region unless both of those execution environments have the same key. This allows data to be kept secret between execution environments. Note that the format of the key input may be different between regions as compared to between execution environments. For example, each region may use a different key, whereas each execution environment may use the same key but with different tweakable bits. The key input for each region is selected when the device is first turned on, e.g., during the boot sequence. In contrast, the key input for each execution environment is dynamic so that it can be changed during operation of the device. In practice, the number of execution environments is variable, and key inputs can be created and deleted as needed as new execution environments are added and old execution environments are terminated. At least two regions include a subdivided realm and a root realm. In some cases, there are at least three regions, including one of a secure realm or a less secure realm (described in more detail below).

[0026] In some embodiments, the apparatus includes a memory translation circuit configured to translate memory addresses from virtual to physical memory addresses and provide an encryption environment identifier used to generate a key input, and the memory access request is forwarded from the memory translation circuit to the memory protection circuit along with the encryption environment identifier. The memory translation circuit may take the form of, for example, a Memory Management Unit (MMU). In addition to providing a translation from virtual to physical addresses (which may be an intermediate physical address), a encryption environment identifier associated with the physical address is provided. This encryption environment identifier is ultimately used to generate a key input, which is used to perform encryption (for write memory accesses) or decryption (for read memory accesses). Once the encryption environment identifier is determined, it is provided to the memory protection circuit along with the access request. The use of the encryption environment identifier means that the entire key input (such as the entire key) does not need to be provided. Instead, a smaller encryption environment identifier can be used instead, thereby reducing system overhead in bus and cache line width expansion required to carry the identifier.

[0027] In some embodiments, the memory translation circuitry is configured to store a plurality of page table entries and, in response to performing a lookup for a virtual memory address on the plurality of page table entries, indicate an encryption environment identifier. Thus, each page table entry includes an indication of the encryption environment identifier to be used to perform encryption of that particular page of memory. The indication may be the encryption environment identifier itself, or, for example, an indication as to which of the plurality of encryption environment identifiers should be used to encrypt and / or decrypt data stored in the page. The page table entry may also include access permissions indicating which regions and / or execution environments may access the memory address.

[0028] In some embodiments, the memory translation circuitry comprises a plurality of encryption environment identifier registers, the memory translation circuitry being configured to indicate which of the environment identifier registers should be used to provide the encryption environment identifier. A different encryption environment identifier is stored in each of the plurality of encryption environment identifier registers. An indicator stored in each page table entry indicates which value of the register should be provided. As a result, the number of additional bits required in the page table entry is kept small, thereby reducing the amount of storage required. By providing a plurality of registers, it is possible to enable an execution environment to use two or more different encryption environment identifiers simultaneously, each of the identifiers being used to access data at a different location. For example, an execution environment can use one encryption environment identifier to access its own private data and a second encryption environment identifier to access data shared with a second execution environment. The second execution environment can have a third encryption environment identifier to access its own private data, or the second encryption environment identifier can be the only encryption environment identifier used by the second execution environment, for example. Of course, more complex sharing schemes are possible.

[0029] In some embodiments, the encryption environment identifier is shared among a subset of the execution environments. By sharing the encryption environment identifier, multiple execution environments that share the encryption environment identifier can each access the same region of memory, thereby allowing data to be shared among the execution environments that is not accessible to other execution environments without the encryption environment identifier.

[0030] In some embodiments, the memory protection circuitry is configured to obtain a key input for one of the execution environments by performing a lookup using a cryptographic environment identifier provided by the memory access request. Thus, the memory protection circuitry may include a table or access a table using the cryptographic environment identifier to obtain a key input corresponding to that cryptographic environment identifier.

[0031] In some embodiments, the result of the lookup is a keystroke in the form of a key, such that the lookup can provide the key from the result of the lookup, in some embodiments, the lookup is performed specifically for keystrokes associated with the execution environment, rather than keystrokes associated with the domain.

[0032] In some embodiments, the result of the lookup is a key entry in the form of a contribution used to perform the encryption or decryption. As previously mentioned, the key entry may be a (secret) contribution to performing the encryption or decryption, such as an adjustment value or part of a key. In some embodiments, the lookup is performed specifically against a key entry associated with the execution environment, rather than a key entry associated with the domain.

[0033] In some embodiments, the memory address at which the memory access request is issued is a physical memory address in one of a plurality of physical address spaces, each of which is associated with one of at least two regions. Depending on the permissions provided in the underlying architecture, a region may be able to access (perhaps in a limited manner) data stored in a region with which it is not associated.

[0034] In some embodiments, the memory address at which the memory access request is issued is a physical memory address in one of a plurality of physical address spaces, each of which is associated with precisely one of the at least two regions.

[0035] In some embodiments, the at least two regions include a root region for managing switching between the at least two regions, and the plurality of physical address spaces includes a root physical address space associated with the root region that is separate from the physical address spaces associated with the plurality of other regions. Providing a dedicated root region for controlling switching can help maintain security by limiting the extent to which code running in one region can trigger a switch to another region. For example, the root region can perform various security checks when a region switch is requested. The root region has its own physical address space allocated to it, rather than using one of the physical address spaces associated with one of the other regions. Providing a dedicated root physical address space that is isolated from the physical address spaces associated with the other regions can provide stronger security guarantees for data or code associated with the root region, which may be considered paramount to security in view of managing entry to the other regions. Providing a dedicated root physical address space that is distinct from the physical address spaces of the other regions can also simplify system development, as allocation of physical addresses within each physical address space may be simplified to a specific unit of hardware memory storage. For example, by identifying a separate root physical address space, it may be simpler for data or program code associated with the root region to be preferentially stored in protected on-chip memory rather than less secure off-chip memory, and the overhead of determining which portions are associated with the root region is less than if the code or data of the root region were stored in a common address space shared with another region.

[0036] In some embodiments, the at least two regions may include at least a secure region associated with a secure physical address space and a less secure region associated with a less secure physical address space. The less secure physical address space is accessible from the less secure secure region, the secure region and the root region, and the secure physical address space is accessible from the secure region and the root region and is inaccessible from the less secure region. This therefore allows code running in the secure region to protect its code or data from access by code operating in the less secure region with stronger security guarantees than if page tables were used as the only security control mechanism. For example, parts of code requiring stronger security may be executed in a secure region managed by a trusted operating system that is separate from a non-secure operating system running in the less secure region. An example of a system supporting such secure and less secure regions may be a processing system that operates according to a processing architecture that supports TrustZone® architecture features provided by Arm® Limited (Cambridge, UK). In conventional TrustZone™ implementations, the monitor code for managing switching between secure and less secure regions uses the same secure physical address space used by the secure region. In contrast, providing a root region for managing switching between other regions, as described above, and allocating dedicated root physical address space for use by the root region helps improve security and simplify system development.

[0037] In some embodiments, all of the multiple physical address spaces are accessible from the root region. Since code running in the root region must be trusted by any party that provides code that runs in one of the other regions, and the code in the root region is involved in switching to that particular region in which that party's code is running, the root region can essentially be trusted to access any physical address space. By allowing the root region to access all of the physical address spaces, functions such as transitioning memory areas into and out of regions, for example, copying code and data into and providing services to regions during boot, can be performed. In these subexamples where multiple encryption environment identifier registers are provided, one of the registers can be used to store an encryption environment identifier associated with one of the regions that software running in the root region wishes to access. This allows software in the root region to encrypt / decrypt data in the realm regions. The root region does not have a primary MECID for root PAS access. Instead, a default MECID value of 0 is used in the root region. The root region uses an alternate MECID register 96 to store an alternate MECID for access to the realm PAS.

[0038] In some embodiments, one of the regions is a realm region that is associated with a realm physical address space, the realm physical address space being subdivided into a variable number of sub-area physical address spaces, such that each realm can be given its own physical address space within the overall realm address space.

[0039] In some embodiments, the less secure physical address space is accessible from the realm region, and the realm physical address space is accessible from the realm region and the root region, but not from the less secure region. Thus, the realm region may be considered more secure than the less secure region, but equally secure relative to the secure region.

[0040] In some of these embodiments, the secure area may be accessible from the realm area. However, in other embodiments, the realm physical address space is not accessible from the secure area and the secure physical address space is not accessible from the realm area. It is increasingly desirable to provide a secure computing environment that limits the need for software providers to trust other software providers in relation to other software running on the same hardware platform. There may be applications in several fields, such as mobile payments and banking, implementing anti-cheat or anti-piracy mechanisms in computer games, security extensions of operating system platforms, secure virtual machine hosting in cloud systems, confidential computing, etc., where the provider of software code may be unwilling to trust the provider of the operating system or hypervisor (components that may traditionally be considered trustworthy). In systems such as those based on the TrustZone® architecture described above, which support secure and less secure areas with their respective physical address spaces, as secure components operating in the secure area become more and more prevalent, the set of software that typically operates in the secure area will expand to include several software that may be provided by a different number of software providers, including, for example, the following parties: An original equipment manufacturer (OEM) that assembles a processing device (such as a mobile phone) from components including silicon integrated circuit chips provided by a particular silicon provider, an operating system vendor (OSV) that provides the operating system that runs on the device, and a cloud platform operator (or cloud host) that maintains a server farm that provides the server space to host virtual machines on the cloud.Thus, if the realms were implemented in a strict order of increasing privilege, problems could arise because application providers providing application-level code wishing to provide a secure computing environment may not want to trust parties (such as OSVs, OEMs, cloud hosts, etc.) that may traditionally have provided the software that runs in the secure realm, but similarly, parties providing code that runs in the secure realm are unlikely to want to trust application providers providing code that runs in a higher privilege realm that is given access to data associated with a less privileged realm. These embodiments recognize that a strict hierarchy of realms of successively increasing privilege may not be appropriate, and thus realm realms may be considered orthogonal to the secure realm, although realm realms and secure realms each have access to a less secure physical space, and neither realm can access the other's physical space.

[0041] In some embodiments, the less secure physical address space is accessible from all of at least two regions. This is useful because it facilitates sharing of data or program code between software running in different regions. If a particular item of data or code is to be accessible in different regions, it can be allocated in the less secure physical address space so that it can be accessed from any region.

[0042] The description of the examples below may also be relevant.

[0043] In some embodiments, an apparatus is provided comprising: a processing circuit configured to execute processing in one of a fixed number of at least two regions, where one of the regions is subdivided into a variable number of execution environments; a memory translation circuit configured to determine a given encryption environment identifier associated with one of the execution environments in response to a memory access request to a given memory address and to forward the memory access request with the given encryption environment identifier; and a storage circuit configured to store a plurality of entries each associated with an associated encryption environment identifier and an associated memory address, the storage circuit including a determination circuit configured to determine, in at least one enabled operating mode, whether the given encryption environment identifier differs from an associated encryption environment identifier associated with one of the entries associated with the given memory address.

[0044] The at least two regions may be at least three regions. For example, these may include a secure region, a non-secure region (which does not mean no security, just less security than the secure region), and a realm region, where the realm region may be a region subdivided into multiple execution environments. Access to resources between different regions may be controlled. For example, data belonging to a process executing in a secure region may not be accessible to resources operating in a non-secure region. On the other hand, resources in a secure region may not be accessed by resources operating in a realm region (and vice versa), and both the realm region and the secure region may access resources in a non-secure region. In these embodiments, each of the execution environments operating in the subdivided realms has an associated encryption environment identifier that identifies an area of ​​memory that may be used to store resources used by the execution environment in an encrypted manner. In this way, resources belonging to each execution environment may be isolated and protected from each other. A storage circuit (e.g., a cache) may be used to store entries (e.g., cache lines). Each cache line may be associated with an execution environment identifier, i.e., a cache line is not necessarily encrypted, but may be stored unencrypted in association with its encryption environment identifier. Each of the entries (e.g., cache lines) also has an associated memory address (e.g., a location in an area of ​​memory to which the cache line pertains). When a memory access request is issued on behalf of one of the execution environments, the memory access request obtains an encryption environment identifier associated with that execution environment. The memory access request then travels through the memory hierarchy toward the main memory. In some cases, the data to which the memory access request pertains is already stored in a storage circuit (e.g., a cache). Under normal circumstances, in order for the data to be returned, the encryption environment identifier associated with the memory access request is required to match the encryption environment identifier associated with the entry of the storage circuit that contains the data to which the memory access request is issued.In these embodiments, a decision circuit is provided to make that decision. The device may be capable of switching between enabled operating modes, or the current mode may be fixed.

[0045] In some embodiments, the device comprises a memory protection circuit configured to use a key input to perform encryption or decryption on data of a memory access request in response to the data not being present in the storage circuit, the key input being based on a given encryption environment identifier. The key input for each of the regions is fixed at boot time of the device, and the key input for each of the execution environments is dynamic. The encryption environment identifier is used by the memory protection circuit to determine a key input (e.g., a key, a portion of a key, one or more tunable bits). The key input is used to achieve encryption and decryption for the execution environments and the regions. It is therefore different for each region and for each execution environment. The key input for a region is fixed when the device boots up. On the other hand, the key input for an execution environment is dynamically determined because the execution environments may start, stop, and change dynamically.

[0046] In some embodiments, the storage circuitry is configured to, in at least one error mode of operation, execute an error action in response to a given cryptographic environment identifier differing from an associated cryptographic environment identifier associated with a given memory address. In these embodiments, the device responds to a mismatch when a mismatch occurs, the response to the mismatch being to execute the error action. In situations where a mismatch does not occur, the memory access request proceeds normally.

[0047] In some embodiments, the at least one enabled operating mode includes a poison operating mode in which the storage circuitry is configured to poison one of the entries in response to an associated memory address of one of the entries being a given memory address when the given cryptographic environment identifier differs from an associated cryptographic environment identifier associated with one of the entries. The entry is poisoned, thereby making the entry a poisoned entry. That is, some or all of the entry generates an error if and when it is subsequently consumed by the processing circuitry. By poisoning the entry and postponing any errors that may occur, it is possible to prevent reading of data that may be private (in the case of a read request) or use of corrupted data (in the case of a write request). However, if the data is never read from it again, no error will occur. Under-consumption may occur, for example, in the case of pre-fetching or speculative execution. In the case of a memory access request that is a memory read request, the data may be consumed almost immediately, leading to an immediate error.

[0048] In some embodiments, the processing circuitry is configured to generate an exception when one of the entries is received by the processing circuitry, such that the poisoned entry remains unusable by the processing circuitry.

[0049] In some embodiments, when the memory access request is a write memory access request, a portion of one of the entries that is accessed by the (mismatched) write memory access request is modified and the remaining portion of one of the entries is poisoned. A write request may modify only a portion of one of the entries (e.g., a cache line) of a storage circuit (e.g., a cache). In these situations, the portion of the entry that is intended to be modified by the memory access request is modified, while the other portion of the entry that is not intended to be modified by the access request remains intact but is poisoned so that an exception will be generated by the processing circuit if it is accessed in the future.

[0050] In some embodiments, the at least one enabled operating mode includes an alias operating mode in which the storage circuit is configured to treat entries of the storage circuit as different when the associated memory addresses of the entries match and when the associated cryptographic environment identifiers of the entries mismatch. In these embodiments, rather than generating an error in response to a mismatch, each entry effectively treats the cryptographic environment identifier as part of its address. Two entries to the same address with different cryptographic environment identifiers are treated as two separate and distinct entries. Thus, if memory access requests to one memory address with one cryptographic environment identifier are attempting to access different data items to entries associated with the same memory address and different cryptographic environment identifiers, a mismatch "cannot" occur.

[0051] In some embodiments, the at least one enabled operating mode includes a delete operating mode in which the storage circuitry is configured to delete and invalidate one of the entries in response to an associated memory address of one of the entries being a given memory address when the given encryption environment identifier differs from an associated encryption environment identifier associated with one of the entries. In these embodiments, the mismatch is handled by writing an existing entry in the storage circuitry back to a point in the memory hierarchy where encryption is performed using the encryption environment identifier (e.g., main memory). The entry in the storage circuitry is then invalidated so that it cannot be accessed.

[0052] In some embodiments, in a delete mode of operation, in response to the associated memory address of one of the entries being a given memory address when the given cryptographic environment identifier differs from an associated cryptographic environment identifier associated with one of the entries, the storage circuitry is further configured to treat the memory access request as a miss in the storage circuitry. Writing back (deleting) the data and then invalidating it, the memory access request can be treated as missing in the storage circuitry. Thus, the request can be reissued to the memory hierarchy, and the requested data is eventually returned to the storage circuitry for storage.

[0053] In some embodiments, the processing circuitry is configured to speculatively issue memory access requests as speculative read requests while in the speculative mode of operation, and the speculative mode of operation is disabled unless the memory circuitry is in the enabled mode of operation. For example, a speculative read request generated when speculating on the outcome of a branch instruction may occur using an incorrect cryptographic environment identifier. For example, a speculative read request may be permitted for a regular memory (e.g., DRAM) address that has a valid MMU mapping. However, such an access should not be permitted if the MECID is "wrong" for that location. A conventional hypervisor may have its own virtual address mapping for all of the DRAM, independent of the mappings of virtual machines monitored by that hypervisor. For a realm with a MECID, if any realm manager had a mapping to all of the DRAM, the CPU could infer those addresses regardless of whether the "correct" realm MECID value existed. As a result, a secure system should either prevent speculative read requests from occurring or implement one of the aforementioned enabled modes in which mismatches between cryptographic environment identifiers are caught (or prevented entirely).

[0054] In some embodiments, the at least one enabled operational mode includes an erase operational mode in which the storage circuitry is configured to perform an erase of one of the entries in response to an associated memory address of one of the entries being a given memory address when the given encryption environment identifier differs from an associated encryption environment identifier associated with one of the entries. Erasing an entry differs from invalidation, where data is simply marked as invalid and inaccessible, in that data actually stored in the storage circuitry is removed. There are several ways in which this may be accomplished.

[0055] In some embodiments, the storage circuit is configured to perform the erasure by zeroing or randomizing one of the entries. By zeroing the data, the data is replaced by a predetermined sequence (typically a bit "0", but the use of a bit "1" can also be referred to as "zeroing"), where the predetermined sequence has no appreciable meaning. Another alternative is to scramble or randomize the data of the entry. In either case, the original meaning of the data is removed and can no longer be determined.

[0056] In some embodiments, when the memory access request is a write memory access request, the storage circuitry is further configured to update the associated cryptographic identifier to correspond to the given cryptographic environment identifier. In addition to the actions described above, an operation to write to a particular entry where a mismatch occurs may cause the cryptographic environment identifier associated with that entry to be overwritten by the cryptographic environment identifier associated with the memory access request.

[0057] In some embodiments, the decision circuit is configured to prevent, in at least one disabled operating mode, a determination of whether a given cryptographic environment identifier is different from an associated cryptographic environment identifier associated with one of the entries associated with a given memory address. The disabled operating mode thereby disables or prevents mismatch detection from taking place. Thus, any mismatch that may occur is essentially ignored. This may therefore result in the plaintext of any entry of the storage circuit being leaked to other execution environments. In practice, there may be other protection mechanisms that prevent this from happening. For example, the management system may prevent a memory access request associated with another execution environment from being issued.

[0058] In some embodiments, in at least some of the at least one enabled operating mode, in response to an associated memory address of one of the entries being a given memory address when the given cryptographic environment identifier differs from an associated cryptographic environment identifier associated with one of the entries, the storage circuitry is configured to generate an asynchronous exception. With an asynchronous exception, the exception is out of sync with the code leading to the exception; that is, an exception may be raised; however, it may not be handled until a (possibly non-deterministic) time later. However, this allows for debugging or detection of leaked data. There are several ways in which an asynchronous exception may be raised.

[0059] In some embodiments, the storage circuitry is configured to generate an asynchronous exception and store details of the memory access request in one or more registers accessible to the processing circuitry. For example, data related to the memory access request may be stored in these registers, such as the memory address at which the access is made, the type of access (read or write), the execution environment in which the memory access request is made, and the execution environment with which the memory address is associated in the storage circuitry. This may be used to achieve debugging and / or detection of conditions that lead to a cryptographic environment mismatch.

[0060] Specific embodiments will now be described with reference to the drawings.

[0061] FIG. 1 illustrates in schematic form an example of a data processing system 2 having at least one requester device 4 and at least one completer device 6. An interconnect 8 provides communication between the requester device 4 and the completer device 6. A requester device may issue a memory access request that requests memory access to a particular addressable memory system location. A completer device 6 is a device that is responsible for processing memory access requests directed to it. Although not illustrated in FIG. 1, some devices may be capable of functioning as both a requester device and a completer device. The requester device 4 may include, for example, a processing element such as a central processing unit (CPU) or a graphics processing unit (GPU), or other master devices such as bus master devices, network interface controllers, display controllers, etc. A completer device may include a memory controller responsible for controlling access to a corresponding memory storage device, a peripheral controller for controlling access to peripheral devices, etc. Although FIG. 1 illustrates an exemplary configuration of one of the requester devices 4 in more detail, it will be understood that the other requester devices 4 may have a similar configuration. Alternatively, other requestor devices may have a different configuration than requestor device 4 shown on the left side of FIG.

[0062] The requester device 4 has a processing circuit 10 for performing data processing responsive to instructions by referencing data stored in registers 12. The registers 12 may include general purpose registers for storing operands and results of processed instructions, as well as control registers for storing control data for configuring how processing is performed by the processing circuit. For example, the control data may include a current domain indication 14 used to select which operational domain is the current domain, and a current exception level indication 15 indicating which exception level is the current exception level at which the processing circuit 10 is operating.

[0063] Processing circuitry 10 may be capable of issuing memory access requests that specify a virtual address (VA) that identifies an addressable location to be accessed, and a region identifier (region ID or “security state”) that identifies the current region. Address translation circuitry 16 (e.g., a memory management unit (MMU)) translates the virtual address to a physical address (PA) through one of many stages of address translation based on page table data defined in a page table structure stored in the memory system. Translation lookaside buffer (TLB) 18 functions as a lookup cache to cache a portion of the page table information for faster access than if the page table information had to be fetched from memory each time an address translation is needed. In this embodiment, in addition to generating a physical address, address translation circuitry 16 also selects one of a number of physical address spaces associated with the physical address, outputs a physical address space (PAS) identifier that identifies the selected physical address space, and also provides a MECID, the purpose of which is explained in more detail below.

[0064] The PAS filter 20 functions as a requester filtering circuit to ascertain, based on the translated physical address and the PAS identifier, whether the physical address is permitted to be accessed within the physical address space identified and specified by the PAS identifier. This lookup is based on granule protection information stored in a granule protection table structure stored within the memory system. The granule protection information may be cached within a granule protection information cache 22, similar to the caching of page table data within the TLB 18. In the example of FIG. 1, the granule protection information cache 22 is shown as a structure separate from the TLB 18, but in other examples, these types of lookup caches may be combined as a single lookup cache structure, such that a single lookup of an entry in the combined structure provides both page table information and granule protection information. The granule protection information defines information that limits the physical address space in which a given physical address may be accessed, and based on this lookup, the PAS filter 20 determines whether to allow the memory access request to proceed to be issued to one or more of the caches 24 and / or the interconnect 8. If the specified PAS for a memory access request is not permitted to access the specified physical address, the PAS filter 20 may block the transaction and signal a failure.

[0065] Although FIG. 1 shows an example of a system having multiple requester devices 4, the features shown for a single requesting device on the left side of FIG. 1 may also be included in a system where there is only one requester device, such as a single core processor.

[0066] While FIG. 1 shows an example in which the selection of a PAS for a given request is performed by the address translation circuitry 16, in other examples, information for determining which PAS to select along with the PA can be output by the address translation circuitry 16 to the PAS filter 20, which can select a PAS and verify whether the PA can be accessed within the selected PAS.

[0067] The provision of PAS filter 20 helps support a system that can operate in several operating domains, each associated with its own isolated physical address space, where for at least a portion of the memory system (e.g., for some caches or coherency enforcement mechanisms such as snoop filters), separate physical address spaces are treated as if they refer to separate sets of addresses that identify entirely separate memory system locations, even if the addresses in those address spaces actually refer to the same physical location in the memory system. This can be useful for security purposes.

[0068] FIG. 2 shows examples of different operating states and domains in which processing circuit 10 can operate, as well as examples of the types of software that may be executed at different exception levels and domains (it will, of course, be understood that the particular software installed on a system is selected by the party in charge of that system, and therefore is not an essential feature of the hardware architecture).

[0069] The processing circuit 10 is operable at a number of different exception levels 80, in this embodiment four exception levels labeled EL0, EL1, EL2, and EL3, where in this embodiment EL3 refers to the most privileged exception level and EL0 refers to the least privileged exception level. It will be appreciated that other architectures may choose the reverse numbering, with the exception level having the highest number being considered to be the least privileged. In this embodiment, the least privileged exception level EL0 is for application level code, the next most privileged exception level EL1 is used for operating system level code, the next most privileged exception level EL2 is used for hypervisor level code that manages switching between several virtual operating systems, and the most privileged exception level EL3 is used for monitor code that manages switching between the respective domains and the allocation of physical addresses to the physical address space, as will be described below.

[0070] When an exception occurs at a particular exception level while processing the software, for some types of exceptions, the exception is accepted to a higher (more privileged) exception level, and the particular exception level at which the exception is accepted is selected based on attributes of the particular exception that occurred. However, in some circumstances, other types of exceptions may be accepted at the same exception level as the exception level associated with the code that was being processed when the exception was accepted. When an exception is accepted, information that characterizes the state of the processor at the time the exception was accepted may be saved, including, for example, the current exception level at the time the exception was accepted. Thus, once an exception handler has been processed to address the exception, processing may return to the previous processing, and the saved information may be used to identify the exception level to which processing should return.

[0071] In addition to the different exception levels, the processing circuit also supports a number of operating regions, including a root region 82, a secure (S) region 84, a less secure region 86, and a realm region 88. For ease of reference, the less secure region is described below as a "non-secure" (NS) region, although it will be understood that this is not intended to imply a particular level (or lack thereof) of security. Instead, "non-secure" simply indicates that the non-secure region is intended for code that is less secure than code operating in the secure region. The root region 82 is selected when the processing circuit 10 is at the highest exception level EL3. When the processing circuit is at one of the other exception levels EL0-EL2, the current region is selected based on a current region indicator 14, which indicates which of the other regions 84, 86, 88 is active. For each of the other regions 84, 86, 88, the processing circuit may be at either exception level EL0, EL1, or EL2.

[0072] At boot time, a number of boot codes (e.g., BL1, BL2, OEM boot) may be executed, for example, within the more highly privileged exception levels EL3 or EL2. Boot codes BL1, BL2 may be associated, for example, with the root region, and OEM boot code may operate in the secure region. However, once the system is booted, during execution, processing circuitry 10 may be considered to operate in one of regions 82, 84, 86, and 88 at a time. Each of regions 82-88 is associated with its own associated physical address space (PAS). This allows for isolation of data from different regions within at least a portion of the memory system, as will be described in more detail below.

[0073] Non-secure area 86 may be used for normal application level processing, and operating system and hypervisor activity for managing such applications. Thus, within non-secure area 86 there may be application code 30 running at EL0, operating system (OS) code 32 running at EL1, and hypervisor code 34 running at EL2.

[0074] The secure region 84 allows certain system-on-chip security, media, or system services to be isolated in a separate physical address space from the physical address space used for non-secure processing. The secure and non-secure regions are not equivalent in the sense that the secure region can access both secure and non-secure resources while the non-secure region code cannot access resources associated with the secure region 84. An example of a system that supports such a partition of secure and non-secure regions 84, 86 is a system based on the TrustZone® architecture provided by Arm® Limited. The secure region can run trusted applications 36 at EL0, trusted operating systems 38 at EL1, and optionally a secure partition manager 40 at EL2. EL2 can use two page tables to support isolation between different trusted operating systems 38 running within the secure region 84, in a manner similar to how the hypervisor 34 can manage isolation between virtual machines or guest operating systems 32 running within the non-secure region 86.

[0075] Extending systems to support secure areas 84 has become common in recent years because it allows a single hardware processor to support isolated secure processing and avoids the need for processing to occur on a separate hardware processor. However, as the use of secure areas has grown in popularity, many practical systems with such secure areas now support a relatively highly mixed environment of services provided within the secure area by a wide range of different software providers. For example, the code running within the secure area 84 may include different software, including (among other things) silicon providers that have manufactured integrated circuits, original equipment manufacturers (OEMs) that assemble integrated circuits provided by the silicon providers into electronic devices such as mobile phones, operating system vendors (OSVs) that provide operating systems 32 for the devices, and / or cloud platform providers that manage cloud servers that support services for a number of different customers via the cloud.

[0076] However, there is an increasing demand for providers of user-level code (which may normally be expected to run as applications 30 in the non-secure area 86) to be provided with a secure computing environment that can be trusted not to leak information to other parties running code on the same physical platform. Such a secure computing environment may desirably be dynamically allocable during execution and guaranteed and provable so that a user can verify whether sufficient security guarantees are provided on the physical platform before entrusting the device with processing potentially sensitive code or data. Users of such software may not want to trust providers of feature-rich operating systems 32 or hypervisors 34 that may normally run in the non-secure area 86 (or, even if those providers themselves are trustworthy, users may want to protect themselves from the operating systems 32 or hypervisors 34 being compromised by attackers). And while the secure area 84 can be used for such user-provided applications that require secure processing, in practice this creates problems for both users who provide code that requires a secure computing environment and providers of existing code that runs in the secure area 84. For providers of existing code running in secure area 84, the addition of arbitrary user-provided code within the secure area increases the attack surface for potential attacks against their code. This may be undesirable, and therefore users may be strongly discouraged from allowing code to be added to secure area 84. On the other hand, users providing code that require a secure computing environment may be reluctant to trust access to their data or code to all of the different code providers running in secure area 84, as it may be difficult to audit and certify all of the separate code provided by different software providers running in secure area 84, if certification and attestation of the code running in a particular area is required as a prerequisite for the user-provided code to perform operations.This may limit the opportunities for third parties to offer more secure services.

[0077] Thus, as shown in FIG. 2, an additional area 88, called a realm area, is provided that may be used by such user-introduced code to provide a secure computing environment orthogonal to any secure computing environment associated with components operating in the secure realm 24. In the realm area, the software executing may include any number of realms (or execution environments), each of which may be isolated from the other realms by a realm management module (RMM) 46 operating at exception level EL2. The RMM 46 may control the isolation between the respective realms 42, 44 executing in the realm area 88 by, for example, defining access permissions and address mappings in page table structures, similar to how the hypervisor 34 manages the isolation between different components operating in the non-secure realm 86. In this example, the realms include an application-level realm 42 that executes at EL0 and an encapsulated application / operating system realm 44 that executes across exception levels EL0 and EL1. It will be appreciated that it is not mandatory to support both EL0 and EL0 / EL1 type realms, and multiple realms of the same type may be established by the RMM 46.

[0078] The realm region 88, like the secure region 84, has its own physical address space assigned to it, but is orthogonal to the secure region 84 in the sense that the realm region and the secure region 88, 84 can each access the non-secure PAS associated with the non-secure region 86, and the realm region and the secure region 88, 84 cannot access each other's physical address space. This means that the code running in the realm region 88 and the secure region 84 have no dependencies on each other. Code in the realm region only needs to trust the hardware, the RMM 46, and the code running in the root region 82 that manages the switching between the regions, which means that proofs and guarantees are more feasible. Proofs allow a given software to request verification that the code installed on the device matches certain expected characteristics. This can be done by checking whether a hash of the program code installed on the device matches an expected value that is signed by a trusted party using a cryptographic protocol. The RMM 46 and monitor code 29 may be attested, for example, by checking that a hash of this software matches an expected value signed by a trusted party, such as a silicon provider who manufactured the integrated circuit that includes the processing system 2, or an architecture provider who designed a processor architecture that supports region-based memory access control. This allows the user-provided code 42, 44 to verify that the integrity of the region-based architecture can be trusted before performing any secure or sensitive functions.

[0079] Thus, as shown by the dotted lines indicating the gaps in the non-secure realm where these processes would have previously been executed, it can be seen that code associated with realms 42, 44 that would previously have been executing in the non-secure realm 86 can now be moved to a realm realm where they may have stronger security guarantees since their data and code are not accessible by other code running in the non-secure realm 86. However, due to the fact that the realm realm 88 and the secure realm 84 are orthogonal and therefore cannot see each other's physical address space, this means that providers of code in the realm realm do not need to trust providers of code in the secure realm, and vice versa. Code in the realm realm can simply trust the firmware that provides the root realm 82 and the monitor code 29 of the RMM 46, which may be provided by the silicon provider, or by the provider of the instruction set architecture supported by the processor. These providers may need to be inherently trusted from the start when the code is running on their device, so that no other further trust relationships with other operating system vendors, OEMs, or cloud hosts are required by the user in order for the user to be provided with a secure computing environment.

[0080] This may be useful, for example, for a variety of purpose applications and use cases including, for example, mobile wallet and payment applications, fraud and piracy prevention mechanisms in gaming, operating system platform security extensions, secure virtual machine hosting, confidential computing, networking, or gatewaying for the Internet of Things. It will be appreciated that users may find many other applications in which Realm support is useful.

[0081] To support the security assurances provided to a realm, a processing system may support an attestation reporting feature whereby firmware images and configurations, e.g., monitor code images and configurations, or RMM code images and configurations are measured at boot time or during run time. Realm contents and configurations are measured during run time, allowing a realm owner to trace back relevant attestation reports to known implementations and assurances, and make trust decisions about whether to operate on that system.

[0082] As shown in FIG. 2, a separate root region 82 is provided to manage region switching, with the root region having its own isolated root physical address space. Creating a root region and isolating resources from the secure region allows for a more robust implementation, even for systems with only non-secure and secure regions 86, 84 and no realm region 88, but can also be used for implementations that support realm region 88. The root region 82 can be implemented using monitor software 29 provided (or certified) by the silicon provider or architect and can be used to provide secure boot functionality, trusted boot measurements, system-on-chip configuration, debug control and firmware update management for firmware components provided by other parties, such as OEMs. Code in the root region can be developed, certified and deployed by the silicon provider or architect without any dependency on the final device. In contrast, the secure region 84 can be managed by the OEM to implement certain platform and security services. Management of the non-secure realm 86 may be controlled by the operating system 32, which provides operating system services, while the realm realm 88 is isolated from the existing secure software environment in the secure realm 84 while allowing the development of new forms of trusted execution environments that may be dedicated to user or third-party applications.

[0083] FIG. 3 shows a schematic of another example of a processing system 2 for supporting these techniques. Elements that are the same as in FIG. 1 are indicated with the same reference numerals. FIG. 3 shows more details of the address translation circuit 16, including a stage 1 memory management unit 50, and a stage 2 memory management unit 52. The stage 1 MMU 50 may be involved in translating either virtual addresses to physical addresses (when translation is triggered by EL2 or EL3 code) or to intermediate physical addresses (when translation is triggered by EL0 or EL1 code, with further stage 2 translation by the stage 2 MMU 52 being necessary). The stage 2 MMU may translate the intermediate physical addresses to physical addresses. The stage 1 MMU may be based on page tables controlled by the operating system for translations starting from EL0 or EL1, on page tables controlled by the hypervisor for translations from EL2, or on page tables controlled by the monitor code 29 for translations from EL3. Alternatively, stage 2 MMU 52 may be based on page table structures defined by hypervisor 34, RMM 46 or secure partition manager 14, depending on which region is being used. Separating the translation into two stages in this way allows operating systems to manage address translation for themselves and for applications, under the assumption that they are the only operating systems running on the system, while RMM 46, hypervisor 34 or SPM 40 can manage isolation between different operating systems running within the same region.

[0084] As shown in Figure 3, the address translation process using the address translation circuitry 16 can return security attributes 54 that, in combination with the current exception level 15 and the current region 14 (or security state), allow a section of a particular physical address space (identified by a PAS identifier or "PAS TAG") to be accessed in response to a given memory access request. The physical address and PAS identifier can be looked up in a granular protection table 56 that provides the granular protection information described above, or this can originate from the address translation circuitry. In this example, the PAS filter 20 is shown as a granular memory protection unit (GMPU) that verifies whether the selected PAS is authorized to access the requested physical address, and if so, allows the transaction to be passed to any caches 24 or interconnects 8 that are part of the system fabric of the memory system.

[0085] The GMPU 20 allows memory to be allocated to separate address spaces while at the same time providing strong hardware-based isolation guarantees, providing efficient sharing schemes as well as spatial and temporal flexibility in how physical memory is allocated to these address spaces. As previously mentioned, the execution units in the system are logically divided into virtual execution states (regions or "worlds"), of which there is one execution state (root world) called the "root world" and located at the highest exception level (EL3), and the root world manages the allocation of physical memory to these worlds.

[0086] A single system physical address space is virtualized into multiple "logical" or "architectural" physical address spaces (PAS), where each such PAS is an orthogonal address space with independent coherency properties. A system physical address is mapped into a single "logical" physical address space by extending it with a PAS tag.

[0087] A given world is allowed access to a subset of the logical-physical address space. This is enforced by a hardware filter 20 that can be attached to the output of the memory management unit 16.

[0088] The world defines the security attributes of the access (PAS tag) using fields in the translation table descriptor of the page table used for address translation. The hardware filter 20 has access to a table (granule protection table 56, or GPT) that defines granule protection information (GPI) for each page in the system physical address space, which indicates the PAS tag with which it is associated, and (optionally) other granule protection attributes.

[0089] The hardware filter 20 checks the world ID and security attributes for the GPI of a granule to determine if access can be granted, thus forming a Granular Memory Protection Unit (GMPU).

[0090] GPT56 can reside, for example, in on-chip SRAM or off-chip DRAM. If stored off-chip, GPT56 can be integrity protected by an on-chip memory protection engine, which can use encryption, integrity, and freshness mechanisms to maintain the security of GPT56.

[0091] Locating the GMPU 20 at the requester side of the system (e.g., on the MMU output) rather than at the completer side allows the Interconnect 8 to assign access permissions at page granularity while allowing hashing / striping of pages continuously across multiple DRAM ports.

[0092] The transaction remains tagged with the PAS TAG as it propagates throughout the system fabric 24, 8 until it reaches a location defined as a point 60 of the physical alias. This allows the filter to be localized on the requester r without weakening security guarantees compared to completer filtering. As the transaction propagates throughout the system, the PAS TAG can be used as a security mechanism in-depth for address isolation. For example, a cache can add the PAS TAG to an address tag in the cache to prevent accesses made with an incorrect PAS TAG for the same PA from hitting the cache, thereby improving side-channel resistance. The PAS TAG can also be used as a context selector for a protection engine attached to the memory controller 68 that encrypts data before it is written to external DRAM.

[0093] The Point of Physical Aliasing (PoPA) is the location in the system where the PAS TAG is stripped and addresses are converted from logical physical addresses back to system physical addresses. The PoPA can be located below the cache on the completer side of the system, where accesses to the physical DRAM are made (using the cryptographic context resolved via the PAS TAG). Alternatively, it may be located above the cache to simplify the system implementation at the expense of weakening security.

[0094] At any point, a world may request to transition a page from one PAS to another. The request is made at EL3 to monitor code 29, which inspects the current state of the GPI. EL3 may allow only a specific set of transitions to occur (e.g., non-secure PAS to secure PAS, not realm PAS to secure PAS). To provide a clean transition, a new instruction, "Delete data and invalidate to point of physical alias", is supported by the system and EL3 may submit it before transitioning the page to the new PAS. This ensures that any residual state associated with the previous PAS is flushed from any caches upstream of PoPA 60 (closer to the requester).

[0095] Another property that can be achieved by attaching the GMPU 20 to the master side is efficient sharing of memory between worlds. It may be desirable to allow shared access to a physical granule to a subset of the N worlds while preventing other worlds from accessing it. This can be achieved by adding a "limited sharing" semantic to the granule protection information and forcing it to use a specific PAS TAG. As an example, a GPI can indicate that a physical granule can only be accessed by the "realm world" 88 and the "secure world" 84 while it is tagged with the PAS TAG of the secure PAS 84.

[0096] The example properties above result in rapid changes in the visibility properties of a particular physical granule. Consider the case where each world is assigned a private PAS that is only accessible to that world. For a particular granule, a world can request that it be made visible to the non-secure world at any time, without changing the PAS association, by changing their GPI from "exclusive" to "limitedly shared with non-secure world". In this way, the visibility of that granule can be increased without requiring costly cache maintenance or data copy operations.

[0097] Also shown in the embodiment of FIG. 3 is a MECID consumer 64, which together with the PAS TAG stripper 60 collectively form a memory protection circuit 62. The MECID consumer 64 consumes MECIDs provided by the memory translator 16, each of which is associated with a different realm or execution environment. The MECID consumer 64 provides a key input that is used to encrypt data past the point of encryption (PoE) based on the MECID. This encryption may be separate from the encryption performed based on the PAS. Thus, each realm (which may each be associated with a different MECID) can individually encrypt its own data such that the data cannot be accessed by other realms. Thus, even if there is an error, misconfiguration, or attack on the RMM 46 that allows one realm to access the physical address space of another realm, the data belonging to the other realms has no meaning to it.

[0098] Note that in this embodiment, PoE and MECID consumer 64 are shown as being combined with PAS TAG stripper 60. In principle, PoE could be anywhere between the provider of MECID (e.g., address translation circuit 16) and PoPA 60, and the two elements 60, 64 could be executed sequentially rather than together. Elements of the memory hierarchy occurring between the requester device and PoE store data in an unencrypted manner and with a corresponding MECID. In contrast, other elements of the memory hierarchy (past PoE) store data in an encrypted manner without a corresponding MECID.

[0099] FIG. 4 illustrates how the system physical address space 64 can be divided into chunks allocated for access within a particular architectural physical address space 61 using a granule protection table 56. The granule protection table (GPT) 56 defines which portions of the system physical address space 65 are accessible from each architectural physical address space 61. For example, the GPT 56 can contain a number of entries, each corresponding to a granule of a particular size of physical address (e.g., 4K pages), and can define the assigned PAS for that granule, which can be selected from among non-secure, secure, realm, and root regions. By design, if a particular granule or set of granules is assigned to a PAS associated with one of the regions, it can only be accessed within the PAS associated with that region and not within the PAS of the other regions. However, note that even though a granule allocated to (for example) the secure PAS cannot be accessed from within the root PAS, the root region 82 can still access that granule of physical addresses by specifying in its page table PAS selection information to ensure that virtual addresses associated with pages mapped to that area of ​​physically addressed memory are translated to physical addresses in the secure PAS instead of the root PAS. Thus, sharing of data between regions (to the extent permitted by the accessibility rules defined in the table above) can be controlled at the time of selecting a PAS for a given memory access request.

[0100] However, in some implementations, in addition to allowing access to granules of physical addresses within the assigned PAS defined by the GPT, the GPT can use other GPT attributes to mark certain regions of an address space (e.g., address space associated with a region of lower or orthogonal privilege that would not normally be permitted to select the assigned PAS for an access request for that region) as shared with another address space. This can facilitate temporary sharing of data without having to change the assigned PAS for a given granule. For example, in FIG. 4, realm PAS region 70 is defined in the GPT as being assigned to the realm region and is normally inaccessible from non-secure region 86 because non-secure region 86 cannot select the realm PAS for its access request. Non-secure code would not normally be able to see data in region 70 because non-secure region 26 does not have access to the realm PAS. However, if a realm wishes to temporarily share some of its data in its allocated region of memory with the non-secure region, it can request that monitor code 29 running in root region 82 update GPT 56 to indicate that region 70 is shared with non-secure region 86, thereby making region 70 accessible from the non-secure PAS shown on the left side of FIG. 4 without having to change which regions are allocated to region 70. When a realm region indicates a region of its address space as being shared with the non-secure region, a memory access request issued from the non-secure region and targeting that region may initially specify the non-secure PAS, but PAS filter 20 can remap the PAS identifier of the request to instead specify the realm PAS. Downstream memory system components will then treat the request as if it had originally issued from the realm region. This sharing can improve performance because the operations to assign different regions to specific memory regions may be more performance intensive, involving a higher degree of cache / TLB invalidation and / or data zeroing in memory or copying data between memory regions.This may not be justifiable where sharing is expected to be only temporary.

[0101] A portion of the realm PAS is assigned to each realm currently active in the system. Access to these sub-areas 90, 92 within the realm PAS may be restricted / controlled by the RMM 46 as previously described. In addition to this, however, the contents of each sub-area 90, 92 may be encrypted differently depending on the realm associated with that sub-area within the realm PAS. For example, a first sub-area 90 is associated with a first realm R0 and therefore the data within that sub-area is encrypted in a different manner than the contents of a second sub-area 92 associated with a second realm R1. In addition to this, each PAS is encrypted. In the example of FIG. 4, each PAS is encrypted using a different key and then the sub-areas are further encrypted using different individual keys. In other embodiments, the realm area itself may not have one overall key and instead the individual realms themselves may be encrypted. Each realm / execution environment has its own encrypted area of ​​the PAS that other execution environments (realms) cannot access. In addition, as a result of the encryption, a realm cannot access a secure realm or a root realm.

[0102] Note that a memory access request is given a PAS along with the physical address. As a result, the memory treats two requests to the same physical address (with different PASs) as requests to different physical addresses, even though the same physical address is actually being accessed. This "aliasing" is important because it provides more secure isolation of the PASs. As a result, cache timing attacks, in which inferences are made about privileged data based on whether it is present in the cache, become infeasible, if not impossible. Data stored in the cache of one PAS is not accessible to another PAS.

[0103] FIG. 5 summarizes the operation of the address translation circuit 16 and the PAS filter. The PAS filtering 20 can be considered as an additional stage 3 check performed after the stage 1 (and optionally stage 2) address translation performed by the address translation circuit. Also note that the EL3 translation is based on the page table entry providing a selection information (labeled NS, NSE in the example of FIG. 5) based on two bits of address, while the single bit of selection information "NS" is used to select the PAS of the other state. The security state shown in FIG. 5 as an input to the granule protection check refers to a region ID that identifies the current region of the processing element 4. The MECID is provided by the stage 2 MMU for EL0 and EL1, whereas the MECID is provided by the stage 1 MMU for software executing at EL2 and EL3.

[0104] In practice, it is not necessary for the translation table to directly store the MECID, and in fact, doing so would dramatically increase the size of the page table entry. Instead, as shown in FIG. 6, each entry 98 of the page table in the address translation circuit 16 includes attributes 100 that can indicate access permissions, memory type, access, and dirty status, etc. A PAS indicator field 102 is used to indicate which PAS is used for the entry. For the EL3 stage 1 translation table, the NS and NSE bits are used to define the PAS (i.e., whether the root region, secure region, realm region, or non-secure region is referenced). In the EL2 stage 1 table and the EL1&0 stage 2 table, it is the NS bit that is used by itself in the realm (and sometimes secure) state. This allows access to either target, i.e., NS=0 refers to the realm (or secure) PAS and NS=1 refers to the non-secure PAS. In addition, an AMEC flag 104 is provided. This indicates which MECID storage register should be used to provide the MECID value. In this case, the AMEC field is one bit (0 or 1), thus indicating whether the value stored in the first register 94 or the value stored in the second register 96 should be used. Of course, other numbers of registers may be provided, in which case the size of the AMEC field would be larger. Finally, each page table entry 98 includes the (physical) page number 106 to which the entry 98 pertains. Once the MECID is established, it is provided as part of an outgoing memory request, where it is ultimately consumed by the MECID consumer 64 to perform the encryption / decryption.

[0105] By storing the MECID itself in the registers 94, 96 and storing only an indicator in the page table entry 98, the page table can be kept smaller. By providing a set of registers 94, 96, multiple MECIDs can be used simultaneously, for example, if a particular realm has its own MECID but also has access to an area of ​​physical space shared with another realm, both the realm's own MECID and the MECID of the shared space can be stored simultaneously. Similarly, the hypervisor 34 can use the alternate MECID register 96 to load the MECID of the realm to access the physical address space belonging to that particular realm. By providing multiple MECID registrars 94, 96, it is also possible to keep the MECID size independent of the page table format. It is also possible to use large MECIDs (e.g., values ​​across multiple registers). In some embodiments, multiple registers can be used to store different MECIDs for different virtual-to-physical translation regimes. For example, for each of the different exception levels shown in FIG. 5, a different MECID register can be used for that realm.

[0106] It will be appreciated that the RMM 46 and / or the hypervisor 34 are responsible for loading the correct MECID value into the registers 94, 96 during a context switch operation. That is, the MECID to be used by the newly active realm is loaded into these registers 94, 96.

[0107] 7 shows an embodiment of a MECID consumer in PoE 64 working with a PAS TAG stripper in PoPA 60, where an incoming memory access request is received with a MECID and a PAS, so that the incoming memory access request will miss or be discarded from other caches 24 in the memory hierarchy.

[0108] When a memory access arrives at the MECID consumer 64, the PAS is used to look up the corresponding first key input, and the MECID is used to look up the corresponding second key input. If multiple MECID tables are provided (e.g., for each PAS), the PAS is also used to select which MECID table to use to obtain the second key input. The key inputs may be keys that perform the first and second stages of encryption. Alternatively, the key inputs may be keys that are mathematically combined together (e.g., hashing or appending bits together) to form a further key. In some embodiments, each of the two inputs is a part of a key (a default value is used or provided for worlds other than the realm world). The key input may also or alternatively be an adjustable bit. Other possibilities or combinations of these possibilities will also be understood by those skilled in the art. In any case, the key input is passed to the encryption / decryption unit 108, which uses the key input and the data itself to perform encryption (for memory write requests) or decryption (for memory read requests). In the case of a memory write request, the encrypted data is written to memory, and in the case of a memory read request, the decrypted data is returned to the requester device 4 .

[0109] Thus, the MECID itself is not a key (or keystroke), and may in fact be much larger than the MECID. This avoids the need to transmit a much larger key over the system fabric 24, 8. However, in embodiments where the PoE is much closer to the generator of the MECID (e.g., the address translation circuitry 16), it may be more practical to transmit the keystroke itself.

[0110] When a realm is created, a new entry is added to the MECID consumer 64. Similarly, when a realm is deleted or terminated, the entry in the storage of the MECID consumer 64 is deleted and the data belonging to that process can no longer be read. In contrast, the key inputs provided for each PAS are static and are determined when the device boots up.

[0111] In embodiments where PoE and POPA are separate, a first stage of encryption (eg, for a particular realm) and a second stage of encryption (for the PAS) may be performed.

[0112] The above discussion has focused on the use of MECID and encryption / decryption with respect to realms, however the same techniques may also be applied within other worlds or regions, such as secure worlds / regions.

[0113] FIG. 8 shows a flow chart 110 according to some of the above embodiments in which a key is derived using two inputs (as opposed to an input that is used directly in the encryption / decryption stage). In step 112, a memory access request is received. In step 114, a first key input is obtained for a particular domain (e.g., based on a PAS). This key input is fixed at boot time. In step 116, a second key input is obtained based on the current execution environment. This key input is dynamic in the sense that it exists as long as there is an associated execution environment for it. In step 118, a key is obtained using the key input. Then, in step 120, it is determined whether the memory access request is a write request. If not, in step 122, the data obtained from memory is decrypted using the key. If not, in step 124, the data is encrypted using the key and stored in memory. As explained above, the explicit step 118 of obtaining a key may be omitted and the key input is used directly in the decryption step 122 and / or the encryption step 124.

[0114] FIG. 9 illustrates a simulator implementation that may be used. While the above embodiments implement the invention in terms of apparatus and methods for operating specific processing hardware that supports the techniques, it is also possible to provide an instruction execution environment according to the embodiments described herein implemented by the use of a computer program. Such computer programs are often referred to as simulators insofar as the computer program provides a software-based implementation of a hardware architecture. Various simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on a host processor 430, which optionally runs a host operating system 420 and supports the simulator program 410. In some arrangements, there may be multiple layers of simulation between the hardware and the instruction execution environment provided, and / or there may be multiple different instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations that run at reasonable speeds, but certain approaches may be justified in certain situations, such as when it is desired to run code native to another processor for compatibility or reuse reasons. For example, a simulator implementation may provide an instruction execution environment that has additional functionality not supported by the host processor hardware, or that is typically associated with a different hardware architecture. An overview of simulation is given in "Some Efficient Architecture Simulation Techniques", Robert Bedichek, Winter 1990 USENIX Conference, pp. 53-63.

[0115] Although embodiments have been described above with reference to particular hardware constructs or features, in the simulated embodiments, equivalent functionality may be provided by suitable software constructs or features. For example, particular circuits may be implemented as computer program logic in the simulated embodiments. Similarly, memory hardware such as registers or caches may be implemented as software data structures in the simulated embodiments. In arrangements in which one or more of the hardware elements referenced in the foregoing embodiments reside in host hardware (e.g., host processor 430), some simulated embodiments may use the host hardware where suitable.

[0116] The simulator program 410 may be stored in a computer-readable storage medium (which may be a non-transitory medium) and provides a program interface (an instruction execution environment) to the target code 400 (which may include applications, an operating system, and a hypervisor) that is the same as the interface of the hardware architecture being modeled by the simulator program 410. Thus, the program instructions of the target code 400 may be executed from within the instruction execution environment using the simulator program 410, so that a host computer 430 that does not actually have the hardware features of the apparatus 2 described above can emulate these features. This may be useful, for example, to enable testing of target code 400 being developed for a new version of a processor architecture before a hardware device that actually supports that architecture is available, since the target code may be tested by being executed within a simulator that runs on a host device that does not support the new version of the processor architecture.

[0117] The simulator code includes processor program logic 412 that emulates the behavior of the processing circuit 10, including, for example, instruction decode program logic that decodes instructions of the target code 400, maps the instructions to a sequence of corresponding instructions in a native instruction set supported by the host hardware 430, and performs functions equivalent to the decoded instructions. The processor program logic 412 also simulates the processing of code at different exception levels and domains, as described above. The register emulate program logic 413 maintains data structures in the host address space of the host processor that emulate architectural register state defined according to the target instruction set architecture associated with the target code 400. Thus, rather than such architectural state being stored in hardware registers 12 as in the embodiment of FIG. 1, it is instead stored in the memory of the host processor 430, and the register emulate program logic 413 maps register references of instructions of the target code 400 to corresponding addresses for retrieving the simulated architectural state data from the host memory. This architectural state may include the aforementioned MECID register 94 and ALT MECID register 96 as well as the aforementioned current region indication 14 and current exception level indication 15 .

[0118] The simulation code includes memory protection program logic 416 and address translation program logic 414 that emulate the functions of MECID consumer 64 and address translation circuit 16, respectively. Thus, address translation program logic 414 translates virtual addresses specified by target code 400 into simulated physical addresses (which point to physical locations in memory from the perspective of the target code) in one of the PASs, but these simulated physical addresses are actually mapped onto the (virtual) address space of the host processor by address space mapping program logic 415. Memory protection program logic 416 "consumes" the MECID provided as part of the memory access request and provides one or more key inputs that are used to encrypt / decrypt data from memory.

[0119] FIG. 10 illustrates the location of the encryption points and the extent to which the delete and invalidate operations extend within the system. As previously mentioned, address translation circuitry 16 in the form of one or more stages 50, 52 of the MMU is used to translate virtual addresses (VA) to physical addresses (PA) and, if an access is made by an execution environment, to a MECID. The MECID is one embodiment of an encryption environment identifier used to encrypt data passing through the PoE for a particular execution environment. The granular memory protection unit 20 is used to provide a PAS TAG associated with the physical address space to be accessed (although this may also be provided directly from the address translation circuitry 16). The combination of the PAS TAG PA, and (if appropriate) the MECID, is used to access data held within a particular physical address space (identified by the PAS TAG) at a particular physical address (identified by the PA) associated (if applicable) with a particular execution environment (using the MECID). Once the PoE is passed, any MECID is consumed and used to perform encryption / decryption for storage circuits beyond the PoE.

[0120] As shown in FIG. 10, the PoE may be anywhere in the cache hierarchy 24. The closer to the processor, the more caches store encrypted data. The closer the PoE is to the memory, the fewer caches store encrypted data and instead store unencrypted data with the MECID. For example, in the embodiment of FIG. 10, the cache hierarchy 24 is comprised of a level 1 cache 130, a level 2 cache 132, and a level 3 cache 134. If the PoE is between the level 1 cache 130 and the level 2 cache 132, the data is stored unencrypted in the level 1 cache 130 and encrypted in the level 2 cache 132, the level 3 cache 134, and the main memory. In contrast, if the PoE is between the level 2 cache 132 and the level 3 cache 134, the data is stored encrypted in the level 3 cache 134 and the main memory, but unencrypted in the level 1 cache 130 and the level 2 cache 132.

[0121] If the PoE and PoPA are different, encryption for the MECID is done by the PoE, while further encryption may be done in the PoPA (for a different PAS). Note that in some embodiments, encryption of data in the PAS that is already encrypted in the PoE does not occur.

[0122] During operation of the cache, maintenance operations may need to be performed. This includes delete and invalidate operations (which are a specific type of invalidation operation) that may be performed as a result of changes in memory allocation (such as removal from an execution environment or allocation to a new execution environment). At least some of these cache maintenance operations are performed only up to the PoE and not beyond the PoE. For example, when an execution environment expires, data belonging to that execution environment must continue to be protected. Past the PoE, data is encrypted and if the key for encryption is deleted, the data can no longer be accessed. However, before the PoE in the cache hierarchy 24, data is stored unencrypted and therefore should be removed from the cache to prevent a different execution environment with the same MECID from accessing the data (the MECID identifier space is small and therefore likely to be reused). To achieve this, cache maintenance operations are performed up to the PoE, whereby the data is invalidated (and therefore no longer accessible). In some embodiments, the actual operations are delete and invalidate operations, even though deleting data (writing it back to memory) does not affect the expired execution environment.

[0123] 11 illustrates the relationship between the cache hierarchy 24, the PoE 64, and the PoPA 60. As previously described, the PoE 64 may be anywhere in the cache hierarchy 24. For example, the PoE 64 may occur after all caches in the cache hierarchy 24 and before the main memory, or may exist before the cache hierarchy 24. The PoPA 60 may be above or after the PoE 64. For example, the PoE 64 and the PoPA 60 may be located at the same point somewhere in the cache hierarchy 24. In some embodiments, the PoE 64 and the PoPA 60 may be located at alternating ends of the cache hierarchy, i.e., the PoE 64 may occur before the cache hierarchy 24 and the PoPA 60 may occur at the end of the cache hierarchy 24. Thus, there are potentially three different "zones" within the memory hierarchy: a first set of storage circuits where data is not encrypted and is aliased (up to PoE64), a second set of storage circuits where data is encrypted and is aliased (between PoE64 and PoPA60), and a third set of storage circuits where data is encrypted and is not aliased (past PoPA60).

[0124] As explained above, cache maintenance operations related to memory allocation changes are issued to caches in the cache hierarchy before but not past PoE 64. Other cache maintenance operations (such as moving data or memory pages from one region to another) can pass through PoE 64 to PoPA 60, and still other cache maintenance operations can permeate the entire memory hierarchy.

[0125] FIG. 12 shows a flow chart 140 illustrating the behavior of cache maintenance in more detail from the perspective of a particular cache. In step 142, a cache maintenance operation (CMO) is received by the cache. The cache maintenance operation includes a target and an indication of the location of the PoE64. The target can be, for example, a physical address associated with a particular region of memory to be transferred (e.g., belonging to an execution environment that is known to be out of date), or it can target the MECID itself depending on the architecture of the cache. In step 144, the cache determines whether it is before the PoE64 in the hierarchy. If not, nothing further is done and the process ends (or returns to the start). If not, in step 146, the target is deleted and invalidated. Then, in step 148, a new CMO is issued to the cache at the next cache level. The new CMO includes the same target and the same indication of the PoE64. In this way, only the cache lines targeted by the CMO are invalidated. However, this only occurs up to the PoE64. Past this point, this kind of CMO is ignored and not forwarded. Since the data belonging to the target cache lines is encrypted past PoE64, invalidating these cache lines is not strictly necessary.

[0126] Other cache maintenance operations may also be applied and are processed in the normal manner, i.e., they may still be propagated through PoE64 (if appropriate).

[0127] As a result, different cache maintenance instructions can be issued for each movement of memory between execution environments, and for moving memory between regions. Further instructions can be provided for other cache maintenance operations.

[0128] FIG. 13A illustrates targeting of cache maintenance operations. In this example, the maintenance operations are triggered by allocation of memory to an execution environment. This may occur, for example, due to the expiration and / or creation of a particular execution environment. The realm management module (RMM) causes cache maintenance operations to be performed for addresses 0x2132 and 0xC121. However, the hypervisor 32 or the operating system may be responsible for such operations being performed as well. In either case, the cache maintenance operations are sent to the level 1 cache 130. These cause corresponding entries in the cache to be invalidated. The cache maintenance operations are then sent up the memory hierarchy to, but not beyond, PoE64. In this case, this includes the level 1 cache 130 and the level 2 cache 132. However, the level 3 cache 134 is not affected, i.e., the cache maintenance operations are not forwarded. In either case, entries tagged with addresses 0xC121 or 0x2132 (which are not encrypted because the associated cache is before PoE64) are invalidated (or deleted and invalidated). These entries are encrypted and inaccessible, so no invalidation occurs past PoE64 (as a result of these CMOs).

[0129] FIG. 13B illustrates the targeting of a cache maintenance operation. In this example, the maintenance operation is triggered by the expiration of an execution environment (0xF1). The expiration of the execution environment is managed by the realm management module (RMM) 46, but a similar cache maintenance operation could alternatively be issued by, for example, the hypervisor 34. In either case, an instruction is issued indicating that the memory associated with this execution environment should be invalidated. The MECID of the corresponding execution environment is retrieved. Again, this could be performed by the RMM or the hypervisor 34, but could also be determined by another component. An invalidation instruction is then sent, referencing the particular MECID associated with the expired execution environment (0xF14E in this case). As explained above, this invalidation instruction is sent up the memory hierarchy to, but not beyond, the PoE64. In this case, this includes the level 1 cache 130 and the level 2 cache 132. In either case, entries tagged with MECID 0xF14E (which are not encrypted because the associated cache is before PoE) are invalidated (or removed and invalidated).

[0130] Note that the lookup performed between the execution environment and the MECID allows for MECIDs that are not associated with any single execution environment, thereby allowing for sharing of data. In these situations, the entries belonging to such MECIDs may be invalidated when a particular one of the associated execution environments is terminated (e.g., when one of the execution environments acts as a "master" for the MECID) or may be invalidated when all of the associated execution environments are terminated. A further reason for separating the execution environment identifiers and the MECIDs is to limit the reuse of MECIDs and allow more MECIDs to exist simultaneously than are currently active. For example, execution environments can be made dormant (inactive) but their data can remain in the system. In this embodiment, for example, there may only be 256 execution environments that can be executed simultaneously (since the execution environment identifiers are 8 bits). However, the MECID identifiers are larger (16 bits), and therefore execution environments can be swapped in and out.

[0131] As a result of the above, the cache hierarchy is less affected by cache maintenance operations. This is because certain cache maintenance operations (e.g., operations that invalidate expired execution environments) do not need to occur past the PoE. Thus, the impact of invalidation requests on the system can be reduced. This does not compromise security, since once past the PoE, the data is encrypted and therefore incomprehensible if another execution environment attempts to access those memory items. Thus, cache maintenance operations that invalidate these data entries serve no useful purpose.

[0132] FIG. 14 illustrates a method of data processing 140 according to some embodiments. In step 142, processing is performed in one of a plurality (e.g., two or more, such as three or more) of domains. One of these domains is subdivided into several execution environments (e.g., realms). The processing accesses a memory in the memory hierarchy. In step 144, a point of encryption is defined in the memory hierarchy. This divides the memory hierarchy into encrypted components (where data is stored in encrypted form) and unencrypted components (where data is not encrypted). Then, in step 146, at least some maintenance operations to be issued (such as those associated with expired execution environments) are prevented from being issued at or beyond PoE64. These cache maintenance operations are not issued to storage circuits where data is stored in encrypted form.

[0133] FIG. 15 illustrates a simulator implementation that may be used. While the above embodiments implement the invention in terms of apparatus and methods for operating specific processing hardware that supports the technique, it is also possible to provide an instruction execution environment according to the embodiments described herein implemented by the use of a computer program. Such computer programs are often referred to as simulators insofar as the computer program provides a software-based implementation of a hardware architecture. Various simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on a host processor 430, which optionally runs a host operating system 420 and supports the simulator program 410. In some arrangements, there may be multiple layers of simulation between the hardware and the instruction execution environment provided, and / or there may be multiple different instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide a simulator implementation that runs at a reasonable speed, but such an approach may be justified in certain situations, such as when it is desired to run code native to another processor for compatibility or reuse reasons. For example, a simulator implementation may provide an instruction execution environment that has additional functionality not supported by the host processor hardware, or that is typically associated with a different hardware architecture. An overview of simulation is given in "Some Efficient Architecture Simulation Techniques", Robert Bedichek, Winter 1990 USENIX Conference, pp. 53-63.

[0134] Although embodiments have been described above with reference to particular hardware constructs or features, in the simulated embodiments, equivalent functionality may be provided by suitable software constructs or features. For example, particular circuits may be implemented as computer program logic in the simulated embodiments. Similarly, memory hardware such as registers or caches may be implemented as software data structures in the simulated embodiments. In arrangements in which one or more of the hardware elements referenced in the foregoing embodiments reside in host hardware (e.g., host processor 430), some simulated embodiments may use the host hardware where suitable.

[0135] The simulator program 410 may be stored in a computer-readable storage medium (which may be a non-transitory medium) and provides a program interface (an instruction execution environment) to the target code 400 (which may include applications, an operating system, and a hypervisor) that is the same as the interface of the hardware architecture being modeled by the simulator program 410. Thus, the program instructions of the target code 400 may be executed from within the instruction execution environment using the simulator program 410, so that a host computer 430 that does not actually have the hardware features of the apparatus 2 described above can emulate these features. This may be useful, for example, to enable testing of target code 400 being developed for a new version of a processor architecture before a hardware device that actually supports that architecture is available, since the target code may be tested by being executed within a simulator that runs on a host device that does not support the new version of the processor architecture.

[0136] The simulator code includes processor program logic 412 that emulates the behavior of the processing circuit 10, including, for example, instruction decode program logic that decodes instructions of the target code 400, maps the instructions to a sequence of corresponding instructions in a native instruction set supported by the host hardware 430, and performs functions equivalent to the decoded instructions. The processor program logic 412 also simulates the processing of code at different exception levels and domains, as described above. The register emulate program logic 413 maintains data structures in the host address space of the host processor that emulate architectural register state defined according to the target instruction set architecture associated with the target code 400. Thus, rather than such architectural state being stored in hardware registers 12 as in the embodiment of FIG. 1, it is instead stored in the memory of the host processor 430, and the register emulate program logic 413 maps register references of instructions of the target code 400 to corresponding addresses for retrieving the simulated architectural state data from the host memory. This architectural state may include the aforementioned current region indication 14 and current exception level indication 15, along with the aforementioned MECID register 94 and ALT MECID register 96. Similarly, storage circuit emulate program logic 148 maintains data structures within the host address space of the host processor that emulates the memory hierarchy. Thus, instead of data being stored in level 1 cache 130, level 2 cache 132, level 3 cache 134, and memory 150 as in the embodiment of FIG. 10 (for example), data is instead stored in the memory of host processor 430, and storage circuit emulate program logic 148 maps memory addresses of instructions of target code 400 to corresponding addresses to obtain simulated memory addresses from host memory.

[0137] The simulation code includes memory protection program logic 416 and address translation program logic 414 that emulate the functions of the MECID consumer 64 and address translation circuit 16, respectively. Thus, the address translation program logic 414 translates the virtual addresses specified by the target code 400 into simulated physical addresses (pointing to physical locations in memory from the perspective of the target code) in one of the PASs, but these simulated physical addresses are actually mapped by the address space mapping program logic 415 into the virtual memory structures 130, 132, 134, 150 emulated by the storage circuit emulation program logic 148. The memory protection program logic 416 "consumes" the MECID provided as part of the memory access request and provides one or more key inputs that are used to encrypt / decrypt data from memory as described above. The storage circuit emulation logic 148 can also emulate the functions of cache maintenance operations, both the encryption point 64 and the physical alias point 60, as described above.

[0138] FIG. 16 illustrates an example system according to some embodiments. In these embodiments, a memory access request to a virtual address is issued by an execution environment (realm) executing within a subdivided world / area (i.e., realm area) on processing circuitry 10. The memory access request is received by memory translation circuitry (e.g., address translation circuitry) 16, where the virtual address (VA) is translated to a physical address (PA). Additionally, a PAS is determined and a MECID is determined, as previously described. The memory access request is then sent to the memory hierarchy to locate the requested data. In this embodiment, it is received by storage circuitry in the form of a level 1 cache 130. Since the level 1 cache, in this embodiment, precedes PoE64, the contents of level 1 cache 130 are not encrypted. Thus, each unencrypted cache line entry stores data associated with the address of the cache line, the PAS, and the MECID.

[0139] A "hit" occurs on the cache if / when a physical memory address (PA) corresponds to one of the cache lines stored in the cache. In this situation, the requested data is returned and therefore the memory access request does not need to proceed to main memory 150. In contrast, a "miss" occurs when none of the cache lines correspond to the requested physical memory address (none of the cache lines store the data being requested). In this situation, the memory access request is forwarded further up the memory hierarchy towards main memory 150. Once the requested data is located (which may be in main memory), the data may be stored in lower level cache 130 so that it may be accessed again more quickly in the future.

[0140] As previously mentioned, each cache line in cache 130 stores data associated with the physical address of the cache line, the data of the cache line, an identification of the physical address space (PAS) with which the data is associated, and a MECID, the latter being one example of a cryptographic environment identifier that may be associated with a subset of execution environments (often one particular execution environment). This is the execution environment (or environments) that "owns" the data. In response to a hit, decision circuit 180 determines whether there is a match between the MECID of the hit entry in storage circuit 130 and the MECID provided in the memory access request.

[0141] FIG. 17 illustrates an example of a MECID mismatch. This can occur for a number of reasons. For example, a table in memory translation circuit 16 may contain multiple entries for the same PA, each belonging to a different MECID. Also, insufficient cache maintenance operations may be performed when a MECID is reallocated. The MECID may be too large for the system, resulting in repeated components of the MECID that are actually used. In some circumstances, a mismatch may occur due to insufficient translation lookaside buffer (TLB) maintenance and barriers when the MECID register is updated.

[0142] In any event, this example illustrates a memory read request issued from the memory translation circuit. The request is directed to physical address 0xB1432602. This is comprised of a cache line address 0xB14326 and an offset into the cache line of 02, which is the particular portion of the cache line that the memory read request is seeking to read. The request is also directed to a PAS of 01 (which in this example refers to the realm PAS) and a MECID of 0xF143, which is the MECID associated with the execution environment or realm from which the memory read request is issued. This is received by storage circuitry 130, which determines whether there is a hit at the memory address being accessed. In this case, there is a hit because storage circuitry 130 contains an entry with cache line address 0xB14326. The PAS(01) also matches. However, in this case, even though there is a hit, the decision circuitry can determine (by comparison) that the MECIDs are mismatched. In particular, the MECID sent with the request is 0xF143, while the MECID stored for the cache line is 0xF273. Thus, the request is being issued by an execution environment that should not have access to the line. Thus, an error action can be taken.

[0143] Recall that a MECID may be associated with several execution environments (in situations where data is shared between execution environments), so a MECID does not necessarily identify a particular execution environment.

[0144] There are several configurations that can be employed to prevent mismatches from occurring, as well as several error actions that can be taken.

[0145] FIG. 18 illustrates a poison mode of operation that poisons an associated cache line in response to a mismatch. Here, a memory write request is issued that targets a particular portion of the cache line. However, a mismatch occurs in the MECID. In this example, the targeted portions of the cache line are overwritten / modified by the write request. These portions are expected to be correct and therefore are not poisoned. However, other portions of the cache line are recorded as being poisoned. If these poisoned portions of the cache line are read by the processing circuitry in the future (e.g., as a result of a later memory read request to these portions of the cache line), a poison notation is returned to the processing circuitry. This causes a synchronization error by the processing circuitry.

[0146] By not raising an error immediately, such an error can be avoided entirely: if, for example, future memory read requests are never issued to other parts of the cache line, the poison representation is never sent back to the processor and the error never occurs; thus, the mismatch would never have occurred, but it would have no effect.

[0147] In other embodiments, an entire cache line may be poisoned as a result of a write to any portion of the cache line, since the overwritten data may be said to have caused the original data to be corrupted. In other embodiments, a memory read request may result in some or all of the cache line being poisoned and immediately returned to the processing circuitry, which may (almost immediately) cause a synchronization error. In some embodiments, all or part of the data returned from the cache as part of the read request is poisoned, but the cache line itself remains unmodified.

[0148] Also, in this embodiment, in addition to poisoning the mismatched cache line, the MECID of the cache line is updated to the MECID provided in the memory access request.

[0149] FIG. 19 shows an example implementation in which an alias mode of operation is shown. Here, it is determined whether a memory read request is a hit or a miss based on the PA, PAS, and MECID. That is, all three components are used to form an "effective address." In this example, a first read request is directed to address 0xB1432620 and uses a MECID of 0x2170. On the surface, the address should hit entry 182 in cache 130 since the PA matches. However, since the MECID, PAS, and PA are treated as a collectively valid "address," and all three do not match (the MECID of entry 182 is 0xF273 compared to the MECID of the request, which is 0x2170), there is a miss. This can be determined by decision circuit 180, which looks for a match for each of the PA, PAS, and MECID.

[0150] In contrast, a second memory read request made to the exact same PA and PAS with a different MECID of 0xF273 will result in a hit since both the PA, PAS, and MECID match.)

[0151] This therefore prevents a mismatch from occurring since a memory access request to a physical address can only hit a cache line if the PA, PAS, and MECID are all the same in the cache line and the memory access request.

[0152] There are several actions that can be taken in response to a miss. In some embodiments, a mismatch in MECID alone can be used to prevent the request from proceeding further. In other embodiments, the miss is forwarded up the memory hierarchy. When the request reaches the PoE, it is highly likely that an incorrect MECID will be used to select a keystroke, thus resulting in an incorrect decryption of the requested data (in the case of a read request) or an incorrect encoding of the provided data (in the case of a write request). However, in both cases, the goal of maintaining confidentiality of the data is maintained.

[0153] 20 illustrates an embodiment of a delete mode of operation, where when a mismatch is detected, the mismatched cache line in cache 130 is deleted (written back further up the memory hierarchy, such as past the point of encryption, such as to memory), the mismatched line is then invalidated, and the requested line is fetched from memory.

[0154] Thus, in this example, a memory access to read at address 0xB1432620 with MECID 0xF273 is a mismatch on cache line address 0xb14326 with MECID 0x2170. Thus, the cache line is written back (deleted) to memory and invalidated (the "V" flag is changed from 0 to 1). The subject of the request (address 0xB1432620) is then fetched from memory with MECID 0x2170. In fact, this memory access request may still fail if the MECID is incorrect in the memory hierarchy. In particular, past the point of encryption, if the MECID is incorrect, an incorrect keystroke will be selected for decryption and garbage will be returned by the memory access request. In any case, the fetched data is then stored in cache 130 along with the MECID of the new access request.

[0155] 21 illustrates an example of an erase mode of operation. In this mode of operation, when a mismatch is detected, the data of the mismatched line in cache 130 is zeroed, scrambled, or randomized so that it is no longer understandable. This makes the line unusable. Note that this is distinct from the operation of invalidating a cache line (e.g., by setting the validity flag "V" to 0).

[0156] Thus, in this example, a mismatch is caused by a memory request to address 0x94130001, which hits the cache line at address 0x941300. However, a mismatch occurs because the request has a MECID of 0x2142, while the cache line has a MECID of 0x7D04. Thus, the cache line with address 0x941300 is zeroed (in this case) by setting all bits of the data to 0. The cache line can then be returned. As a result, the unencrypted data is inaccessible. Note that in this example, the cache line is not invalidated (although such an operating mode may further set the cache line as invalid).

[0157] FIG. 22 shows an example of the overall process in the form of a flow chart 190. In step 192, a memory access request is received by the storage circuit 130. In step 194, it is determined whether there was a hit. If not, in step 196, the request is forwarded further up the memory hierarchy, e.g., towards the main memory 150. The process then returns to step 192. If not, in step 198, it is determined (e.g., by the decision circuit 180) whether a mismatch has occurred between the MECID of the memory access request and a hit entry in the storage circuit 130. If not, in step 200, the memory access request is completed (e.g., by reading or writing the associated entry in the storage circuit 130) and the request then proceeds to step 192. If not, in step 202, it is determined in which mode the system is operating. If the system is in a poison mode of operation, in step 202, the entry in the storage circuit 130 is poisoned as described above. The process then proceeds to step 210. If the system is in a delete mode of operation, then in step 204, the entry in the storage circuit 130 is deleted and disabled, and the process proceeds to step 210. If the system is in an erase mode of operation, then in step 206, the entry in the storage circuit 130 is zeroed or scrambled. The process then proceeds to step 210. These are all examples of error modes of operation, in that a mismatch causes an error. In contrast, the alias mode of operation (shown in FIG. 19) is not an error mode, since it actively prevents a mismatch from occurring in the first place. Collectively, the error mode and the alias mode form an enabled mode of operation. The other modes of operation of the decision circuit 180 are disabled modes of operation, and in step 208, the mismatch is simply ignored and the request is completed. The process then proceeds to step 210.

[0158] After some number of operational modes have been enabled, the process proceeds to step 210, where it is determined whether a synchronous mode is also enabled. If so, then in step 212, an asynchronous exception is also generated (e.g., by writing to a register 12 associated with processing circuit 10 regarding the mismatch). In either case, the process then returns to step 192.

[0159] In each of the error operating modes, the MECID of the non-matching entry may also be updated to the MECID of the incoming memory access request.

[0160] The device may be capable of switching at run-time between each or a subset of the enabled operating modes and a disabled mode. Each of the enabled operating modes is of course dependent, and the system may include any combination of these. A disabled mode may or may not also be present.

[0161] FIG. 23 illustrates the interaction between enabled modes and speculative execution in the form of a flow chart 214. In speculative execution, instructions are executed before it is known whether they should be executed or not (e.g., pending the outcome of a branch instruction). Speculative reads can occur using an incorrect MECID (as explained above), and therefore, in order to perform speculative execution, it is desirable for one of the enabled modes to be active. In step 216, it is determined whether the decision circuit 180 is in an enabled mode of operation. If so, in step 218, the speculative mode of operation is enabled, allowing speculative reads and writes to be performed. If not, the speculative mode of operation is disabled. This prevents speculative read operations from being performed (in some embodiments, speculative write operations may also be prevented). In either case, the process then returns to step 216.

[0162] As an alternative to this process, rather than continually "polling" the current operating mode of the decision circuit 180, the speculative operating mode can be enabled / disabled whenever the operating mode of the decision circuit 180 is changed.

[0163] FIG. 24 illustrates a simulator implementation that may be used. While the above embodiments implement the invention in terms of apparatus and methods for operating specific processing hardware that supports the techniques, it is also possible to provide an instruction execution environment according to the embodiments described herein implemented by the use of a computer program. Such computer programs are often referred to as simulators insofar as they provide a software-based implementation of a hardware architecture. Various simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on a host processor 430, which optionally runs a host operating system 420 and supports the simulator program 410. In some arrangements, there may be multiple layers of simulation between the hardware and the instruction execution environment provided, and / or there may be multiple different instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide a simulator implementation that runs at a reasonable speed, but such an approach may be justified in certain situations, such as when it is desired to run code native to another processor for compatibility or reuse reasons. For example, a simulator implementation may provide an instruction execution environment that has additional functionality not supported by the host processor hardware, or that is typically associated with a different hardware architecture. An overview of simulation is given in "Some Efficient Architecture Simulation Techniques", Robert Bedichek, Winter 1990 USENIX Conference, pp. 53-63.

[0164] Although embodiments have been described above with reference to particular hardware constructs or features, in the simulated embodiments, equivalent functionality may be provided by suitable software constructs or features. For example, particular circuits may be implemented as computer program logic in the simulated embodiments. Similarly, memory hardware such as registers or caches may be implemented as software data structures in the simulated embodiments. In arrangements in which one or more of the hardware elements referenced in the foregoing embodiments reside in host hardware (e.g., host processor 430), some simulated embodiments may use the host hardware where suitable.

[0165] The simulator program 410 may be stored in a computer-readable storage medium (which may be a non-transitory medium) and provides a program interface (an instruction execution environment) to the target code 400 (which may include applications, an operating system, and a hypervisor) that is the same as the interface of the hardware architecture being modeled by the simulator program 410. Thus, the program instructions of the target code 400 may be executed from within the instruction execution environment using the simulator program 410, so that a host computer 430 that does not actually have the hardware features of the apparatus 2 described above can emulate these features. This may be useful, for example, to enable testing of target code 400 being developed for a new version of a processor architecture before a hardware device that actually supports that architecture is available, since the target code may be tested by being executed within a simulator that runs on a host device that does not support the new version of the processor architecture.

[0166] The simulator code includes processor program logic 412 that emulates the behavior of the processing circuit 10, including, for example, instruction decode program logic that decodes instructions of the target code 400, maps the instructions to a sequence of corresponding instructions in a native instruction set supported by the host hardware 430, and performs functions equivalent to the decoded instructions. The processor program logic 412 also simulates the processing of code at different exception levels and domains, as described above. The register emulate program logic 413 maintains data structures in the host address space of the host processor that emulate architectural register state defined according to the target instruction set architecture associated with the target code 400. Thus, rather than such architectural state being stored in hardware registers 12 as in the embodiment of FIG. 1, it is instead stored in the memory of the host processor 430, and the register emulate program logic 413 maps register references of instructions of the target code 400 to corresponding addresses for retrieving the simulated architectural state data from the host memory. This architectural state may include the aforementioned current region indication 14 and current exception level indication 15, along with the aforementioned MECID register 94. Similarly, the storage circuit emulate program logic 148 maintains data structures within the host address space of the host processor that emulates the memory hierarchy. Thus, instead of data being stored in level 1 cache 130, level 2 cache 132, level 3 cache 134, and memory 150 as in the embodiment of FIG. 10 (for example), data is instead stored in the memory of the host processor 430, and the storage circuit emulate program logic 148 maps the memory addresses of instructions of the target code 400 to the corresponding addresses to obtain the simulated memory addresses from the host memory.

[0167] The simulation code includes address translation program logic 414 that emulates the function of the address translation circuit or memory translation circuit 16, respectively. Thus, the address translation program logic 414 translates the virtual addresses specified by the target code 400 into simulated physical addresses (pointing to physical locations in memory from the perspective of the target code) in one of the PASs, but these simulated physical addresses are actually mapped by the address space mapping program logic 415 onto the virtual memory structures 130, 132, 134, 150 emulated by the storage circuit emulation program logic 148. The determination program logic 151 can determine whether the MECID provided as part of a simulated memory access request to a memory address matches the MECID associated with an entry in the simulated memory hierarchy 148 for that memory address, thereby performing the function of the determination circuit 180 described above. The storage circuit emulation logic 148 can emulate the encryption point 64 and the physical alias point 60, as described above. The decision program logic 151 may determine whether a difference is detected between the encryption environment identifiers, as previously described.

[0168] In this application, the term "configured to..." is used to mean that an element of an apparatus has a configuration that is capable of performing a defined operation. In this context, "configuration" refers to a manner of arrangement or interconnection of hardware or software. For example, an apparatus may have dedicated hardware that provides the defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not imply that an apparatus element needs to be modified in any way to provide the defined operation.

[0169] Although illustrative embodiments of the invention have been described in detail herein with reference to the accompanying drawings, it will be understood that the invention is not limited to those precise embodiments, and that various changes, additions, and modifications may be made by those skilled in the art without departing from the scope and spirit of the invention as defined by the appended claims. For example, various combinations of the features of the dependent claims may be made with the features of the independent claims without departing from the scope of the invention.

Claims

1. It is a device, A processing circuit configured to execute processing in one of at least two fixed regions, wherein one of the regions is subdivided into a variable number of execution environments, and one of the execution environments is a management execution environment configured to manage the execution environments. A memory protection circuit that defines an encryption point after at least one unencrypted storage circuit in the memory hierarchy and before at least one encrypted storage circuit in the memory hierarchy, The at least one encrypted memory circuit is configured to use a key input to encrypt or decrypt data of a memory access request issued from within the current region of the region, The aforementioned key input differs for each of the aforementioned regions and each of the aforementioned execution environments. A device wherein the management execution environment is configured to prevent issuing maintenance operations to the at least one encrypted storage circuit in the memory hierarchy.

2. The apparatus according to claim 1, wherein the management execution environment is configured to issue the maintenance operation to at least one unencrypted storage circuit in the memory hierarchy in response to a change in memory allocation made to one of the execution environments.

3. The apparatus according to claim 2, wherein the maintenance operation is a deactivation operation.

4. The apparatus according to claim 2, wherein the maintenance operation is a deletion and a deactivation operation.

5. The apparatus according to claim 3, wherein the maintenance operation is configured to invalidate an entry in the at least one unencrypted memory circuit associated with one of the execution environments.

6. The apparatus according to claim 2, wherein the change in the allocation is the allocation of memory to one of the execution environments.

7. The apparatus according to claim 6, wherein the maintenance operation is configured to invalidate an entry in the at least one encrypted memory circuit associated with an expired execution environment among the execution environments.

8. Each of the execution environments is associated with a cryptographic environment identifier used to generate the key input, The apparatus according to any one of claims 1 to 7, wherein the maintenance operation is configured to invalidate an entry in the memory hierarchy associated with the encryption environment identifier.

9. The memory address to which the memory access request is issued is a physical memory address in one of multiple physical address spaces. The apparatus according to any one of claims 1 to 7, wherein each of the physical address spaces is associated with one of the at least two regions.

10. The memory protection circuit defines a physical alias point located after at least one aliased storage circuit in the memory hierarchy and before at least one unaliased storage circuit in the memory hierarchy. The apparatus according to claim 9, wherein the at least one aliased memory circuit treats physical addresses from different physical address spaces corresponding to the same memory system resource as if the physical addresses corresponded to different memory system resources.

11. The apparatus according to claim 10, wherein the point of the physical alias is the encryption point or thereafter.

12. The apparatus according to claim 9, wherein the point of the physical alias is the point of encryption.

13. The apparatus according to claim 9, wherein, in response to a memory transition request that requests the transfer of memory from a source physical address space to a destination physical address space, the maintenance operation is configured to invalidate at least some entries in the at least one aliased memory circuit.

14. The apparatus according to claim 13, wherein at least some of the entries are assigned to one of the at least two regions associated with the origin physical address space.

15. It is a method, Processing is performed in one of at least two fixed-number regions, where one of the regions is subdivided into a variable number of execution environments, and one of the execution environments is a management execution environment configured to manage the execution environments. Defining an encryption point after at least one unencrypted memory circuit in the memory hierarchy and before at least one encrypted memory circuit in the memory hierarchy, To prevent issuing a maintenance operation to at least one encrypted storage circuit in the memory hierarchy, This includes using key input to encrypt or decrypt data of a memory access request issued for a memory address from within the current region of the said region, The aforementioned key input differs for each of the aforementioned regions and each of the aforementioned execution environments. A method wherein the management execution environment is configured to prevent issuing maintenance operations to at least one encrypted storage data structure in the memory hierarchy.

16. A computer program for controlling a host data processing device to provide an instruction environment for executing target code, wherein the computer program is A processing program logic configured to simulate the processing of the target code in at least one of two regions, one of which is subdivided into a variable number of execution environments, and one of which is a management execution environment configured to manage the execution environments; A memory protection program logic configured to define an encryption point after at least one unencrypted stored data structure in the memory hierarchy and before at least one encrypted stored data structure in the memory hierarchy, The at least one encrypted storage data structure is configured to use key input to perform encryption or decryption on data of a memory access request issued from within the current region of the region, The aforementioned key input differs for each of the aforementioned regions and each of the aforementioned execution environments. A computer program configured to prevent the management execution environment from issuing maintenance operations to the at least one encrypted storage data structure in the memory hierarchy.