Apparatus and method using a plurality of physical address spaces

By employing an address translation circuit and a requester-side filtering circuit for granularity protection lookups in data processing systems, the solution effectively manages memory access across multiple physical address spaces, enhancing security and preventing information leaks.

JP7691995B2Active Publication Date: 2025-06-12ARM LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022557074
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-24
Filing Date
2021-01-26
Publication Date
2025-06-12
Estimated Expiration
2041-01-26

AI Technical Summary

Technical Problem

Existing data processing systems lack effective mechanisms for securely managing memory access across multiple physical address spaces, leading to potential information leaks and security breaches.

Method used

The implementation of an address translation circuit and a requester-side filtering circuit that perform granularity protection lookups to determine whether a memory access request can be processed, based on granularity protection information indicating permitted physical address spaces.

Benefits of technology

This solution enhances security by isolating memory access to different physical address spaces, reducing the reliance on page table permission information and preventing unauthorized access, thereby improving overall system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691995000002
    Figure 0007691995000002
  • Figure 0007691995000003
    Figure 0007691995000003
  • Figure 0007691995000004
    Figure 0007691995000004
Patent Text Reader

Abstract

The address translation circuit (16) translates a virtual address specified by a memory access request issued by a requester circuit into a target physical address (PA). The requester-side filtering circuit (20) performs a granule protection lookup based on the target PA and a selected physical address space (PAS) to determine whether the memory access request is allowed to be passed to a cache or an interconnect. In the granule protection lookup, the requester-side filtering circuit obtains granule protection information that corresponds to a target granule of a physical address that includes the target PA and indicates at least one permitted PAS associated with the target granule, and blocks the memory access request when the granule protection information indicates that the selected PAS is not a permitted PAS.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technique relates to the field of data processing.

[0002] A data processing system may have an address translation circuit for translating a virtual address of a memory access request into a physical address corresponding to the accessed location within the memory system.

[0003] In at least some examples, an apparatus can include an address translation circuit for translating a target virtual address specified by a memory access request issued by a requester circuit into a target physical address, and a requester-side circuit that performs a granularity protection lookup based on the target physical address and a selected physical address space associated with the memory access request to determine whether the memory access request is to be passed to a cache or passed to an interconnect to communicate with a completer device to process the memory access request, wherein the selected physical address space is one of a plurality of physical address spaces, and in the granularity protection lookup, the requester-side filtering circuit obtains granularity protection information corresponding to the target granularity of the physical address including the target physical address, the granularity protection information indicating at least one permitted physical address space associated with the target granularity, and is configured to block the memory access request when the granularity protection information indicates that the selected physical address space is not one of the at least one permitted physical address spaces.

[0004] In at least some examples, converting a target virtual address specified by a memory access request issued by a requester circuit to a target physical address, and in a requester-side filtering circuit, performing a granularity protection lookup based on the target physical address and a selected physical address space associated with the memory access request to determine whether the memory access request is to be passed to a cache to process the memory access request, or passed to an interconnect to communicate with a completer device, wherein the selected physical address space is one of a plurality of physical address spaces, and in the granularity protection lookup, the requester-side filtering circuit obtains granularity protection information corresponding to a target granularity of a physical address including the target physical address, the granularity protection information indicating at least one permitted physical address space associated with the target granularity, and blocks the memory access request when the granularity protection information indicates that the selected physical address space is not one of the at least one permitted physical address spaces, a data processing method is provided.

[0005] At least some embodiments provide a computer program for controlling a host data processing apparatus that provides an instruction execution environment for executing target code. The computer program includes address translation program logic for translating a simulated virtual address of a target specified by a memory access request into a simulated physical address of the target, the simulated physical address of the target, and filtering program logic for performing a granularity protection lookup based on the selected simulated physical address space associated with the memory access request to determine that the memory access request can be processed. The selected simulated physical address space is one of a plurality of simulated physical address spaces. In the granularity protection lookup, the filtering program logic obtains granularity protection information corresponding to a target granularity of a simulated physical address including the simulated physical address of the target, the granularity protection information indicating at least one permitted simulated physical address space associated with the target granularity, and is configured to prevent the memory access request from being processed when the granularity protection information indicates that the selected simulated physical address space is not one of the at least one permitted simulated physical address spaces. A computer program is provided.

[0006] At least some examples provide a computer-readable storage medium storing the computer program described above. The computer-readable storage medium can be a non-transitory storage medium or a transitory storage medium.

Brief Description of the Drawings

[0007] Further aspects, features, and advantages of the present technology will become apparent from the following description of examples read in conjunction with the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

[0008] Control of access to the physical address space A data processing system can support the use of virtual memory and is provided with an address translation circuit that translates a virtual address specified by a memory access request into a physical address associated with a location within the memory system being accessed. The mapping between the virtual address and the physical address can be defined by one or more page table structures. Page table entries within the page table structure can also define some access permission information that can control whether a given software process executing on the processing circuit can access a particular virtual address.

[0009] In some processing systems, all virtual addresses can be mapped by an address translation circuit onto a single physical address space used by the memory system to identify locations within the memory being accessed. In such a system, control over whether a particular software process can access a particular address is provided only based on the page table structure used to provide the virtual-to-physical address translation mapping. However, such a page table structure can typically be defined by an operating system and / or a hypervisor. If the operating system or hypervisor is accessed illegally, it may cause an information leak where an attacker may be able to access confidential information.

[0010] Thus, in some systems where a particular process needs to be securely executed in isolation from other processes, the system can support operations within several regions and can support several distinct physical address spaces, and for at least some components of the memory system, memory access requests in which virtual addresses are translated to physical addresses within different physical address spaces are treated as if accessing completely different addresses in memory, even if the physical addresses within each physical address space actually correspond to the same location in memory. As seen in some memory system components, by isolating accesses to different physical address spaces from different operating regions of the processing circuit, stronger security guarantees can be provided that do not depend on page table permission information set by the operating system or hypervisor.

[0011] The processing circuit can support processing within a root region involved in managing switches between other regions in which the processing circuit can operate. By providing a dedicated root region for controlling the switch, it can help maintain security by limiting the extent to which code running in one region can trigger a switch to another region. For example, the root region can perform various security checks when a region switch is requested.

[0012] Thus, the processing circuit can support processing being executed in at least three regions, namely the root region and at least one of two other regions. The address translation circuit can translate the virtual address of a memory access executed from the current region to a physical address within one of a plurality of physical address spaces selected based at least on the current region.

[0013] In the embodiments described below, the plurality of physical address spaces include a root physical address space associated with a root region, separate from the physical address spaces associated with other regions. Thus, the root region has its own physical address space allocated to it, rather than using one of the physical address spaces associated with one of the other regions. By providing a dedicated root physical address space isolated from the physical address spaces associated with other regions, stronger security guarantees can be provided for the data or code associated with the root region, which can be considered most important for security in view of managing entry into other regions. Also, by providing a dedicated root physical address space distinct from the physical address spaces of other regions, the allocation of physical addresses within each physical address space can be simplified to specific units of the hardware memory storage, thus simplifying system development. For example, by identifying a separate root physical address space, it can be simpler for the data or program code associated with the root region to be preferentially stored in protected on-chip memory rather than less secure off-chip memory, and the overhead for determining the portion associated with the root region is smaller than when the code or data of the root region is stored in a common address space shared with other regions.

[0014] The root physical address space may be exclusively accessible from the root region. Thus, when the processing circuit is operating in one of the other regions, the processing circuit may not be able to access the root physical address space. This ensures that data or program code on which the root region depends cannot be tampered with by code being executed in one of the other regions, for managing switching between regions or controlling the rights that the processing circuit has in one of the other regions, improving security.

[0015] On the one hand, the root region may be able to access all of a plurality of physical address spaces. The code running within the root region must be trusted by any party providing code that operates in one of the other regions, and since the root region code is involved in switching to that particular region in which the party's code is running, essentially the root region can be trusted to access any physical address space. By making all physical address spaces accessible from the root region, functions such as migrating memory regions in and out of regions, for example copying code and data into a region during startup, and providing services to that region can be performed.

[0016] The address translation circuit may limit which physical address space is accessible depending on the current region. If a particular physical address space is accessible from the current region, this means that the address translation circuit can translate the virtual address specified for a memory access issued from the current region into a physical address within that particular physical address space. This does not necessarily mean that the memory access is permitted. Because even if, in a particular memory access, the virtual address is translated into a physical address of a particular physical address space, there may be further checks performed to determine whether that physical address is actually permitted to be accessed within that particular physical address space. This will be discussed further below with reference to the granularity protection information defining the partitioning of physical addresses among the respective physical address spaces. Still, stronger security guarantees can be provided by restricting which subset of the physical address space is accessible from the current region.

[0017] In some examples, the processing circuit can support two additional regions in addition to the root region. For example, the other regions can include a secure region associated with a secure physical address space and a less secure region associated with a less secure physical address space. The less secure physical address space can be accessible from each of the less secure region, the secure region, and the root region. The secure physical address space can be accessible from the secure region and the root region, but may not be accessible from the less secure region. The root region can be accessible to the root region, but the less secure region and the secure region may not be accessible to the root region. Thus, this provides stronger security guarantees than when the page table is used as the sole security control mechanism, such that code running within the secure region can protect that code or data from access by code operating within the less secure region. For example, portions of the code that require stronger security can be run within a secure region managed by a trusted operating system separate from the non-secure operating system operating within the less secure region. An example of a system that supports such a secure region and a less secure region can be a processing system operating according to a processing architecture that supports the TrustZone trademark architecture feature provided by Arm Limited (Cambridge, UK). In a conventional TrustZone implementation, the monitor code for managing the switch between the secure region and the less secure region uses the same secure physical address space used by the secure region. In contrast, as described above, providing a root region for managing the switch between other regions and allocating a dedicated root physical address space for use by the root region can help improve security and simplify system development.

[0018] However, in other examples, other regions may include at least three other regions in addition to a further region, such as a root region. These regions can include the secure regions and less secure regions described above, but may also include at least one further region associated with a further physical address space. The less secure physical address space may also be accessible from the further region, while the further physical address space may be accessible from the further region and the root region, but may not be accessible from the less secure region. Thus, similar to the secure regions, the further regions can be considered more secure than the less secure regions, enabling further partitioning of code into respective worlds associated with separate physical address spaces and restricting their interactions.

[0019] In some examples, the respective regions are hierarchically associated with privilege levels that increase as the system ascends from a less secure region, through the secure regions and the further regions, to the root region, and the further regions are considered to have higher privilege than the secure regions and thus may have access to the secure physical address space.

[0020] However, there is an increasing desire for a secure computing environment that limits the need for a software provider to trust other software providers related to other software running on the same hardware platform. For example, in mobile payment and banking, the implementation of anti-cheating or anti-copyright infringement mechanisms in computer games, the security extension of operating system platforms, secure virtual machine hosting in cloud systems, confidential computing, etc., there can be applications in several fields where the provider of software code may be negative about trusting the provider of the operating system or hypervisor (components that may have been considered trustworthy in the past). In systems such as those based on the above TrustZone (registered trademark) architecture that support secure and less secure regions with each having its own physical address space, as secure components operating in the secure region become more prevalent, the set of software that normally operates in the secure region has expanded to include several software provided by different software providers, including, for example, the following parties. An original equipment manufacturer (OEM) that assembles a processing device (such as a mobile phone) from components including a silicon integrated circuit chip provided by a specific silicon provider, an operating system vendor (OSV) that provides an operating system to be executed on the device, and a cloud platform operator (or cloud host) that maintains a server farm providing server space for hosting virtual machines on the cloud. Therefore, there can be problems when implementing the regions in a strictly privilege-increasing order. That is, an application provider providing application-level code desires a secure computing environment but may not want to trust the parties (OSV, OEM, or cloud host) that would have provided software to run in the secure region conventionally.However, similarly, a party providing code that operates in a secure area would not likely want to entrust an application provider with providing code that operates in a higher privilege area that has been granted access to data associated with a lower privilege area. Therefore, it has been recognized that a strict hierarchical area where privileges are continuously increased may not be appropriate.

[0021] Therefore, in the more detailed examples below, additional areas can be considered orthogonal to the secure area. The additional area and the secure area can each access a less secure physical address space, but the additional physical address space associated with the additional area is inaccessible from the secure area, and at the same time, the secure physical address space associated with the secure area is inaccessible from the additional area. The root area can still access the physical address space associated with both the secure area and the additional area.

[0022] Therefore, in this model, the additional area (an example of which is the realm area described in the examples below) and the secure area are independent of each other and do not need to trust each other. The secure area and the additional area only need to trust the root area, and the root area is inherently trusted because it manages access to other areas.

[0023] The examples below illustrate a single instance of an additional area (the realm area), but it will be understood that the principle of the additional area being orthogonal to the secure area can be extended to provide multiple additional areas. As a result, each of the secure area and at least two additional areas can access a less secure physical address space, cannot access the root physical address space, and cannot access the physical address space associated with each other.

[0024] A less secure physical address space may be accessible from all regions supported by the processing circuitry. This is useful for facilitating sharing of data or program code between software running in different regions. If a particular item of data or code is to be made accessible from different regions, it can be allocated to the less secure physical address space so that it is accessible from any region.

[0025] When translating a virtual address to a physical address, the address translation circuitry may perform the translation based on at least one page table entry. If at least the current region is one of a subset of at least three regions supported by the processing circuitry, the address translation circuitry, based on the current region and physical address space selection information specified within at least one page table entry used for translating the virtual address to a physical address, can select which physical address space should be used as the physical address space to which the physical address for a given memory access is to be translated. Thus, the information defined within the page table structure can affect which physical address space is selected for a memory access when the given memory access is issued from the current region. In some regions, this selection based on the physical address space selection information specified in the page table entry may not be necessary. For example, if the current region is the less secure region described above, since the less secure region is not accessible from all other address spaces, the less secure physical address space can be selected regardless of any information specified within at least one page table entry used for address translation.

[0026] However, for other regions, it is possible for that region to be selected between two or more different physical address spaces. Thus, for these regions, it may be useful to define in the page table entry information about a given block of addresses indicating which physical address space should be used for that access, whereby different portions of the virtual address space seen by a given software can be mapped onto different physical addresses.

[0027] For example, when the current region is the root region, the address translation circuit may translate a virtual address to a physical address based on a root region page table entry that includes at least 2 bits of physical address space selection information for selecting between at least 3 physical address spaces accessible from the root region. For example, in an implementation that supports a root region, a less secure region, and a secure region, the physical address space selection information in the root region page table entry can select between any of these 3 physical address spaces. In an implementation that also has at least 1 further region, the physical address space selection information can select between any of the root physical address space, the secure physical address space, the less secure physical address space, and at least 1 further physical address space.

[0028] On the one hand, when the current region is a secure region or a further region, the selection of the physical address space can be more restricted, and thus less physical address space selection information in terms of the number of bits may be required compared to the root region. For example, in a secure region, the physical address space selection information can be selected between a secure address space and a less secure address space (since the root physical address space and the further physical address space may be inaccessible). When the current region is a further region, the physical address space selection information can be used to select between a further physical address space and a less secure physical address space since the secure physical address space and the root physical address space may be inaccessible. In the case of a page table entry used to select the physical address space used when the current region is a secure region or a further region, the physical address space selection indicator used to make this selection can be encoded at the same position within at least one page table entry regardless of whether the current region is a secure region or a further region. This makes the encoding of the page table entry more efficient, enables the hardware to interpret that part of the page table entry as being reused for both the secure region and the further region, and reduces the circuit area.

[0029] The memory system may include a physical alias point (PoPA) which is the point at which an aliased physical address from a different physical address space corresponding to the same memory system resource is mapped to a single physical address that uniquely identifies that memory system resource. The memory system may include at least one PoPA pre-memory system component provided upstream of the PoPA, which treats the aliased physical addresses as if they corresponded to different memory system resources.

[0030] For example, at least one PoPA pre-memory system component can include a cache or translation lookaside buffer, which can cache data, program code, or address translation information for aliased physical addresses in separate entries, so that when the same memory system resource is required to be accessed from different physical address spaces, the access causes a different cache or TLB entry to be allocated. Also, the PoPA pre-memory system component can include a coherence control circuit such as a coherent interconnect, snoop filter, or other mechanism for maintaining coherence between cache information at each master device. The coherence control circuit can assign separate coherence states to each aliased physical address within different physical address spaces. Thus, aliased physical addresses are treated as separate addresses for the purpose of maintaining coherence, even if they actually correspond to the same underlying memory system resource. At first glance, tracking coherence separately for aliased physical addresses might seem to cause a problem of coherence loss, but in fact, this is not a problem because if processes operating in different regions are intended to actually share access to a particular memory system resource, they can access that resource using a less secure physical address (or use the limited sharing feature described below to access the resource using one of the other physical address spaces). Another example of a PoPA pre-memory system component can be a memory protection engine provided to protect data stored in off-chip memory from loss of confidentiality and / or tampering. Such a memory protection engine can, for example, separately encrypt data associated with a particular memory system resource using different encryption keys depending on the physical address space from which the resource was accessed, treating the aliased physical address as if it corresponded to a different memory system resource (e.g., an encryption scheme that depends on the address can be used, and the physical address space identifier can be considered part of the address for this purpose).

[0031] Regardless of the form of the PoPA pre-memory system components, it may be useful to treat such PoPA memory system components as if the aliased physical addresses corresponded to different memory system resources. This is because it provides hardware-implemented isolation between accesses issued to different physical address spaces, and as a result, information associated with one region does not leak into another region due to characteristics such as cache timing side channels or side channels being triggered by the coherence control circuit and accompanied by changes in coherence.

[0032] In some implementations, the aliased physical addresses within different physical address spaces may be able to be represented using different numerical physical address values for each different physical address space. This approach may require a mapping table in PoPA to determine which different physical address values correspond to the same memory system resource. However, this overhead of maintaining the mapping table may be considered unnecessary, and in some implementations, it may be simpler if the aliased physical addresses include physical addresses represented using the same numerical physical address value in each of the different physical address spaces. When this approach is taken, at the physical aliasing point, it may be sufficient to simply discard the physical address space identifier that identifies which physical address space is being accessed using the memory access, and then provide the remaining physical address bits downstream as the unaliased physical address.

[0033] Accordingly, in addition to the PoPA front memory system components, the memory system may also include PoPA memory system components configured to de-alias a plurality of aliased physical addresses to obtain an aliased physical address provided to at least one downstream memory system component. The PoPA memory system component can be a device that accesses a mapping table to find the de-aliased address corresponding to the aliased address within a particular address space, as described above. However, the PoPA component may simply be a location within the memory system where the physical address tag associated with a given memory access is discarded, such that the physical address provided downstream uniquely identifies the corresponding memory system resource regardless of which physical address space it was provided from. Alternatively, in some cases, the PoPA memory system component may still provide a physical address space tag to at least one downstream memory system component (e.g., for the purpose of enabling completer-side filtering as further discussed below). However, PoPA marks a point within the memory system such that downstream memory system components no longer treat aliased physical addresses as different resources and can map the same memory system resource considering each of the aliased physical addresses. For example, if a memory controller or a hardware memory storage device downstream of PoPA receives the physical address tag and physical address of a given memory access request, if the physical address corresponds to the same physical address as a previously seen transaction, any hazard checks or performance improvements performed for each transaction accessing the same physical address (such as merging accesses to the same address) can be applied, even if each transaction specified a different physical address space tag. In contrast, for memory system components upstream of PoPA, such hazard checks or performance improvement steps taken for transactions accessing the same physical address may not be triggered if these transactions specify the same physical address within different physical address spaces.

[0034] As described above, at least one pre-PoPA memory system component may include at least one pre-PoPA cache. This can be a data cache, an instruction cache, or a unified level 2, level 3, or system cache.

[0035] The processing circuit can support a cache invalidation instruction up to a PoPA that specifies a target address (which can be a virtual address or a physical address). In response to the cache invalidation instruction up to the PoPA, the processing circuit can issue at least one invalidation command to require that at least one pre-PoPA cache invalidate one or more entries associated with a target physical address value corresponding to the target address. In contrast, when at least one invalidation command is issued, at least one post-PoPA cache located downstream of the PoPA may be permitted to retain one or more entries associated with the target physical address value. For at least one pre-PoPA cache, the cache can invalidate one or more entries associated with the target physical address value specified by at least one invalidation command regardless of which physical address space is associated with those entries. Thus, even if physical addresses having the same address value within different physical address spaces were being treated as if they represented different physical addresses by the pre-PoPA cache, for the purpose of processing the invalidation triggered by the cache invalidation instruction up to the PoPA, the physical address space identifier can be ignored.

[0036] Therefore, it is possible to define a form of cache invalidation instruction that causes the processing circuit to require that any cache entry associated with a specific physical address corresponding to a target virtual address be invalidated in any cache up to the physical aliasing point. This form of invalidation instruction may differ from other types of invalidation instructions that require the invalidation of cache entries that affect caches up to other points in the memory system, such as the coherence point (the point at which all observers (e.g., processor cores, direct memory access engines, etc.) are guaranteed to see the same copy of data associated with a given address). Providing a dedicated instruction form that requires invalidation up to the physical aliasing point can be useful, especially in the case of root region code that can manage changes to the address assignment for each region. For example, when updating the granularity protection information that defines which physical addresses are accessible within a given physical address space, or when redistributing a specific block of physical addresses to a different physical address space, the code for the root region can use a cache invalidation instruction up to the PoPA to ensure that any data, code, or other information present in the cache that depends on the old granularity protection information for accessibility is invalidated and subsequent memory accesses are correctly controlled based on the new granularity protection information. In some examples, in addition to invalidating the cache entry, at least one pre-PoPA cache can also delete the data from that cache entry and write any dirty divergences of the data associated with the invalidated entry to a location within the memory system beyond the PoPA. In some cases, different versions of the cache invalidation instruction up to the PoPA are supported and can indicate whether deletion is required.

[0037] A memory encryption circuit that responds to a memory access request specifying a selected physical address space and a target physical address within the selected physical address space may be provided to encrypt or decrypt data associated with a protected area based on one of several encryption keys selected according to the selected physical address space when the target physical address is within a protected address area. In some examples, the protected address area may be the entire physical address space, although in other examples, encryption / decryption may be applied only to a specific sub-area as the protected address area. By allocating a dedicated root physical address space separate from the physical address space associated with other areas, it simplifies for the memory encryption circuit to select a different encryption key for the root area compared to other areas to improve security. Similarly, by selecting different encryption keys for all other areas, it enables stronger isolation of code or data assets associated with a specific area.

[0038] In one particular example, the device may have at least one on-chip memory on the same integrated circuit as the processing circuit, and all valid physical addresses within the root physical address space may be mapped to the at least one on-chip memory such that they are separate from the off-chip memory. This helps to improve the security of the root area. It will be understood that information from other areas can also be stored in the on-chip memory. Providing a separate root physical address space simplifies the memory allocation. In an example where the root area shares a secure physical address space with a secure area, it may be difficult to hold all the data associated with the secure area in the on-chip memory as there may be too much of it, and it may be difficult to determine which specific data is associated with the root area. In contrast, it is much easier to split and retrieve that data (or code) when the data (or code) of the root area is flagged with a different physical address space identifier.

[0039] However, in other examples, some addresses within the root physical address space may be mapped to off-chip memory. To protect the root region data stored off-chip, mechanisms for memory encryption, integrity, and freshness may be used.

[0040] The above technology can be implemented in a hardware device having hardware circuit logic for implementing the functions as described above. Therefore, the processing circuit and the address conversion circuit may include hardware circuit logic. However, in other examples, a computer program that controls a host data processing device to provide an instruction execution environment for executing target code may provide processing program logic and address conversion program logic that execute functions equivalent to the above-described processing circuit and address conversion circuit in software. This may be useful, for example, to enable target code written for a particular instruction set architecture to be executed on a host computer that may not support that instruction set architecture. Thus, simulation software can emulate the functions expected by an instruction set architecture not provided by the host computer by providing an equivalent instruction execution environment for the target code as would be expected if the target code were executed on a hardware device that actually supports the instruction set architecture. Therefore, a computer program that provides simulation can include processing program logic that simulates processing in at least one of the aforementioned at least three regions, and address conversion program logic that converts a virtual address to a physical address within one of several simulated physical address spaces selected based at least on the current region. Similar to in a hardware device, the at least three regions may include a root region for managing switching between other regions, and the root region may have an associated root simulated physical address space separate from the simulated physical address spaces associated with other regions. In the case of the method by which simulation of an architecture is provided, each physical address space selected by the address conversion program logic is a simulated physical address space because, although they do not actually correspond to the physical address spaces identified by the hardware components of the host computer, they are mapped to addresses within the virtual address space of the host.Providing such a simulation can be useful for various purposes, for example, making old code written for one instruction set architecture executable on different platforms that support different instruction set architectures, or assisting in the software development of new software that is to be executed for a version of an instruction set architecture when hardware devices that support the new version of the instruction set architecture are not yet available (thereby making it possible to start the development of software for the new version of the architecture in parallel with the development of hardware devices that support the new version of the architecture).

[0041] Granule protection lookup In a system that can map a virtual address of a memory access request to a physical address in one of two or more separate physical address spaces, granule protection information can be used to restrict which physical addresses are accessible within a particular physical address space. This can be useful to ensure that access to a particular physical memory location implemented in either on-chip or off-chip hardware can be restricted to a particular physical address space, or, if desired, a particular subset of the physical address space.

[0042] In one approach for managing such restrictions, implementations where a given physical address can be accessed regardless of the particular physical address space can be implemented using a host-side filtering circuit provided in or near a host device for processing memory access requests. For example, the host-side filtering circuit can be associated with a memory controller or a peripheral controller. In such an approach, issuing a memory access request to a cache or an interconnect for routing transactions from a requester device to a host device may not depend on any lookup of information defining which physical addresses are accessible within a given physical address space.

[0043] In contrast, in the embodiments described below, a granularity protection lookup is performed by a requester-side filtering circuit that checks whether a memory access request can be passed to a cache or an interconnect based on a lookup of granularity protection information indicating at least one permitted physical address space associated with a target granularity of the accessed physical address. The granularity of the physical address space defined by each item of granularity protection information may be of a particular size that may be the same as or different from the size of a page used in a page table structure used by an address translation circuit. In some cases, the granularity may be larger than the size of a page that defines an address translation mapping of an address translation circuit. Alternatively, the granularity protection information may be defined at the same page-level granularity as the address translation information within a page table structure. Defining the granularity protection information at the page-level granularity may be convenient as it allows for finer control over which regions of the memory storage hardware are accessible from a particular physical address space and thus from a particular operating region of the processing circuit.

[0044] Accordingly, the apparatus can have an address translation circuit for translating a target virtual address specified by a memory access request issued by a requester circuit into a target physical address, and a requester-side circuit for performing a granularity protection lookup based on the target physical address and a selected physical address space associated with the memory access request, thereby determining whether to allow the memory access request to be passed to a cache or passed to an interconnect to communicate with a completer device to process the memory access request. The selected physical address space can be one of a plurality of physical address spaces. In the granularity protection lookup, the requester-side filtering circuit obtains granularity protection information corresponding to the target granularity of the physical address including the target physical address, the granularity protection information indicating at least one permitted physical address space associated with the target granularity, can be configured to block the memory access request when the granularity protection information indicates that the selected physical address space is not one of at least one permitted physical address spaces.

[0045] The advantage of performing the granularity protection lookup on the requester side rather than on the completer side of the interconnect is that it allows for more fine-grained control over which physical addresses are accessible from a given physical address space than is practical on the completer side. This is because the completer side may typically have relatively limited ability to access the entire memory system. For example, the memory controller of a given memory unit may be made accessible only to locations within that memory unit and may not be able to access other regions of the address space. Providing more fine-grained control may rely on a more complex table of granularity protection information that can be stored in the memory system, and it may be more practical to access such a table from the requester side, which has more flexibility in issuing memory access requests to a wider subset of the memory system.

[0046] Also, performing the granule protection lookup on the requester side can help enable the ability to dynamically update granule protection information during execution, which may not be practical for a completer side filtering circuit that may be limited to accessing a relatively small amount of statically defined data defined at startup.

[0047] Another advantage of the requester side filtering circuit is that it allows the interconnect to distribute different addresses within the same granule to different completer side ports that communicate with different completer side devices (e.g., different DRAM (Dynamic Random Access Memory) units), which is performance efficient, but would be impractical if the entire granule had to be directed to the same completer unit so that the granule protection lookup can be performed on the completer side to verify if a memory access is possible.

[0048] Therefore, there can be many advantages to performing the granule protection lookup, which determines whether access can be made from a particular physical address space to a particular physical address for a given memory access request, on the requester side rather than on the completer side.

[0049] Granule protection information can be represented in various ways. In one example, granule protection can be defined by a single linearly indexed table stored in a single contiguous address block, where a particular entry accessed within the block is selected based on the target physical address. However, in practice, granule protection information may not be defined for the entire physical address space, and thus, it may be more efficient to use a multi-level table structure for storing granule protection information, similar to the multi-level page tables used for address translation. In such a multi-level structure, a level 1 granule protection table entry can be selected that uses a portion of the target physical address to provide a pointer identifying a location in memory where a further level of granule protection table is stored. Then, another portion of the target physical address can be used to select which entry of that further granule protection table should be retrieved. After sequentially repeating between one or more levels beyond the first level of the table, ultimately, a granule protection table entry can be retrieved that provides the granule protection information associated with the target physical address.

[0050] Regardless of the particular structure selected for the table storing granule protection information, granule protection information can be represented in several ways as to which portions of the physical address space are at least one permitted physical address. One approach can be to provide a series of fields each indicating whether one of the corresponding physical address spaces can access the granule of the physical address containing the target physical address. For example, a bitmap may be defined within the granule protection information, and each bit of the bitmap indicates whether the corresponding physical address space is a permitted physical address space for that granule or a non-permitted physical address space.

[0051] However, in practice, in most usage cases, the likelihood that a significant number of physical address spaces will be permitted to access a given physical address is likely to be relatively low. As discussed in the section before controlling access to the physical address space, less secure physical address spaces are available for selection in all regions and can thus be used when data or code is shared among several regions. Therefore, it may not be necessary for a particular physical address to map to all or many parts of the available physical address space.

[0052] Therefore, a relatively efficient approach is that the granularity protection information can specify the assigned physical address space assigned to the target granularity of the physical address, and at least one permitted physical address space can include at least the assigned physical address space indicated by the granularity protection information for that particular target granularity. In some implementations, the granularity protection information can specify a single physical address space as the assigned physical address space. Thus, in some cases, the granularity protection information can include an identifier of one particular physical address space that functions as the assigned physical address space permitted to access the target granularity of the physical address.

[0053] In some implementations, the only physical address space permitted to access the target granularity of a physical address may be the allocated physical address space, and the target granularity of the physical address may not be permitted to be accessed from any other physical address space. This approach can be efficient for maintaining security. Access to a particular physical address space from different regions can instead be controlled through an address translation function, in which case the address translation circuitry can choose which particular physical address space should be used for a given memory access, so there may be no need to share the granularity of the physical address among multiple physical address spaces. If only the allocated physical address space is permitted to access the target granularity of the physical address, it may be necessary to update which physical address space is the allocated physical address space in order to enable that granularity of the physical address to be accessed from other physical address spaces. For example, this may require the aforementioned root region to perform some process for switching the physical address space allocated for a given granularity of a physical address. This process may have a particular performance cost. That may include, for example, overwriting each location within a given granularity of the physical address (for security) with NULL data or other data unrelated to the previous contents of those physically addressed locations. Thereby, it can be ensured that a process accessing the newly allocated physical address space cannot learn anything from the data previously stored at locations associated with a given granularity of the physical address.

[0054] Therefore, another approach can be that not only the allocated physical address space is identified, but the granularity protection information can also include sharing attribute information indicating whether at least one other physical address space other than the allocated physical address space is one of the at least one permitted physical address spaces. Thus, when the sharing attribute information indicates that at least one other physical address space is permitted to access the corresponding granularity of the physical address, that granularity of the physical address can be accessed from multiple physical address spaces. This can be useful to temporarily enable code within a region associated with one physical address space to be visible from a region associated with a different physical address space among the allocated granularities of the physical address. This can more efficiently enable temporary sharing of data or code because it does not require potentially costly operations to change which physical address space is the allocated physical address space. The sharing attribute information can be set directly by code running within the region associated with the allocated physical address space or can be set by the root region in response to a request from code running within the region associated with the allocated physical address space.

[0055] When shared attribute information is supported, not only is it used to check whether an address assigned to one physical address space can be accessed by a request specifying a different address space, but the requester-side filtering circuit can also transform the physical address space selected for a memory access request issued to a downstream cache or interconnect based on the shared attribute information. Thus, if the granularity protection lookup determines that the selected physical address space is a physical address space other than the one assigned that is one of the at least one permitted physical address spaces indicated by the shared attribute information, the requester-side filtering circuit can cause the memory access request to be passed to the cache or interconnect specifying the assigned physical address space instead of the selected physical address space. This means that for the purpose of accessing the downstream memory, the component before the PoPA handles the memory access as if it had been issued from the beginning specifying the physical address space to which it is assigned, such that a cache entry or snoop filter entry tagged with that assigned physical address space can be accessed for the memory access request.

[0056] In some implementations, the requester-side filtering circuit may obtain the granularity protection information used for the granularity protection lookup from the memory each time a memory access request is checked against the granularity protection information. This approach can result in lower hardware costs that may be required on the requester side. However, obtaining the granularity protection information from the memory can be relatively time-consuming.

[0057] Accordingly, to improve performance, the requester-side filtering circuit may have access to at least one lookup cache capable of caching the granularity protection information. As a result, the granularity protection lookup can be performed within at least one lookup cache, and if the required granularity information is already stored in at least one lookup cache, there is no need to fetch it from memory. In some cases, the at least one lookup cache may be a cache separate from the translation lookaside buffer (TLB) used by the address translation circuit for cache page table data that provides a mapping between virtual and physical addresses. However, in other examples, the at least one lookup cache can combine a cache of page table data with a cache of granularity protection information. Thus, in some cases, the at least one lookup cache can store at least one combined translation and granularity protection entry that specifies information according to both the granularity protection information and at least one page table entry used by the address translation circuit to map the target virtual address to the target physical address. Whether to implement the TLB and the granularity protection cache as separate structures or as a single combined structure is an implementation choice and either can be used.

[0058] Regardless of which technique is used for at least one look-up cache, at least one look-up cache can respond to at least one look-up cache invalidation command. The look-up cache invalidation command can invalidate a look-up cache entry that stores information depending on the granularity protection information associated with the granularity of physical addresses including the invalidation target physical address by specifying the invalidation target physical address. In a conventional processing system having a TLB, the TLB usually supports an invalidation command that specifies a virtual address, or (in a system supporting two-stage address translation, an intermediate address), but it is not usually necessary for the TLB to be able to identify which entry is invalidated using the physical address. However, when at least one look-up cache is provided to cache granularity protection information, it may be useful to be able to invalidate any entry that depends on that information if the granularity protection information for a given granularity of physical addresses changes. Thus, the command can identify the specific physical address at which entries containing information depending on the granularity protection information are invalidated.

[0059] When the granularity protection information cache is implemented separately from the TLB, the TLB may not need the ability to search for entries by physical address. In this case, the cache of granularity protection information can respond to a cache invalidation command that specifies a physical address, but the command can be ignored by the TLB.

[0060] However, it may be useful to provide an additional scheme for looking up entries based on a physical address when at least one lookup cache includes a combined translation / granule protection cache, whose entries are looked up based on a virtual address or an intermediate address and return both page table information associated with the virtual / intermediate address and granule protection information associated with the corresponding physical address. Thereby, at least one lookup cache invalidation command specifying an invalidation target physical address can be processed. Such a lookup by physical address may not be necessary in the normal lookup of the combined cache. This is because if the entries are combined, a lookup by virtual address or by intermediate address may be sufficient to access all of the combined information for performing both the lookup for address translation and the lookup for granule protection. However, since the cache invalidation command depends on the granule protection information for the specified physical address, the combined cache can be looked up based on the physical address, thereby identifying any entries that need to be invalidated.

[0061] As described above for the previous embodiments, the memory system may have PoPA memory system components, at least one PoPA pre-memory system component, and at least one PoPA post-memory system component. Thus, aliased physical addresses in different physical address spaces may correspond to the same memory system resources identified using the non-aliased physical address when the memory access request exceeds the physical aliasing point, in the same manner as described above. Before PoPA, at least one PoPA pre-memory system component may handle aliased physical addresses from different physical address spaces as if they corresponded to different memory system resources, thereby improving security. Also in this case, theoretically, it may be possible to identify physical addresses in different physical address spaces using different numerical address values, but this may be relatively complex to implement and may be simpler when the aliased physical addresses are represented using the same physical address value in different physical address spaces.

[0062] When at least one PoPA pre-memory system component includes at least one PoPA pre-cache, the processing circuit, in response to a cache invalidation instruction up to the PoPA specifying the target virtual address, can trigger invalidation by the target physical address while allowing any PoPA post-cache (as described above) to hold data having the target physical address for any PoPA pre-cache upstream of the physical aliasing point.

[0063] The selected physical address space associated with the memory access can be selected in various ways. In some examples, the selected physical address space can be selected (by an address translation circuit or a requester-side filtering circuit) based at least in part on the current region of operation of the requester circuit that issued the memory access request. The selection of the selected physical address space can also depend on physical address space selection information specified in at least one page table entry used for translating the target virtual address to the target physical address. The selection of which physical address space is the selected physical address space can be performed as described above for the previous embodiments.

[0064] The regions and physical address spaces available for selection within a given system may be as described above and, as described above, may include less secure regions, secure regions, root regions, and further regions, each having a corresponding physical address space. Alternatively, the regions / physical address spaces may include a subset of these regions. Thus, any features associated with any of the aforementioned regions may be included in a system having a requester-side filtering circuit.

[0065] In an implementation where the root region has a corresponding root physical address space as described above, when the current region is the root region, the requester-side filtering circuit can bypass the granularity protection lookup. Since the root region can be trusted to access all ranges of physical addresses, the granularity protection lookup may not be necessary when the current region is the root region, and thus power can be saved by skipping the granularity protection lookup when in the root region.

[0066] The granule protection information may be modifiable by software executed in the root region. Thus, the granule protection information may be dynamically updatable during execution. This can be an advantage for some of the discussed realm usage scenarios, where a realm providing a secure execution environment is dynamically created during execution and the corresponding area of the allocated memory can be reserved for the realm. Using only requester-side filtering, such an approach can often be impractical. In some implementations, the root region may be the only region permitted to modify the granule protection information, such that if other regions need to modify the implemented granule protection information, the other regions can request that the root region modify the granule protection information, and the root region can then check whether to grant a request made from another region.

[0067] It may be beneficial to provide a requester-side filtering circuit for performing a granule protection lookup on the requester side before a memory access request is passed to the cache or interconnect, although there may also be other scenarios where it may be preferable for the protection information (defining which physical addresses can be accessed from a given physical address space) to be checked instead on the completer side of the interconnect. For example, in some parts of the address space, it may be desirable to provide the fine-grained page-level details of the memory hardware split into areas of physical addresses accessible from different physical address spaces, while for other parts of the memory system, it may be preferable to allocate large blocks of contiguous addresses to a single physical address space, in which case the overhead of accessing the (potentially multi-level) granule protection structure stored in the memory may not be justified. If an entire memory unit (e.g., a particular DRAM module) is allocated to a single physical address space, it may be simpler to handle the enforcement of access restrictions to that memory unit via checks on the completer side.

[0068] Thus, in some implementations, there may be not only a requester-side filtering circuit, but also a completer-side filtering circuit that receives from the interconnect and responds to memory access requests that specify a target physical address and a selected physical address space. The completer-side filtering circuit performs a completer-side protection lookup of completer-side protection information based on the target physical address and the selected physical address space to determine whether the memory access request is permitted to be processed by the completer device. By providing a hybrid approach that allows a portion of the memory to be protected via requester-side filtering and another portion to be protected by completer-side filtering, a balance is enabled between better performance and flexibility in memory usage allocation than can be achieved through either requester-side filtering or completer-side filtering alone.

[0069] Thus, in some implementations, the granularity protection information can specify a pass-through indicator indicating that at least one permitted physical address space should be resolved by the completer-side filtering circuit.

[0070] Therefore, when the granule protection information designates a pass-through indicator, the requester-side filtering circuit can determine whether to pass the memory access request to the cache or the interconnect regardless of any check as to whether the selected physical address space is one of at least one permitted physical address spaces of the target granule of the physical address. On the other hand, when the granule protection information accessed for the target granule does not designate a pass-through indicator, the determination as to whether the memory access request can be passed to the cache or the interconnect may, in this case, depend on a check as to whether the selected physical address space is one of at least one permitted physical address spaces because there is a possibility that subsequent completer-side filtering is not performed after the memory access request is permitted to proceed to the cache or the interconnect. Therefore, the pass-through indicator can control the division of the address space between the granule of the physical address for which the check should be executed on the requester side and the granule for which these checks should be executed on the completer side, providing additional flexibility to the system designer.

[0071] The computer-side protection information used by the computer-side filtering circuit does not necessarily have the same format as the granularity protection information used by the requester-side filtering circuit. For example, the computer-side filtering circuit may be defined with coarser details than the granularity protection information used for granularity protection lookups by the requester-side filtering circuit. The granularity protection information may be defined in a multi-level table structure where each level of the table provides an entry corresponding to a block of memory with a given number of addresses corresponding to a power of two. Thus, an entry at a given level of the table required to check a given target physical address can be indexed simply by adding a plurality of a portion of specific bits from the target physical address to a base address associated with that level of the table, avoiding the need to compare the content of the accessed entry with the target physical address to determine whether it is the correct entry. In contrast, in the computer-side protection information, a smaller number of entries may be defined, and each entry may specify the start and end addresses (or start address and size) of a region of memory that may correspond to a number of addresses other than a power of two. This may be more suitable for defining relatively coarse-grained blocks in the computer-side protection information, but this approach may require comparing the target physical address with the upper and lower limits of each range of physical addresses defined within each computer-side protection information entry to determine whether any of them match the specified physical address. The approach in the indexed multi-level table used for the requester-side granularity protection information can support a relatively large number of distinct entries, thereby supporting a fine-grained mapping of the physical address space of the physical address. This is usually not practical in the case of using the approach of defining the upper and lower limits of each region in the computer-side protection information due to the comparison overhead in looking up each of these entries to check whether the target address is included within the boundaries of that entry.However, the completer-side protection information can be more efficient from the perspective of memory storage, and the penalty when there is a miss in the lookup cache is also lighter, so the fluctuations regarding performance can be smaller. Of course, this is just an example of how the lookup information is implemented on the requester side and the completer side.

[0072] The requester-side protection information can be dynamically updated during execution. The completer-side protection information can be statically defined by the hardware on the system-on-chip, configured at startup, or dynamically reconfigured during execution.

[0073] Regarding the foregoing embodiments, the technology described above for granular protection lookup can be implemented in a system having dedicated hardware logic for performing the functions of the address translation circuit and the requester-side filtering circuit. However, equivalent functions can also be implemented in software within a computer program to control the host data device and provide an instruction execution environment for the execution of target code for the same reasons as described above. Therefore, address translation program logic and filtering program logic can be provided to emulate the functions of the foregoing address translation circuit and requester-side filter. Regarding the foregoing embodiments, at least one of the following can apply to the computer program that provides the instruction execution environment. The granular protection information can be dynamically updated during execution by the target code. The granular protection information is defined at the page-level granularity.

[0074] Description of Embodiments Figure 1 schematically shows an example of a data processing system 2 having at least one requester device 4 and at least one completer device 6. An interconnect 8 provides communication between the requester device 4 and the completer device 6. The requester device can issue a memory access request that requests memory access to a location in a specific addressable memory system. The completer device 6 is a device responsible for processing a memory access request directed thereto. Although not shown in Figure 1, some devices may be able to function as both a requester device and a completer device. The requester device 4 may include, for example, a processing element such as a central processing unit (CPU) or a graphics processing unit (GPU) or other master devices such as a bus master device, a network interface controller, a display controller, etc. The completer device includes, for example, a memory controller responsible for access control to a corresponding memory storage unit, a peripheral controller for controlling access to peripheral devices, and the like. Figure 1 shows a more detailed configuration example of one of the requester devices 4, but it should be understood that other requester devices 4 may have a similar configuration. Alternatively, other requester devices may have a configuration different from the requester device 4 shown on the left side of Figure 1.

[0075] The requester device 4 has a processing circuit 10 for executing data processing in response to an instruction by referring to the data stored in the register 12. The register 12 may include not only a general-purpose register for storing operands and the results of processed instructions, but also a control register for storing control data for configuring how the processing is executed by the processing circuit. For example, the control data may include a current region indication 14 used to select which operating region is the current region, and a current exception level indication 15 indicating which exception level is the current exception level at which the processing circuit 10 is operating.

[0076] The processing circuit 10 may be capable of issuing a memory access request that specifies a virtual address (VA) that identifies an addressable location to be accessed and a region identifier (region ID or "security state") that identifies the current region. The address translation circuit 16 (e.g., a memory management unit (MMU)) converts the virtual address to a physical address (PA) via one of more stages of address translation based on page table data defined in a page table structure stored in the memory system. The translation lookaside buffer (TLB) 18 functions as a lookup cache for caching a portion of the page table information for faster access than would be required if the page table information had to be fetched from memory each time an address translation is needed. In this example, in addition to generating the physical address, the address translation circuit 16 also selects one of several physical address spaces associated with the physical address and outputs a physical address space (PAS) identifier that identifies the selected physical address space. The selection of the PAS is discussed in more detail below.

[0077] The PAS filter 20 functions as a requester-side filtering circuit for verifying whether the physical address can be accessed within the physical address space identified and specified by the PAS identifier based on the converted physical address and the PAS identifier. This lookup is based on the granule protection information stored in the granule protection table structure stored in the memory system. Similar to the cache of the page table data in the TLB 18, the granule protection information can be cached in the granule protection information cache 22. In the example of FIG. 1, the granule protection information cache 22 is shown as a separate structure from the TLB 18, but in other examples, these types of lookup caches can be combined as a single lookup cache structure, and as a result, a single lookup of the entries in the combined structure provides both the page table information and the granule protection information. The granule protection information defines information that restricts the physical address space in which a given physical address can be accessed, and based on this lookup, the PAS filter 20 determines whether to proceed with issuing a memory access request to one or more caches 24 and / or the interconnect 8. If the specified PAS for the memory access request is not permitted to access the specified physical address, the PAS filter 20 can block the transaction and signal an error.

[0078] FIG. 1 shows an example of a state where the system has multiple requester devices 4, but the features shown for one of the requesting devices on the left side of FIG. 1 can also be included in a system with only one requester device, such as a single-core processor.

[0079] FIG. 1 shows an example where the selection of a PAS for a given request is performed by the address translation circuit 16. In other examples, however, information for determining which PAS to select can be output by the address translation circuit 16 to the PAS filter 20 together with the PA, and the PAS filter 20 can select a PAS and check whether the PA can be accessed within the selected PAS.

[0080] The provision of the PAS filter 20 helps support a system that can operate in several operating regions, each associated with its own isolated physical address space. Here, at least a portion of the memory system (e.g., for a coherence implementation mechanism such as some caches or snooping filters) treats separate physical address spaces as if they were a completely different set of addresses identifying locations in a different memory system, even if the addresses within those address spaces actually point to the same physical location within the memory system. This can be useful for security purposes.

[0081] FIG. 2 shows examples of various operating states and regions in which the processing circuit 10 can operate, as well as examples of the types of software that can be executed in various exception levels and regions (it will of course be understood that the specific software installed on the system is selected by the party managing the system and is thus not an essential feature of the hardware architecture).

[0082] The processing circuit 10 is operable at several different exception levels 80, in this example, four exception levels labeled EL0, EL1, EL2, and EL3, and in this example, EL3 refers to the exception level with the highest level of privilege and EL0 refers to the exception level with the lowest privilege. In other architectures, it will be understood that the reverse numbering may be chosen and the exception level with the largest number may be considered to have the lowest privilege. In this example, the lowest privilege exception level EL0 is for application level code, the next highest privilege exception level EL1 is used for operating system level code, the next highest privilege exception level EL2 is used for hypervisor level code that manages switching between several virtual operating systems, and the highest privilege exception level EL3 is used for monitor code that manages switching between the respective regions and the allocation of physical addresses to the physical address space, as will be described later.

[0083] When an exception occurs at a particular exception level while processing software, for some types of exceptions, the exception is accepted at a higher (more privileged) exception level, and the particular exception level at which the exception is accepted is selected based on the attributes of the particular exception that occurred. However, in some situations, other types of exceptions may be accepted at the same exception level as the exception level associated with the code being processed when the exception was accepted. When an exception is accepted, information characterizing the state of the processor at the time the exception was accepted can be saved, including, for example, the current exception level at the time the exception was accepted. Thus, when an exception handler is processed to handle the exception, the processing can return to the previous processing and the saved information can be used to identify the exception level to which the processing should return.

[0084] In addition to the different exception levels, the processing circuit also supports several operating regions including a root region 82, a secure (S) region 84, a less secure region 86, and a realm region 88. For ease of reference, the less secure region is hereinafter described as the "non-secure" (NS) region, although it should be understood that this is not intended to imply a particular level (or lack) of security. Instead, "non-secure" simply indicates that the non-secure region targets code that is less secure than the code operating in the secure region. The root region 82 is selected when the processing circuit 10 is at the highest exception level EL3. When the processing circuit is at one of the other exception levels EL0 to EL2, the current region is selected based on the current region indicator 14, which indicates which of the other regions 84, 86, 88 is active. For each of the other regions 84, 86, 88, the processing circuit can be at any of the exception levels EL0, EL1, or EL2.

[0085] At startup, some boot code (e.g., BL1, BL2, OEM boot) can be executed within a higher privilege exception level, e.g., EL3 or EL2. The boot codes BL1, BL2 can be associated with, for example, the root region, and the OEM boot code can operate in the secure region. However, once the system is started and during execution, the processing circuit 10 can be considered to operate in one of the regions 82, 84, 86, and 88 at a time. Each of the regions 82 to 88 is associated with its own related physical address space (PAS). This enables the isolation of data from different regions within at least a part of the memory system. This will be described in more detail below.

[0086] The non-secure region 86 can be used for normal application-level processing and operating system and hypervisor activities for managing such applications. Thus, within the non-secure region 86, there may be application code 30 operating at EL0, operating system (OS) code 32 operating at EL1, and hypervisor code 34 operating at EL2.

[0087] The secure region 84 enables specific system-on-chip security, media, or system services to be isolated in a physical address space separate from the physical address space used for non-secure processing. The secure region and the non-secure region are not equivalent in the sense that non-secure region code cannot access resources associated with the secure region 84, while the secure region can access both secure and non-secure resources. An example of a system that supports such a partitioning of the secure and non-secure regions 84, 86 is a system based on the TrustZone (registered trademark) architecture provided by Arm (registered trademark) Limited. The secure region can execute trusted applications 36 at EL0, a trusted operating system 38 at EL1, and optionally a secure partition manager 40 at EL2. EL2 can use stage 2 page tables to support isolation between different trusted operating systems 38 running within the secure region 84 in the same way that the hypervisor 34 can manage isolation between virtual machines or guest operating systems 32 running within the non-secure region 86 when secure partitioning is supported.

[0088] Extending the system to support a secure region 84 has become common in recent years because it enables a single hardware processor to support isolated secure processing and avoids the need to perform the processing on another hardware processor. However, as the popularity of using secure regions has grown, many practical systems with such secure regions now support a relatively sophisticated mixed environment of services provided by a wide variety of software providers within the secure region. For example, the code operating within the secure region 84 can include different software, and its providers can include (among others) a silicon provider that manufactures the integrated circuit, an original equipment manufacturer (OEM) that assembles the integrated circuit provided by the silicon provider into an electronic device such as a mobile phone, an operating system vendor (OSV) that provides the operating system 32 for the device, and / or a cloud platform provider that manages cloud servers that support services for a number of different customers via the cloud.

[0089] However, there is an increasing demand for a secure computing environment in which the provider of user-level code (which is typically expected to run as application 30 within non-secure region 86) can be trusted not to leak information to other parties operating code on the same physical platform. Such a secure computing environment is desirably dynamically allocable during execution and provably guaranteed such that a user can verify whether sufficient security assurances are provided on the physical platform before entrusting the device with the processing of potentially confidential code or data. A user of such software may not wish to trust the provider of a feature-rich operating system 32 or hypervisor 34 that can normally operate in non-secure region 86 (alternatively, even if those providers themselves are trustworthy, the user may wish to protect themselves from the operating system 32 or hypervisor 34 being illegally accessed by an attacker). Also, while secure region 84 is available for such user-provided applications that require secure processing, in practice, this causes problems for both the user providing code that requires a secure computing environment and the provider of the existing code operating within secure region 84. For the provider of the existing code operating within secure region 84, the addition of arbitrary user-provided code within the secure region increases the area of potential attacks against those codes. This can be undesirable and thus strongly discouraged from allowing the user to add code to secure region 84. On the other hand, a user providing code that requires a secure computing environment may be reluctant to entrust all providers of different codes operating within secure region 84 with access to their data or code if the guarantee and proof of the code operating in a particular area are required as a prerequisite for the user-provided code to execute processing.It may be difficult to audit and prove all of the distinct code provided by different software providers operating in the secure region 84, which may limit the opportunity for third parties to provide more secure services.

[0090] Accordingly, as shown in FIG. 2, an additional region 88, referred to as the realm region, is provided. The realm region can be used by such code introduced by the user and provides a secure computing environment that is orthogonal to any secure computing environment associated with components operating in the secure region 24. In the realm region, the software to be executed can include a number of realms, each of which can be isolated from other realms by a realm management module (RMM) 46 operating at exception level EL2. The RMM 46 can control the isolation between the respective realms 42, 44 executing in the realm region 88, for example, by defining access permissions and address mappings within the page table structure, in a similar manner to how the hypervisor 34 manages the separation between various components operating in the non-secure region 86. In this example, the realms include an application-level realm 42 executing at EL0 and a capsule application / operating system realm 44 executing across exception levels EL0 and EL1. It will be understood that it is not essential to support both EL0 and EL0 / EL1 types of realms, and that multiple realms of the same type can be established by the RMM 46.

[0091] The realm area 88 has its own physical address space assigned to it, similar to the secure area 84. However, while the realm area and the secure areas 88 and 84 can each access the non-secure PAS associated with the non-secure area 86, the realm area 88 is orthogonal to the secure area 84 in the sense that the realm area 88 and the secure area 84 cannot access each other's physical address spaces. This means that the code executed in the realm area 88 and the secure area 84 have no dependency on each other. The code within the realm area only needs to trust the hardware, the RMM 46, and the code operating in the root area 82 that manages the switching between areas, which means that the proof and guarantee are more feasible. Proof enables a given software to require verification that the code installed on the device matches certain expected characteristics. This can be done by checking whether the hash of the program code installed on the device matches the expected value signed using a cryptographic protocol by a trusted party. The RMM 46 and the monitor code 29 can be proven, for example, by checking whether the hash of this software matches the expected value signed by a trusted party such as a silicon provider that manufactured the integrated circuit including the processing system 2, or an architecture provider that designed a processor architecture that supports region-based memory access control. Thereby, the user-provided code 42, 44 can verify whether the integrity of the region-based architecture is trustworthy before executing any secure or confidential functions.

[0092] Thus, as indicated by the dotted lines showing the gaps within the non-secure regions where these processes would have been previously executed, the code associated with realms 42, 44 that would have been previously executed in non-secure region 86 can now be moved to realm regions that can have stronger security guarantees since their data and code cannot be accessed by other code operating in non-secure region 86. However, due to the fact that realm region 88 and secure region 84 are orthogonal and thus cannot see each other's physical address spaces, this means that the provider of the code within the realm region need not trust the provider of the code within the secure region and vice versa. The code within the realm region can simply trust the firmware that provides the monitor code 29 of root region 82 and RMM 46, which can be provided by the silicon provider, or the provider of the instruction set architecture supported by the processor. These providers may need to be inherently trusted from the start when the code is executed on their devices, and as a result, no other additional trust relationships with other operating system vendors, OEMs, or cloud hosts are required by the user in order for the user to be provided with a secure computing environment.

[0093] This is useful for a variety of purpose applications and use cases, including, for example, mobile wallets and payment applications, fraud and copyright infringement prevention mechanisms in games, operating system platform security extensions, secure virtual machine hosting, confidential computing, networking, or gateway processing for the Internet of Things. It will be understood that the user may find many other applications for which realm support is useful.

[0094] To support the security guarantees provided to the realm, the processing system can support a proof-reporting function, and during startup or execution, measurements are taken of the firmware image and configuration, e.g., the image and configuration of the monitor code, or the image and configuration of the RMM code. During execution, the content and configuration of the realm are measured, allowing the realm owner to trace relevant proof reports back to known implementations and guarantees and make a trust decision as to whether it is operating on that system.

[0095] As shown in FIG. 2, another root region 82 is provided to manage the switching of regions, and that root region has its own isolated root physical address space. Creating a root region and isolating its resources from the secure region enables a more robust implementation, even in a system that has only non-secure and secure regions 86, 84 and does not have a realm region 88, but can also be used in an implementation that supports a realm region 88. The root region 82 can be implemented using monitor software 29 provided (or guaranteed) by the silicon provider or architecture designer and can be used to provide firmware update management for firmware components provided by other parties such as secure boot functionality, trusted boot measurements, system-on-chip configuration, debug control, and OEMs. The code for the root region can be developed, guaranteed, and deployed by the silicon provider or architecture designer without dependencies on the final device. In contrast, the secure region 84 can be managed by the OEM to implement specific platforms and security services. The management of the non-secure region 86 can be controlled by the operating system 32 to provide operating system services, while the realm region 88 is isolated from the existing secure software environment within the secure region 84 and at the same time enables the development of a new form of a trusted execution environment that can be dedicated to user or third-party applications.

[0096] Figure 3 schematically shows another example of the processing system 2 to support these techniques. Elements that are the same as those in FIG. 1 are denoted by the same reference numerals. FIG. 3 shows more details of the address translation circuit 16 and includes a stage 1 memory management unit 50 and a stage 2 memory management unit 52. The stage 1 MMU 50 may be involved in either the conversion from a virtual address to a physical address (when the conversion is triggered by EL2 or EL3 code) or to an intermediate address (in an operating state where further stage 2 conversion by the stage 2 MMU 52 is required and the conversion is triggered by EL0 or EL1 code). The stage 2 MMU can convert the intermediate address to a physical address. The stage 1 MMU can be based on a page table controlled by the operating system for conversions starting from EL0 or EL1, a page table controlled by the hypervisor for conversions from EL2, or a page table controlled by the monitor code 29 for conversions from EL3. On the other hand, the stage 2 MMU 52 can be based on a page table structure defined by the hypervisor 34, the RMM 46, or the secure partition manager 14 depending on which region is being used. By separating the conversion into two stages in this way, the operating system can manage the address translation for itself and for applications under the assumption that they are the only operating systems running on the system, while the RMM 46, the hypervisor 34, or the SPM 40 can manage the isolation between different operating systems running within the same region.

[0097] As shown in FIG. 3, the address translation process using the address translation circuit 16 can return a security attribute 54 that, in combination with the current exception level 15 and the current region 14 (or security state), enables a section of a particular physical address space (identified by a PAS identifier or “PAS TAG”) to be accessed in response to a given memory access request. The physical address and the PAS identifier can be looked up in a granular protection table 56 that provides the aforementioned granular protection information. In this example, the PAS filter 20 is shown as a detailed (granular) memory protection unit (GMPU) that verifies whether the selected PAS is permitted to access the requested physical address and, if so, permits the transaction to be passed to any cache 24 or interconnect 8 that is part of the system fabric of the memory system.

[0098] The GMPU 20 enables memory to be allocated to different address spaces while at the same time providing a strong hardware-based isolation guarantee, offering not only spatial and temporal flexibility in how physical memory is allocated to these address spaces, but also an efficient sharing scheme. As described above, the execution units within the system are logically divided into virtual execution states (regions or “worlds”) with one execution state (the Root World) located at the highest exception level (EL3) called the “Root World”, and the Root World manages the physical memory allocation to these worlds.

[0099] A single system physical address space is virtualized into multiple “logical” or “architectural” physical address spaces (PASs), and each such PAS is an orthogonal address space with independent coherence attributes. The system physical address is mapped to a single “logical” physical address space by extending it with a PAS tag.

[0100] A given world is permitted access to a subset of the logical physical address space. This is implemented by a hardware filter 20 that can be attached to the output of the memory management unit 16.

[0101] The world defines the security attributes (PAS tags) of access using the fields of the translation table descriptors of the page table used for address translation. The hardware filter 20 has access to a table (granule protection table 56, or GPT) that defines granule protection information (GPI) for each page within the system physical address space, and the granule protection information indicates the PAS TAG it is associated with and (optionally) other granule protection attributes.

[0102] The hardware filter 20 checks the world ID and security attributes against the GPI of the granule to determine whether access can be permitted, thus forming a detailed memory protection unit (GMPU).

[0103] The GPT 56 can exist, for example, in on-chip SRAM or off-chip DRAM. When stored off-chip, the GPT 56 can be integrity protected by an on-chip memory protection engine that can use mechanisms for encryption, integrity, and freshness to maintain the security of the GPT 56.

[0104] Positioning the GMPU 20 on the requester side of the system (e.g., on the MMU output) rather than on the completer side allows for page-level access permission assignment while permitting the interconnect 8 to continue hashing / striping pages across multiple DRAM ports.

[0105] The transaction propagates through the entire system fabric 24, 8 until it reaches a location defined as the physical alias point 60, at which point the PAS TAG remains tagged. This allows the filter to be placed on the master side without weakening security guarantees compared to slave - side filtering. As the transaction propagates through the system, the PAS TAG can be used as a detailed security mechanism for address isolation. For example, the cache can add the PAS TAG to the address tags in the cache to prevent accesses made with an incorrect PAS TAG for the same PA from hitting the cache, thereby improving side - channel resistance. The PAS TAG can also be used as a context selector for a protection engine attached to a memory controller that encrypts data before writing it to external DRAM.

[0106] The physical alias point (PoPA) is a location within the system where the PAS TAG is stripped and the address returns from the logical physical address to the system physical address. The PoPA can be located below the cache on the completer side of the system where accesses to physical DRAM are made (using the encrypted context resolved via the PAS TAG). Alternatively, at the expense of weakening security, it may be located above the cache to simplify the system implementation.

[0107] At any point in time, the world can request that the page be transitioned from one PAS to another. The request is made at EL3 for monitor code 29 and examines the current state of the GPI. EL3 may allow only a specific set of transitions (e.g., from a realm PAS to a secure PAS but not from a non-secure PAS to a secure PAS) to occur. To provide a clean transition, a new instruction "delete data and invalidate up to the physical aliasing point" is supported by the system and EL3 can submit this before transitioning the page to the new PAS. This ensures that any residual state associated with the previous PAS is flushed from any caches upstream of PoPA60 (closer to the requester side).

[0108] Another property that can be achieved by attaching the GMPU20 on the master side is the efficient sharing of memory between worlds. It may be desirable to allow a subset of N worlds to have shared access to a physical granule while preventing other worlds from accessing it. This can be achieved by adding "limited sharing" semantics to the granule protection information while forcing it to use a specific PAS TAG. As an example, the GPI can indicate that a physical granule can be accessed only by the "realm world" 88 and the "secure world" 84 while the PAS TAG of the secure PAS 84 is tagged to the physical granule.

[0109] The above example of properties results in a rapid change in the visibility properties of a specific physical granule. Consider the case where each world has a private PAS that is accessible only to that world. For a specific granule, the world can request that it be made visible to the non-secure world at any point in time without changing the association of the PAS by changing its GPI from "exclusive" to "limited sharing with the non-secure world". In this way, the visibility of that granule can be increased without requiring costly cache maintenance or data copy operations.

[0110] Figure 4 shows the concept of aliasing each physical address space on the physical memory provided in the hardware. As described above, each of regions 82, 84, 86, 88 has its own respective physical address space 61.

[0111] At the time when the physical address is generated by the address translation circuit 16, the physical address has a value within a specific numerical range 62 supported by the system, which is the same regardless of which physical address space is selected. However, in addition to generating the physical address, the address translation circuit 16 can also select a specific physical address space (PAS) based on the current region 14 and / or the information in the page table entry used to derive the physical address. Alternatively, instead of the address translation circuit 16 that performs the selection of the PAS, an address translation circuit (e.g., MMU) can output the physical address and the information derived from the page table entry (PTE) used for the selection of the PAS, and then this information can be used by the PAS filter or GMPU 20 to select the PAS.

[0112] The selection of the PAS for a given memory access request can be limited according to the rules defined in the following table depending on the current region in which the processing circuit 10 issues the memory access request.

Table 1

[0113] For a region where there are multiple physical address spaces available for selection, select from among the available PAS options using the information from the accessed page table entry used to provide the physical address.

[0114] Therefore, when the PAS filter 20 outputs a memory access request to the system fabric 24, 8 (assuming it has passed any filtering checks), the memory access request is associated with a physical address (PA) and a selected physical address space (PAS).

[0115] From the perspective of memory system components (caches, interconnects, snooping filters, etc.) that operate before the physical alias (PoPA) point 60, each physical address space 61 is considered a completely different address range corresponding to different system locations within the memory. This means that from the perspective of the pre-PoPA memory system components, the address range identified by a memory access request is actually four times the size of the range 62 that can be output in address translation. This is because in effect, the PAS identifier is treated as an additional address bit alongside the physical address itself, so that depending on the selected PAS, the same physical address PAx can be mapped to several aliased physical addresses 63 within different physical address spaces 61. These aliased physical addresses 63 all actually correspond to the same memory system location implemented in physical hardware, but the pre-PoPA memory system components treat the aliased addresses 63 as different addresses. Therefore, if there is a pre-PoPA cache or snooping filter that allocates an entry to such an address, the aliased address 63 will be mapped to different entries with different cache hit / miss decisions and different coherence management. This reduces the likelihood or effectiveness of an attacker using a cache or coherence side channel as a mechanism to probe the operation of other areas.

[0116] The system may include two or more PoPAs 60 (e.g., as shown in FIG. 14 discussed below). In each PoPA 60, the aliased physical address is folded into a single non-aliased address 65 within the system physical address space 64. The non-aliased address 65 is provided downstream of any PoPA post-component, such that the system physical address space 64, which actually identifies the location of the memory system, is again the same size as the range of physical addresses that can be output in the address translation performed on the requester side. For example, in the PoPA 60, the PAS identifier may be stripped from the address, and for downstream components, the address can be identified simply using the physical address value without specifying the PAS. Alternatively, in some cases where some complier-side filtering of memory access requests is desired, the PAS identifier may still be provided downstream of the PoPA 60 but may not need to be interpreted as part of the address. As a result, the same physical address appearing in different physical address spaces 60 will be interpreted to refer to the same location in the memory system downstream of the PoPA. However, the supplied PAS identifier can still be used to perform complier-side security checks.

[0117] Figure 5 shows how the system physical address space 64 can be divided into chunks allocated for access within the physical address space 61 on a particular architecture using the granularity protection table 56. The granularity protection table (GPT) 56 defines which portions of the system physical address space 65 are accessible from the physical address space 61 on each architecture. For example, the GPT 56 may include several entries each corresponding to a granularity of physical addresses of a particular size (e.g., 4K pages) and may define the allocated PAS for that granularity, which can be selected from among non-secure, secure, realm, and root regions. By design, if a particular granularity or set of granularities is allocated to a PAS associated with one of the regions, it can only be accessed within the PAS associated with that region and cannot be accessed within the PAS of other regions. However, note that the root region 82 can still ensure that the virtual address associated with the page mapped to that area of physically addressed memory in its page table is translated to a physical address within the secure PAS instead of the root PAS, by specifying PAS selection information for that granularity of physical address, even though the granularity allocated to the (e.g.,) secure PAS cannot be accessed from within the root PAS. Thus, sharing of data between regions (to the extent permitted by the access rules defined in the aforementioned table) can be controlled at the time of selecting the PAS for a given memory access request.

[0118] However, in some implementations, in addition to enabling access to a granularity of physical addresses within the allocated PAS defined by the GPT, the GPT can mark a certain area of the address space as shared with another address space using other GPT attributes (e.g., an address space associated with a lower or orthogonal privilege region where it is not normally permitted to select the PAS allocated for access requests to that region).

[0119] This can facilitate temporarily sharing data without the need to change the PAS assigned to a given granule. For example, in FIG. 5, the region 70 of the realm PAS is defined within the GPT to be assigned to the realm area, and the non-secure area 86 is normally inaccessible from the non-secure area 86 because it cannot select the realm PAS for its access request. Since the non-secure area 26 cannot access the realm PAS, non-secure code could not normally view the data in region 70. However, if the realm desires to temporarily share some of the data in the allocated area of memory with the non-secure area, it can be required that the monitor code 29 operating in the root area 82 update the GPT 56 to indicate that region 70 is shared with the non-secure area 86, whereby region 70 can also be made accessible from the non-secure PAS shown on the left side of FIG. 5 without the need to change which area is assigned to region 70. When the realm area indicates that its address space region is to be shared with the non-secure area, a memory access request issued from the non-secure area and targeting that region can initially specify the non-secure PAS, but the PAS filter 20 can remap it so that the PAS identifier of the request instead specifies the realm PAS. Thereby, downstream memory system components handle the request as if it had been issued from the realm area from the beginning. This sharing can improve performance because the operations for allocating different areas to specific memory regions can be more performance-intensive, involving a higher degree of cache / TLB invalidation and / or zeroing of data in memory or copying of data between memory regions, which may not be justified if the sharing is expected to be only temporary.

[0120] FIG. 6 is a flowchart showing how to determine the current operating region, which can be executed by the processing circuit 10 or by the address translation circuit 16 or the PAS filter 20. In step 100, it is determined whether the current exception level 15 is EL3. If so, then in step 102, it is determined that the current region is the root region 82. If the current exception level is not EL3, then in step 104, it is determined that the current region is one of the non-secure, secure, and realm regions 86, 84, 88 as indicated by at least two region indication bits 14 in the EL3 control register of the processor (since the root region is indicated by the current exception level being EL3, it is not essential to have the code of the region indication bit 14 corresponding to the root region, and thus the code of at least one region indication bit can be reserved for other purposes). The EL3 control register is writable when operating at EL3 and cannot be written from other exception levels EL2 - EL0.

[0121] FIG. 7 shows an example of a page table entry (PTE) in a page table structure that can be used by the address translation circuit 16 for mapping from a virtual address to a physical address, from a virtual address to an intermediate address, or from an intermediate address to a physical address (depending on whether the translation is actually performed in an operating state where stage 2 translation is required, and if stage 2 translation is required, which of stage 1 and stage 2 the translation is). In general, a given page table structure may be defined as a multi-level table structure where the first level of the page table is implemented as a page table tree identified based on a base address stored in the processor's translation table base address register, and the index for selecting a particular level 1 page table entry within the page table is derived from a subset of the bits of the input address for which the translation look-up is to be performed (the input address can be a virtual address for stage 1 translation of an intermediate address for stage 2 translation). The level 1 page table entry can be a "table descriptor" 110 that provides a pointer 112 to the next level of the page table, from which further page table entries can be selected based on a further subset of the bits of the input address. Finally, after one or more look-ups to successive levels of the page table, block or page descriptor PTEs 114, 116, 118 that provide an output address 120 corresponding to the input address can be identified. The output address can be an intermediate address (for stage 1 translation in an operating state where further stage 2 translation is also performed), or a physical address (for stage 2 translation, or for stage 1 translation when stage 2 is not required).

[0122] To support the separate physical address spaces described above, the page table entry format can specify some additional state for use in physical address space selection, in addition to the next level page table pointer 112 or output address 120, and any attributes 122 for controlling access to the corresponding block of memory.

[0123] For the table descriptor 110, the PTEs used by any region other than the non-secure region 86 include a non-secure table indicator 124 that indicates which level of page table should be accessed from either the non-secure physical address space or the physical address space of the current region. This helps to facilitate more efficient management of the page table. In many cases, the page table structure used by the root, realm, or secure region 24 may only need to define special page table entries for a portion of the virtual address space. And for other portions, the same page table entries used by the non-secure region 26 can be used. Thus, by providing the non-secure table indicator 124, a higher-level realm / secure dedicated table descriptor can be provided in the page table structure, while at some point in the page table tree, the root realm or secure region can switch to using page table entries from the non-secure region for portions of the address space where higher security is not required. Other page table descriptors in other parts of the page table tree can still be fetched from the associated physical address space related to the root, realm, or secure region.

[0124] On the one hand, the block / page descriptors 114, 116, 118 may include physical address space selection information 126 depending on which region they are associated with. The non-secure block / page descriptor 118 used within the non-secure region 86 does not include any PAS selection information because the non-secure region can only access the non-secure PAS. However, for other regions, the block / page descriptors 114, 116 include PAS selection information 126 that is used to select the PAS for translating the input address. For the root region 22, the EL3 page table entry may have PAS selection information 126 that includes at least 2 bits to indicate the PAS associated with any one of the four regions 82, 84, 86, 88 as the selected PAS to which the corresponding physical address is translated. In contrast, for the RELM and secure regions, the corresponding block / page descriptor 116 only needs to include 1-bit PAS selection information 126, whereby for the RELM region, a selection is made between the RELM and the non-secure PAS, and for the secure region, a selection is made between the secure PAS and the non-secure PAS. To improve the efficiency of circuit implementation and avoid increasing the size of the page table entry, for the RELM and secure regions, the block / page descriptor 116 can encode the PAS selection information 126 at the same position within the PTE regardless of whether the current region is RELM or secure, so that the PAS selection bits 126 can be shared.

[0125] Accordingly, FIG. 8 is a flowchart showing a method of selecting a PAS based on information 124, 126 from the block / page PTE that is used to generate a physical address for a current region and a given memory access request. The PAS selection can be performed by the address translation circuit 16, or in the case where the address translation circuit sends the PAS selection information 126 to the PAS filter 20, it can be performed by a combination of the address translation circuit 16 and the PAS filter 20.

[0126] In step 130 of FIG. 8, the processing circuit 10 issues a memory access request by designating a given virtual address (VA) as the target VA. In step 132, the address translation circuit 16 looks up any page table entry (or cache information derived from such a page table entry) in its TLB 18. If any necessary page table information is not available, the address translation circuit 16 starts a page table walk to memory to fetch the necessary PTE (potentially stepping through each level of the page table structure and / or multiple stages of address translation to obtain the mapping from VA to intermediate physical address (IPA) and from IPA to PA, which requires a series of memory accesses). Any memory access request issued by the address translation circuit 16 in the page table walk operation itself can be subject to address translation and PAS filtering. Thus, note that the request received in step 130 can be a memory access request issued to request a page table entry from memory. When relevant page table information is identified, the virtual address is translated to a physical address (possibly in two stages via the IPA). In step 134, the address translation circuit 16 or the PAS filter 20 determines which region is the current region using the technique shown in FIG. 6.

[0127] If the current region is a non-secure region, in step 136, the output PAS selected for this memory access request is a non-secure PAS.

[0128] If the current region is a secure region, in step 138, the output PAS is selected based on the PAS selection information 126 included in the block / page descriptor PTE that provided the physical address, and the output PAS is selected as either a secure PAS or a non-secure PAS.

[0129] If the current area is the realm area, in step 140, the output PAS is selected based on the PAS selection information 126 included in the block / page descriptor PTE from which the physical address has been derived. In this case, the output PAS is selected as either the realm PAS or the non-secure PAS.

[0130] If it is determined in step 134 that the current area is the root area, in step 142, the output PAS is selected based on the PAS selection information 126 within the root block / page descriptor PTE 114 from which the physical address has been derived. In this case, the output PAS is selected as any one of the physical address spaces associated with the root, realm, secure, and non-secure areas.

[0131] Figure 9 shows an example of a GPT56 entry for a given granularity of a physical address. The GPT entry 150 includes an assigned PAS identifier 152 that identifies the PAS assigned to the granularity of the physical address, and optionally may include additional attributes 154 including, for example, the aforementioned shared attribute information 156 that is visualized with one or more other PASs other than the PAS to which the granularity of the physical address is assigned. The setting of the shared attribute information 156 may be performed by the root region in response to a request from code executed within the region associated with the assigned PAS. Also, the attribute may include a pass-through indicator field 158 indicating whether a GPT check (for determining whether the PAS selected for a memory access request can access the granularity of the physical address) should be performed on the requester side by the PAS filter 20 or by a completer-side filtering circuit on the completer device side of the interconnect, as further described below. If the pass-through indicator 158 has a first value, the requester-side filtering check may be required by the PAS filter 20 on the requester side, and if these fail, the memory access request may be blocked and a fault signaled. However, if the pass-through indicator 158 has a second value, the requester-side filtering check based on GPT56 may not be necessary for a memory access request specifying a physical address within the granularity of the physical address corresponding to that GPT entry 150, in which case the memory access request may be passed through to the cache 24 or the interconnect 8 regardless of whether the selected PAS is checked to be one of the permitted PASs that are permitted to access that granularity of the physical address, and instead any such PAS filtering check may then be performed later on the completer side.

[0132] FIG. 10 is a flowchart showing a requester-side PAS filter check executed by the PAS filter 20 on the requester side of the interconnect 8. In step 170, the PAS filter 20 receives a memory access request associated with a physical address and an output PAS that can be selected as shown in FIG. 8 described above.

[0133] In step 172, the PAS filter 20 issues a request to the memory to fetch the required GPT entry from the granularity protection information cache 22 if available, or from the table structure stored in the memory, thereby obtaining the GPT entry corresponding to the specified PA. When the required GPT entry is obtained, in step 174, the PAS filter determines whether the output PAS selected for the memory access request is the same as the assigned PAS 152 defined in the GPT entry obtained in step 172. If so, in step 176, the memory access request (specifying the PA and the output PAS) can be passed to the cache 24 or the interconnect 8.

[0134] If the output PAS is not the assigned PAS, in step 178, the PAS filter determines whether the output PAS is indicated in the shared attribute information 156 from the obtained GPT entry as an allowed PAS that is accessible at the granularity of the address corresponding to the specified PA as an allowed PAS. If so, again in step 176, the memory access request is permitted to be passed to the cache 24 or the interconnect 8. The shared attribute information can be encoded as a unique bit (or set of bits) within the GPT entry 150, or as one or more encodings of a field of the GPT entry 150 where other encodings of the same field may indicate other information. In step 178, if it is determined that the shared attribute permits an output PAS other than the assigned PAS to access the PA, in step 176, the PAS specified in the memory access request passed to the cache 24 or the interconnect 8 is the assigned PAS rather than the output PAS. The PAS filter 20 transforms the PAS specified by the memory access request to match the assigned PAS, such that the downstream memory system component treats it the same as a request issued specifying the assigned PAS.

[0135] If the output PAS is not indicated in the shared attribute information 156 as being permitted to access a particular physical address (alternatively, in an implementation where the shared attribute information 156 is not supported, step 178 is skipped), at step 180, it is determined whether the pass - through indicator 158 in the GPT entry obtained for the target physical address, regardless of the check performed by the requester - side PAS filter 20, identifies that the memory access request can be passed through to the cache 24 or the interconnect 8. If the pass - through indicator is specified, at step 176, the memory access request can be advanced again (by designating the output PAS as the PAS associated with the memory access request). Alternatively, if any of the checks in steps 174, 178, and 180 do not identify that the memory access request is permitted, at step 182, the memory access request is blocked. Thus, the memory access request is not passed to the cache 24 or the interconnect 8, an error can be signaled, and exception handling to address the error can be triggered.

[0136] Steps 174, 178, 180 are shown sequentially in FIG. 10, but these steps can also be performed in parallel or in a different order if desired. Also, steps 178 and 180 are not essential, and it should be understood that some implementations may not support the use of the shared attribute information 156 and / or the pass - through indicator 158.

[0137] Figure 11 summarizes the operation of the address translation circuit 16 and the PAS filter. PAS filtering 20 can be regarded as an additional stage 3 check that is executed after the address translation of stage 1 (and optionally stage 2) executed by the address translation circuit. Also, the EL3 translation is based on a page table entry that provides selection information (labeled NS, NSE in the example of Figure 11) based on a 2-bit address, while the single-bit selection information "NS" is also noted to be used to select PAS in other states. The security state shown in Figure 11 as input to the granularity protection check refers to the region ID that identifies the current region of the processing element 4.

[0138] Figure 12 is a flowchart showing the processing of a stage 3 lookup cache invalidation instruction that can be used by the monitor code 29 operating in the root region 82 to trigger the invalidation of any lookup cache entry that depends on the GPT entry associated with a particular physical address. This can be useful when the root region changes the allocation of which physical address space a given system PA is assigned to, so that any lookup cache 22 does not hold stale information.

[0139] Therefore, at step 200, the processing circuit 10 within a given processing element (requester device) 4 can execute a stage 3 lookup cache invalidation instruction. The instruction specifies a physical address.

[0140] In step 202, in response to the stage 3 lookup cache invalidation instruction, the processing circuit 10 checks whether the current exception level is EL3, and if not, in step 204, it can reject the instruction and / or signal an exception (such as an undefined instruction exception). This limits the execution of the stage 3 lookup cache invalidation instruction to the monitor code 29 associated with the root region, preventing a malicious party from forcing the invalidation of granularity protection information from the lookup cache 22 by being triggered at other exception levels, thereby preventing a performance degradation.

[0141] If the current exception level is EL3, then in step 206, in response to the instruction executed in step 200, the processing element issues at least one lookup cache invalidation command that can include information according to the granularity protection table entry associated with the physical address identified by the instruction to any lookup cache 18, 22 that can contain information. These caches can include not only the granularity protection information cache 22 as shown in FIG. 1, but in some implementations, a combined TLB / granularity protection cache that combines page table structure and information from the GPT into a single entry. For such a combined TLB / GPT cache, the combined cache may require the ability to be looked up by both virtual and physical addresses.

[0142] In step 208, in response to the issued command, the GPT cache 22 or the combined TLB / GPT cache invalidates any entry that depends on the granularity protection information from the GPT entry associated with the granularity of the physical address corresponding to the physical address specified by the lookup cache invalidation command.

[0143] Since the cache in the memory system can tag an entry with the identification of the PAS associated with the cache located before the PoPA, even if the root region code 29 changes which PAS is associated with a particular granule of the physical address by updating the GPT56, there may still be data cached in the PoPA-previous cache that is tagged with the wrong PAS for that granule of the physical address. To prevent subsequent accesses issued after the GPT update from hitting cache entries that the regions issuing these requests should no longer be able to access, it may be useful to provide an instruction that ensures that any cache entry associated with a particular physical address is invalidated in any cache before PoPA60. This instruction may be a different instruction from other types of cache invalidation instructions that can act on caches within different subsets of the memory system, such as caches before the coherence point or caches local to a particular processing element. Thus, the processing circuit can support a cache invalidation instruction that is observed to be invalidated by the caches in the system and whose scope is distinguished with respect to the scope corresponding to the upstream portion of the PoPA of the memory system up to the PoPA. That is, for this instruction, the PoPA is the boundary of the scope within which cache invalidation by the cache should be respected.

[0144] FIG. 13 is a flowchart showing the processing of cache invalidation instructions up to PoPA. At step 220, cache invalidation instructions up to PoPA are executed by the processing circuit 10 of a given requester device 4. The instruction specifies a virtual address, and at step 222 this virtual address is mapped to a physical address. However, in some cases, similar to the instructions shown in FIG. 12, the execution of instructions on caches invalidated up to PoPA may be restricted to execution within EL3. At step 224, the processing element that executed the instruction issues an invalidation command specifying the physical address, and these commands are sent to any pre-PoPA cache 24 in the system. This may include caches upstream of PoPA not only within the requester device 4 but also within the interconnect 8 or other memory system components where another range of memory locations addressed by a different physical address space 61 or another requester device 4 is treated as a separate range.

[0145] The cache invalidation command up to PoPA can, in some cases, be a "delete and invalidate" form of command that not only requires the data associated with the specified PA to be invalidated from the pre-PoPA cache, but also requires that the data be deleted by writing any dirty data back to a location beyond PoPA60 before invalidation. Thus, in step 226, if the command is a delete and invalidate form of command, the pre-PoPA cache that receives the command triggers the write-back of dirty data from any entry associated with the specified physical address. This data can be written to a cache beyond PoPA or to main memory. If the command is not a delete and invalidate form of command, or if the delete and invalidate form of command is not supported, step 226 can be omitted and the method can proceed directly to step 228, which would be executed if step 226 were executed. In step 228, the pre-PoPA cache that received the command in step 224 invalidates the entry associated with the specified physical address. The cache entries associated with the specified physical address are invalidated regardless of which PAS tags are associated with those entries.

[0146] Thus, the cache invalidation command up to PoPA can be used to ensure that, after an update to GPT, the cache does not continue to tag incorrect PAS identifiers to the entries associated with a given physical address.

[0147] Figure 14 shows a more detailed example of a data processing system that can implement some of the above techniques. Elements that are the same as in previous embodiments are indicated by the same reference numerals. In the example of Figure 14, in addition to the processing circuit 10, the address translation circuit 16, the TLB 18, and the PAS filter 20, the cache 24 is shown in more detail in that it includes a level 1 instruction cache, a level 1 data cache, a level 2 cache, and optionally a shared level 3 cache 24 shared among processing elements, showing the processing element 4 in more detail. The interrupt controller 300 can control the processing of interrupts by each processing element.

[0148] As shown in Figure 14, the processing element 4 that can execute program instructions to trigger access to memory is not the only type of requester device that can include the requester-side PAS filter 20. In other examples, (such as on-chip devices 312 like a network interface controller or a display controller, or off-chip devices 314 that can communicate with the system via a bus, etc., requester devices 312, 314 that do not support their own address translation functions), the system MMU 310 provided to provide an address translation function (for these) can include a PAS filter 20 to perform a requester-side check of GPT entries, similar to that for the PAS filter 20 within the processing element 4. Other requester devices can include a debug access port 316 and a control processor 318, which can also have an associated PAS filter 20, thereby checking whether memory accesses issued by the requester devices 316, 318 to a specific physical address space are permitted under the PAS assignment defined by GPT56.

[0149] The interconnect 8 is shown in more detail in FIG. 14 as a coherent interconnect 8 and includes, in addition to the routing fabric 320, a snooping filter 322 for managing coherence between the caches 24 within each processing element and one or more system caches 324 that can execute a cache of shared data shared between requesting devices. The snooping filter 322 and the system cache 324 can be located upstream of the PoPA 60 and thus can tag their entries using the PAS identifiers selected by the MMUs 16, 310 for a particular master. Requesting devices 316, 318 not associated with an MMU can be assumed to default to always issuing requests for a particular region, such as a non-secure region (or, if they are trusted, the root region).

[0150] FIG. 14 shows, as another example of a pre-PoPA component that handles physical address aliasing within each PAS as if they were referring to different address locations, a memory protection engine (MPE) 330 provided between the interconnect 8 and a given memory controller 6 for controlling access to off-chip memory 340. The MPE 330 can be responsible for encrypting data written to the off-chip memory 340 to maintain confidentiality and decrypting the data when reading. The MPE can also generate integrity metadata when writing data to memory and verify whether the data has changed by using the metadata when the data is read from the off-chip memory, thereby preventing tampering of the data stored in the off-chip memory. When encrypting data or generating a hash of memory integrity, different keys can be used depending on which physical address space is being accessed, even if accessing an aliased physical address that actually corresponds to the same location within the off-chip memory 340. This improves security by further isolating data associated with different operating regions.

[0151] In this example, since PoPA60 is between the memory protection engine 330 and the memory controller 6, when a request reaches the memory controller 6, the physical address is no longer treated as being mapped to different physical locations within the memory 340 depending on the physical address space in which they are accessed.

[0152] FIG. 14 shows another example of the completer device 6, which can be a peripheral bus or a non - coherent interconnect used to communicate with regions of the peripheral device 350 or the on - chip memory 360 (e.g., implemented as a static random access memory (SRAM)). Also, the peripheral bus or non - coherent interconnect 6 can be used to communicate with secure elements 370 such as an encryption unit that executes encryption processing, a random number generator 372, or certain fuses 374 that store information statically embedded in hardware. Also, various power / reset / debug controllers 380 can be accessible via the peripheral bus or non - coherent interconnect 6.

[0153] For the on-chip SRAM 360, it may be useful to provide a PAS filter 400 on the slave side (the completer side). The PAS filter 400 can perform completer-side filtering of memory accesses based on completer-side protection information that defines which physical address spaces are accessible to a given block of physical addresses. This completer-side information can be defined more coarsely than the GPT used by the requester-side PAS filter 20. For example, the slave-side information can simply indicate for other regions that the entire SRAM unit 361 can be dedicated for use by the realm area, another SRAM unit 362 can be dedicated for use by the root area, and so on. Thus, relatively coarsely defined blocks of physical addresses can be addressed to different SRAM units. This completer-side protection information can be statically defined by loading the bootloader code in the information for the completer-side PAS filter at startup that cannot be changed during execution. Therefore, it is not as flexible as the GPT used by the requester-side PAS filter 20. However, for use cases where the division into specific regions accessible to each region of the physical address is known at startup and does not require fine-grained division without being changed, it may be more efficient to use the slave-side PAS filter 400 instead of the requester-side PAS filter 20. This is because it can eliminate the power and performance cost on the requester side of obtaining GPT entries and comparing the assigned PAS and shared attribute information with the information for the current memory access request. Also, if the pass-through indicator 158 can be shown in the top-level GPT entry (or other table descriptor entry at a level other than the final level) in the multi-level GPT structure, access to a further level of the GPT structure (which can be performed to find more fine-grained information regarding the assigned PAS for requests that undergo requester-side checks) can be avoided for requests targeting one of the regions of physical addresses mapped to the on-chip memory 360 monitored by the completer-side PAS filter 400.

[0154] Therefore, supporting a hybrid approach that enables both requester - side and completer - side checks of the protection information can be useful for performance and power efficiency. System designers can define which approach to take for a particular memory region.

[0155] FIG. 15 shows a simulator implementation that can be used. The above-described embodiments implement the present invention in terms of an apparatus and method for operating specific processing hardware that supports the technology, but it is also possible to provide an instruction execution environment according to the embodiments described herein that is implemented using a computer program. Such a computer program is often referred to as a simulator as long as the computer program provides a software-based implementation of a hardware architecture. Various simulator computer programs include binary translators including emulators, virtual machines, models, and dynamic binary translators. Typically, an implementation form of a simulator can optionally execute a host operating system 420 that supports a simulator program 410 and can be executed by a host processor 430. In some configurations, there may be multiple layers of simulation between the hardware and the provided instruction execution environment, and / or there may be multiple different instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations that execute at a reasonable speed, but such an approach can be justified in certain situations, for example, when it is desired to execute native code on another processor for reasons of compatibility or reuse. For example, the simulator implementation may provide an instruction execution environment having additional functions not supported by the host processor hardware, or may typically provide an instruction execution environment associated with a different hardware architecture. An overview of simulation is described in "Some Efficient Architecture Simulation Techniques", Robert Bedichek, 1990 Winter USENIX Conference, pages 53-63.

[0156] So far, embodiments have been described with reference to specific hardware configurations or functions. However, in simulated embodiments, equivalent functions can be provided by appropriate software configurations or functions. For example, a specific circuit may be implemented as computer program logic in a simulated embodiment. Similarly, memory hardware such as registers or caches may be implemented as software data structures in simulated embodiments. In configurations where one or more of the hardware elements referred to in the foregoing embodiments are present in host hardware (e.g., host processor 430), some simulated embodiments may use the host hardware where appropriate.

[0157] The simulator program 410 may be stored in a computer-readable storage medium (which may be a non-transitory medium) and provides a program interface (instruction execution environment) for the target code 400 (which may include applications, operating systems, and hypervisors), and is the same as the interface of the hardware architecture modeled by the simulator program 410. Accordingly, the program instructions of the target code 400 may be executed from within the instruction execution environment using the simulator program 410. Thus, a host computer 430 that does not actually have the hardware features of the aforementioned apparatus 2 can emulate these features. This can be useful, for example, to enable testing of target code 400 being developed for a new version of a processor architecture before a hardware device that actually supports that architecture becomes available. This is because the target code can be tested by being executed within a simulator that runs on a host device that does not support its architecture.

[0158] The simulator code emulates the behavior of the processing circuit 10, for example, including instruction decoding program logic that decodes the instructions of the target code 400, maps the instructions to a sequence of corresponding instructions within the native instruction set supported by the host hardware 430, and executes a function corresponding to the decoded instructions, including processing program logic 412. The processing program logic 412 also simulates the processing of code at different exception levels and regions as described above. The register emulation program logic 413 maintains a data structure within the host address space of the host processor that emulates the architectural register state defined according to the target instruction set architecture associated with the target code 400. Thus, instead of such an architectural state being stored in the hardware register 12 as in the embodiment of FIG. 1, it is stored in the memory of the host processor 430, and the register emulation program logic 413 maps the register references of the instructions of the target code 400 to corresponding addresses for obtaining the emulated architectural state data from the host memory. This architectural state may include the aforementioned current region indication 14 and current exception level indication 15.

[0159] The simulation code includes address translation program logic 414 and filtering program logic 416 that respectively emulate the functions of the address translation circuit 16 and the PAS filter 20 with reference to the same page table structure and GPT56 as described above. Thus, the address translation program logic 414 converts the virtual address specified by the target code 400 into a simulated physical address within one of the PASs (referring to the physical location in memory from the perspective of the target code), but in fact these simulated physical addresses are mapped onto the (virtual) address space of the host processor by the address space mapping program logic 415. The filtering program logic 416 performs a lookup of the granularity protection information to determine whether to allow the memory access triggered by the target code to proceed, similar to the PAS filter described above.

[0160] Further embodiments are described in the following clauses.

[0161] (1) An apparatus comprising a processing circuit for executing processing in at least one of at least three regions, and an address translation circuit for converting a virtual address of a memory access executed from the current region into a physical address within one of a plurality of physical address spaces selected based at least on the current region, wherein at least three regions include a root region for managing switching between a plurality of other regions among the at least three regions, and the plurality of physical address spaces include a root physical address space associated with the root region, which is different from the physical address spaces associated with the plurality of other regions.

[0162] (2) The apparatus according to clause (1), wherein the root physical address space is exclusively accessible from the root region.

[0163] (3) The apparatus according to any one of clauses (1) and (2), wherein all of the plurality of physical address spaces are accessible from the root region.

[0164] (4) The plurality of other regions includes at least a secure region associated with a secure physical address space and a less secure region associated with a less secure physical address space, the less secure physical address space being accessible from the less secure region, the secure region, and the root region, the secure physical address space being accessible from the secure region and the root region and inaccessible from the less secure region, the apparatus according to any one of clauses (1) to (3).

[0165] (5) The plurality of other regions also includes a further region associated with a further physical address space, the less secure physical address space being accessible from the further region, the further physical address space being accessible from the further region and the root region and inaccessible from the less secure region, the apparatus according to clause (4).

[0166] (6) The further physical address space is inaccessible from the secure region, and the secure physical address space is inaccessible from the further region, the apparatus according to clause (5).

[0167] (7) The less secure physical address space is accessible from all of at least three regions, the apparatus according to any one of clauses (4) to (6).

[0168] (8) The address translation circuit is configured to convert a virtual address to a physical address based on at least one page table entry, and when at least the current region is one of a subset of at least three regions, the address translation circuit is configured to select one of a plurality of physical address spaces based on the current region and physical address space selection information specified by at least one page table entry, the apparatus according to any one of clauses (1) to (7).

[0169] (9) When the current region is the root region, the address translation circuit is configured to translate a virtual address to a physical address based on a page table entry of the root region that includes physical address space selection information for selecting between at least three physical address spaces accessible from the root region, the physical address space selection information being at least 2-bit physical address space selection information, the apparatus according to clause (8).

[0170] (10) The address translation circuit is configured to translate a virtual address to a physical address based on at least one page table entry. When the current region is a secure region, the address translation circuit is configured to select, based on a physical address space selection indicator specified within at least one page table entry, whether one of a plurality of physical address spaces is a secure physical address space or a less secure physical address space. When the current region is a further region, the address translation circuit is configured to select, based on a physical address space selection indicator specified within at least one page table entry, whether one of a plurality of physical address spaces is a further physical address space or a less secure physical address space, the apparatus according to clause (6).

[0171] (11) The apparatus according to clause (10), wherein the physical address space selection indicator is encoded at the same position within at least one page table entry regardless of whether the current region is a secure region or a further region.

[0172] (12) At least one pre-PoPA memory system component provided upstream of a physical aliasing point (PoPA) for treating aliased physical addresses from different physical address spaces corresponding to the same memory system resource as if the aliased physical addresses corresponded to different memory system resources, the apparatus according to any one of clauses (1) to (11).

[0173] (13) The apparatus according to clause (12), wherein the aliased physical address includes a physical address represented using the same physical address value within different physical address spaces.

[0174] (14) The apparatus according to any one of clauses (12) and (13), comprising a PoPA memory system component configured to de-alias a plurality of aliased physical addresses in order to obtain an un-aliased physical address to be provided to at least one downstream memory system component.

[0175] (15) The apparatus according to any one of clauses (12) to (14), wherein at least one PoPA pre-memory system component includes at least one PoPA pre-cache, and in response to a cache invalidation command up to a PoPA specifying a target address, the processing circuit issues at least one invalidation command, thereby requiring the at least one PoPA pre-cache to invalidate one or more entries associated with a target physical address value corresponding to the target address.

[0176] (16) The apparatus according to clause (15), wherein when the processing circuit issues at least one invalidation command, at least one PoPA post-cache located downstream of the PoPA is enabled to hold one or more entries associated with the target physical address value.

[0177] (17) The apparatus according to any one of clauses (15) and (16), wherein in response to at least one invalidation command, at least one PoPA pre-cache is configured to invalidate one or more entries associated with the target physical address value, regardless of which of the plurality of physical address spaces is associated with the one or more entries.

[0178] (18) A memory encryption circuit that, in response to a memory access request specifying a selected physical address space and a target physical address within the selected physical address space, encrypts or decrypts data associated with a protected area based on one of a plurality of encryption keys selected according to the selected physical address space when the target physical address is within a protected address area, the apparatus according to any one of clauses (1) to (17) comprising the memory encryption circuit.

[0179] (19) A data processing method including executing processing in one of at least three regions and converting a virtual address of a memory access executed from the current region to a physical address within one of a plurality of physical address spaces selected based at least on the current region, wherein the at least three regions include a root region for managing switching between a plurality of other regions among the at least three regions, and the plurality of physical address spaces include a root physical address space associated with the root region, which is different from the physical address spaces associated with the plurality of other regions.

[0180] (20) A computer program for controlling a host data processing apparatus that provides an instruction execution environment for executing target code, the computer program comprising processing program logic for simulating processing of the target code in one of at least three regions and address conversion program logic for converting a virtual address of a memory access executed from the current region to a physical address within one of a plurality of simulated physical address spaces selected based at least on the current region, wherein the at least three regions include a root region for managing switching between a plurality of other regions among the at least three regions, and the plurality of simulated physical address spaces include a simulated root physical address space associated with the root region, which is different from the simulated physical address spaces associated with the plurality of other regions.

[0181] A computer-readable storage medium storing the computer program according to item (20) of item (21).

[0182] In the present application, the term "configured to..." is used to mean that an element of a device has a configuration capable of performing a defined operation. In this context, "configuration" means a way of arranging or interconnecting hardware or software. For example, a device may have dedicated hardware that provides a defined operation, or a processor or other processing device may be programmed to execute the function. "Configured to" does not mean that an element of the device needs to be modified in any way to provide a defined operation.

[0183] Although exemplary embodiments of the present invention are described in detail herein with reference to the accompanying drawings, it is understood that the present invention is not limited to these exact embodiments, and that various changes and modifications can be made by those skilled in the art without departing from the scope of the present invention as defined by the appended claims.

Claims

1. An address translation circuit configured to convert a target virtual address specified by a memory access request issued by a requester circuit into a target physical address, and a requester-side filtering circuit configured to perform a granularity protection lookup based on the target physical address and a selected physical address space associated with the memory access request to determine whether to pass the memory access request to a cache or pass it to an interconnect for communication with a completer device to process the memory access request A device comprising: The selected physical address space is one of a plurality of physical address spaces, In the granularity protection lookup, the requester-side filtering circuit Obtains granularity protection information corresponding to the target granularity of the physical address including the target physical address, the granularity protection information indicating at least one permitted physical address space associated with the target granularity, Blocks the memory access request when the granularity protection information indicates that the selected physical address space is not one of the at least one permitted physical address spaces Configured as A physical alias point (PoPA) memory system component configured to de-alias a plurality of aliased physical addresses from different physical address spaces corresponding to the same memory system resource and map any of the plurality of aliased physical addresses to an unaliased physical address provided to at least one downstream memory system component, and At least one PoPA pre-memory system component provided upstream of the PoPA memory system component, the at least one PoPA pre-memory system component configured to handle the aliased physical addresses from different physical address spaces such that the aliased physical addresses from different physical address spaces appear to correspond to different memory system resources for the PoPA pre-memory system component A device comprising.

2. The granularity protection information specifies the assigned physical address space assigned to the target granularity of the physical address, The apparatus according to claim 1, wherein the at least one permitted physical address space includes at least the allocated physical address space.

3. The apparatus according to claim 2, wherein the granularity protection information further includes sharing attribute information indicating whether at least one other physical address space other than the allocated physical address space is one of the at least one permitted physical address spaces.

4. The apparatus according to claim 3, wherein when the granularity protection lookup determines that the selected physical address space is a physical address space other than the allocated physical address space that is one of the at least one permitted physical address spaces indicated by the sharing attribute information, the requester-side filtering circuit is configured to pass the memory access request to the cache or the interconnect by designating the allocated physical address space instead of the selected physical address space.

5. The apparatus according to any one of claims 1 to 4, wherein the requester-side filtering circuit is configured to perform the granularity protection lookup in at least one lookup cache configured to cache the granularity protection information.

6. The apparatus according to claim 5, wherein the at least one lookup cache is configured to store at least one entry in which at least one conversion and granularity protection are combined, the entry specifying information that depends on both the granularity protection information and at least one page table entry used by the address conversion circuit to map the target virtual address to the target physical address.

7. The apparatus according to any one of claims 5 and 6, wherein the at least one lookup cache invalidates a lookup cache entry that stores information depending on the granularity protection information associated with the granularity of the physical address including the invalidation target physical address in response to at least one lookup cache invalidation command specifying the invalidation target physical address.

8. The apparatus according to any one of claims 1 to 7, wherein the aliased physical address is represented using the same physical address value within the different physical address spaces.

9. wherein the at least one PoPA pre-memory system component includes at least one PoPA pre-cache, the apparatus includes a processing circuit that issues at least one invalidation command for requiring the at least one PoPA pre-cache to invalidate one or more entries associated with a target physical address value corresponding to the target address in response to a cache invalidation command up to a PoPA specifying the target address, the apparatus according to any one of claims 1 to 8.

10. wherein at least one of the address translation circuit and the requester-side filtering circuit is configured to select the selected physical address space based at least on a current operating region of the requester circuit in which the memory access request was issued, the current operating region including one of a plurality of operating regions, the apparatus according to any one of claims 1 to 9.

11. the address translation circuit is configured to convert the target virtual address to the target physical address based on at least one page table entry, when at least the current operating region is one of a subset of the plurality of operating regions, the at least one of the address translation circuit and the requester-side filtering circuit is configured to select the selected physical address space based on the current operating region and physical address space selection information specified by the at least one page table entry, the apparatus according to claim 10.

12. wherein the plurality of operating regions includes at least a secure region associated with a secure physical address space and a less secure region associated with a physical address space that is less secure than the secure physical address space, when the current operating region is the less secure region or the secure region, the less secure physical address space is selectable as the selected physical address space, The secure physical address space can be selected as the selected physical address space when the current operating area is the secure area, and it is prohibited from being selected as the selected physical address space when the current operating area is the less secure area. The device according to any one of claims 10 and 11.

13. The plurality of operating areas further includes additional operating areas associated with additional physical address spaces, When the current operating area is the additional operating area, the less secure physical address space can be selected as the selected physical address space, When the current operating area is the additional operating area, the additional physical address space can be selected as the selected physical address space, but it is prohibited from being selected as the selected physical address space when the current operating area is the secure area or the less secure area, When the current operating area is the additional operating area, it is prohibited from selecting the secure address space as the selected physical address space. The device according to claim 12.

14. The plurality of operating areas includes a root area for managing switching between other areas, and the root area is associated with a root physical address space. The device according to any one of claims 10 to 13.

15. When the current operating area is the root area, all of the physical address spaces can be selected as the selected physical address space, When the current operating area is an area other than the root area, it is prohibited from selecting the root physical address space as the selected physical address space, The device according to claim 14.

16. The granular protection information can be modified by software executed in the root area. The device according to any one of claims 14 and 15.

17. The granular protection information is defined in page-level detail. The device according to any one of claims 1 to 16.

18. The granular protection information can be dynamically updated during execution. The device according to any one of claims 1 to 17.

19. When the granularity protection information designates a pass-through indicator indicating that the at least one permitted physical address space is resolved by a completer-side filtering circuit, the requester-side filtering circuit determines whether to pass the memory access request to the cache or the interconnect regardless of whether the selected physical address space is one of the at least one permitted physical address spaces. The apparatus according to any one of claims 1 to 18.

20. A completer-side filtering circuit that responds to a memory access request received from the interconnect that designates a target physical address and a selected physical address space, the completer-side filtering circuit performing a completer-side protection lookup of completer-side protection information based on the target physical address and the selected physical address space to determine whether the memory access request is permitted to be processed by the completer device. The apparatus according to any one of claims 1 to 19.

21. An address translation circuit that translates a target virtual address designated by a memory access request issued by a requester circuit into a target physical address, and A requester-side filtering circuit that performs a granularity protection lookup based on the target physical address and a selected physical address space associated with the memory access request to determine whether the memory access request can be passed to the cache or passed to the interconnect to communicate with the completer device in order to process the memory access request A data processing method including: The selected physical address space is one of a plurality of physical address spaces, In the granularity protection lookup, the requester-side filtering circuit Obtains granularity protection information corresponding to a target granularity of a physical address including the target physical address, the granularity protection information indicating at least one permitted physical address space associated with the target granularity When the granularity protection information indicates that the selected physical address space is not one of the at least one permitted physical address spaces, block the memory access request, A physical alias point (PoPA) memory system component that de-aliases a plurality of aliased physical addresses from different physical address spaces corresponding to the same memory system resource and maps any of the plurality of aliased physical addresses to an un-aliased physical address provided to at least one downstream memory system component; At least one PoPA pre-memory system component provided upstream of the PoPA memory system component, handling the aliased physical addresses from different physical address spaces so that the aliased physical addresses from different physical address spaces appear to correspond to different memory system resources for the PoPA pre-memory system component A data processing method comprising.

Citation Information

Patent Citations

  • DATA PROCESSING APPARATUS AND METHOD USING OWNERSHIP TABLE

    JP2018523209A

  • System and technique for fine-grained computer memory protection

    US7287140B1

  • Invalidation of a target realm in a realm hierarchy

    WO2019002810A1