Data processing apparatus and method and memory system component

Through the Physical Alias ​​Point (PoPA) mechanism, the front PoPA memory system component treats alias addresses in different physical address spaces as different resources, while the back PoPA memory system component ensures security. This solves the problem of balancing security and performance in data processing systems isolated across multiple physical address spaces, and enables more efficient data access.

CN114077496BActive Publication Date: 2026-07-31ARM LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ARM LTD
Filing Date
2021-08-19
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing data processing systems struggle to balance security and performance, particularly in terms of isolation and access control across multiple physical address spaces, where security vulnerabilities exist and performance losses are significant.

Method used

The Physical Alias ​​Point (PoPA) mechanism is employed. Before the PoPA, the memory system component treats alias physical addresses from different physical address spaces as alias physical addresses, and after the PoPA, the memory system component treats them as involving the same memory system resources. This provides a pre-PoPA response action to improve performance when a hit occurs, and a no-data response to ensure safety when a miss occurs.

Benefits of technology

It achieves stronger security guarantees under the isolation of multiple physical address spaces, while improving the performance of the data processing system through speculative access and reducing the performance loss when misses occur.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114077496B_ABST
    Figure CN114077496B_ABST
Patent Text Reader

Abstract

This application discloses a data processing apparatus and method, as well as memory system components. A requester circuit issues an access request specifying a target physical address (PA) and a target physical address space (PAS) identifier, the target PAS identifier identifying the target PAS. Before a Physical Alias ​​Point (PoPA), the pre-PoPA memory system component treats alias PAs from different PASs that actually correspond to the same memory system resource as corresponding to different memory system resources. The post-PoPA memory system component treats the alias PAs as relating to the same memory system resource. When the target PA and target PAS of the pre-PoPA request read on a hit are hit in the pre-PoPA cache, a data response is returned to the requester circuit. If the pre-PoPA request read on a hit does not hit in the pre-PoPA cache, a no-data response is returned. The pre-PoPA request read on a hit is issued speculatively and securely while waiting for security checks to be performed, thereby improving performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technology relates to the field of data processing. Background Technology

[0002] The data processing system may have a requester circuit for issuing an access request to the memory system, and one or more memory system components for responding to the access request issued by the requester circuit and providing access to data stored in the memory system. Summary of the Invention

[0003] At least some examples provide an apparatus comprising: a requester circuit that issues an access request to access a memory system, the access request specifying a target physical address and a target physical address space identifier, the target physical address space identifier identifying a target physical address space selected from a plurality of physical address spaces; and a memory system component that responds to the access request issued by the requester circuit, the memory system component comprising: at least one pre-PoPA memory system component before a Physical Alias ​​Point (PoPA), the at least one pre-PoPA memory system component being configured to treat aliased physical addresses from different physical address spaces that actually correspond to the same memory system resource as aliased physical addresses corresponding to different memory system resources; and at least one post-PoPA memory system component after a PoPA, the at least one post-PoPA memory system component being configured to treat aliased physical addresses as involving The same memory system resources; wherein: in response to a hit-on-hit read request for a specified target physical address and target physical address space identifier issued by the requester circuit, at least one memory system component is configured to provide a hit-on-hit read pre-PoPA response action, including: when the hit-on-hit read pre-PoPA request hits at least one pre-PoPA cache prior to PoPA, providing a data response to the requester circuit to return cached data in the hit entry of the at least one pre-PoPA cache corresponding to the target physical address and target physical address space identifier; and when the hit-on-hit read pre-PoPA request misses in the at least one pre-PoPA cache, providing a no-data response to the requester circuit, the no-data response indicating that the data for the target physical address will not be returned to the requester circuit in response to the hit-on-hit read pre-PoPA read request.

[0004] At least some examples provide a method comprising: issuing an access request from a requesting circuit to access a memory system including memory system components, the access request specifying a target physical address and a target physical address space identifier, the target physical address space identifier identifying a target physical address space selected from a plurality of physical address spaces; and responding to the access request issued by the requesting circuit using memory system components, the memory system components comprising: at least one pre-PoPA memory system component before a Physical Alias ​​Point (PoPA), the at least one pre-PoPA memory system component being configured to treat aliased physical addresses from different physical address spaces that actually correspond to the same memory system resource as aliased physical addresses corresponding to different memory system resources; and at least one post-PoPA memory system component after the PoPA, the at least one post-PoPA memory system component being configured to treat the aliased physical address as a different memory system resource. Physical addresses are considered to involve the same memory system resources; wherein: in response to a pre-PoPA read-on-hit request issued by the requester circuit specifying a target physical address and a target physical address space identifier, at least one memory system component of the memory system components provides a pre-PoPA read-on-hit response action, including: when the pre-PoPA read-on-hit request hits at least one pre-PoPA cache prior to PoPA, providing a data response to the requester circuit to return cached data in the hit entry of the at least one pre-PoPA cache corresponding to the target physical address and the target physical address space identifier; and when the pre-PoPA read-on-hit request misses in the at least one pre-PoPA cache, providing a no-data response to the requester circuit, the no-data response indicating that the data for the target physical address will not be returned to the requester circuit in response to the pre-PoPA read-on-hit request.

[0005] At least some examples provide a memory system component comprising: a requester interface receiving an access request from requester circuitry specifying a target physical address and a target physical address space identifier, the target physical address space identifier identifying a target physical address space selected from a plurality of physical address spaces; and control circuitry detecting whether a hit or miss is detected in a lookup of at least one pre-PoPA cache prior to a Physical Alias ​​Point (PoPA), wherein the at least one pre-PoPA cache is configured to treat aliased physical addresses from different physical address spaces that actually correspond to the same memory system resource as aliased physical addresses corresponding to different memory system resources, and the PoPA is such a point after which at least one post-PoPA memory system component is configured to treat aliased physical addresses as related to different memory system resources. and the same memory system resources; wherein: in response to a pre-PoPA read request issued by the requester circuit for a specified target physical address and target physical address space identifier, the control circuit is configured to provide a pre-PoPA read-on-hit response action, including: when a lookup of the at least one pre-PoPA cache detects a hit, controlling the requester interface to provide a data response to the requester circuit, thereby returning cached data in the hit entry of the at least one pre-PoPA cache corresponding to the target physical address and target physical address space identifier; and when a lookup of the at least one pre-PoPA cache detects a miss, controlling the requester interface to provide a no-data response to the requester circuit, the no-data response indicating that the data for the target physical address will not be returned to the requester circuit in response to the pre-PoPA read-on-hit request. Attached Figure Description

[0006] Further aspects, features, and advantages of this technology will become apparent from the following description, taken in conjunction with the accompanying drawings, in which:

[0007] Figure 1 An example of a data processing device is shown;

[0008] Figure 2 Multiple domains in which the processing circuitry can operate are shown;

[0009] Figure 3 An example of a processing system that supports particle protection lookup is shown;

[0010] Figure 4 This schematically illustrates multiple physical address spaces as aliases on the system physical address space that identify locations within the memory system;

[0011] Figure 5 An example is shown of partitioning the effective hardware physical address space so that different architecture physical address spaces have access to the corresponding portions of the system physical address space;

[0012] Figure 6 This is a flowchart illustrating a method for determining the current operating domain of a processing circuit;

[0013] Figure 7 An example of the page table entry format used to translate virtual addresses into physical addresses is shown;

[0014] Figure 8 This is a flowchart illustrating a method for selecting the physical address space to be accessed based on a given memory access request;

[0015] Figure 9 An example of an entry for a particle protection table is shown, which provides particle protection information indicating which physical address spaces are allowed to access a given physical address.

[0016] Figure 10 This is a flowchart illustrating the method for performing particle protection lookup;

[0017] Figure 11 The multiple stages of address translation and particle protection information filtering are illustrated;

[0018] Figure 12 and Figure 13 The response to the previous PoPA request read upon hit is shown;

[0019] Figure 14 It is a flowchart that shows the requester initiating access to the target physical address in the target physical address space, including the use of a pre-PoPA request read on hit.

[0020] Figure 15 An example of the components of the previous PoPA memory system is shown; and

[0021] Figure 16 This is a flowchart illustrating a method for responding to a pre-PoPA request read upon hit using pre-PoPA memory system components. Detailed Implementation

[0022] The data processing system can support the use of virtual memory, providing address translation circuitry to translate the virtual address (VA) specified by a memory access request into a physical address (PA) associated with the location in the memory system to be accessed. The mapping between virtual and physical addresses can be defined in one or more page table structures. Page table entries within the page table structure can also define access permission information that controls whether a given software procedure executing on the processing circuitry is allowed to access a specific virtual address.

[0023] In some processing systems, address translation circuitry maps all virtual addresses to a single physical address space, which the memory system uses to identify locations in memory to be accessed. In such systems, control over whether a particular address is accessible to a specific software process is based solely on the page table structure used to provide the virtual-to-physical address translation mapping. However, this page table structure is typically defined by the operating system and / or hypervisor. If the operating system or hypervisor is compromised, this can introduce security vulnerabilities, allowing attackers to access sensitive information.

[0024] Therefore, for systems that require certain processes to be executed securely in isolation from other processes, the system can support multiple different physical address spaces. For at least some components of the memory system, memory access requests whose virtual addresses are translated into physical addresses in different physical address spaces are treated as completely separate addresses accessing memory. By isolating accesses to different operational domains of the processing circuitry into corresponding different physical address spaces as perceived by some memory system components, this provides stronger security guarantees independent of page table permission information set by the operating system or hypervisor.

[0025] Thus, the device may have requester circuitry for issuing an access request to the memory system, wherein the access request specifies a target physical address (target PA) and a target physical address space identifier, the target physical address space identifier identifying a target physical address space (target PAS) selected from two or more distinct physical address spaces (PAS). It should be noted that the symbol “PAS” (where S is uppercase) used below refers to “physical address space”, while the symbol “PAs” (where s is lowercase) refers to “physical address” (the plural of PA).

[0026] Memory system components may be provided to respond to access requests issued by the requester circuitry. The memory system components include: at least one pre-PoPA memory system component prior to the Physical Alias ​​Point (PoPA), which treats alias PAs from different PASs that actually correspond to the same memory system resource as alias PAs corresponding to different memory system resources; and at least one post-PoPA memory system component after the PoPA, which is configured to treat alias PAs as relating to the same memory system resource. Here, "memory system resource" may refer to an addressable location in memory, a device or peripheral device mapped to a memory address in a PAS, a memory mapping register within the memory system component, or any other resource accessed by issuing a memory access request specifying a physical address mapped to that resource.

[0027] By treating alias PAs from different PASs that actually correspond to the same memory system resource as corresponding to different memory system resources (in memory system components prior to PoPA), this provides stronger security guarantees as mentioned above. For example, when the current PoPA memory system component includes a pre-PoPA cache, a request specifying a given PA and the first PAS may miss cache entries for data associated with the same given PA in the second PAS, thus preventing data from the second PAS from being returned to such software processes that operate on the first PAS, even if the page table does not prevent software processes that operate on the first PAS from accessing the given PA.

[0028] A protocol can be established that defines the format of access requests that can be issued by a requester circuit (or by an upstream memory system component to a downstream memory system component) to control access to memory, and defines the corresponding responses that memory system components take in response to access requests. For example, the protocol can define various types of read requests to read data from the memory system and write requests to write data stored in the memory system.

[0029] In the example discussed further below, the requester circuitry supports issuing a pre-PoPA request for a hit-on-demand read specifying a target PA and a target PAS identifier. In response to the hit-on-demand read pre-PoPA request, at least one memory system component provides a hit-on-demand read pre-PoPA response action, which includes: providing a data response to the requester circuitry to return cached data in the hit entry of the at least one pre-PoPA cache corresponding to the target physical address and the target PAS identifier when the hit-on-demand read pre-PoPA request hits at least one pre-PoPA cache prior to the PoPA; and providing a no-data response to the requester circuitry when the hit-on-demand read pre-PoPA request misses in the at least one pre-PoPA cache, the no-data response indicating that the data for the target physical address will not be returned to the requester circuitry in response to the hit-on-demand read pre-PoPA read request.

[0030] Therefore, a hit-on-demand pre-PoPA request allows the requester circuitry to make a conditional request to return data from the memory system if the request hits a cache prior to the PoPA, but not if the request misses any pre-PoPA cache. This hit-on-demand pre-PoPA request can be used to improve performance because it can be speculatively issued before any checks to determine whether access to a given physical address within a given PAS have been completed. Data already cached in at least one pre-PoPA cache for a given target PA and target PAS identifier can be assumed to have passed such checks and can therefore be safely accessed before such checks are completed, but if a hit-on-demand pre-PoPA request misses any pre-PoPA cache, there is no guarantee that access to the target PA within the target PAS is allowed, and therefore a no-data response may be returned. In contrast, with a regular read request, data is returned to the requester regardless of whether it hit or missed in the pre-PoPA cache. This is because if the request misses in the pre-PoPA cache, a line-filling operation is typically expected to request that data be moved from memory to the pre-PoPA cache and returned to the requester. However, by defining a request type that returns a no-data response if the request misses in at least one of the pre-PoPA caches, this allows the request to be issued while simultaneously inferring the success of any security checks on the target physical address and the target PAS, thus allowing for performance improvements.

[0031] Issuing a read-on-hit pre-PoPA request while presuming a successful security check is one use case. However, once read-on-hit pre-PoPA is included in the memory interface protocol to be provided to the requester circuitry, system designers can find other use cases where this request might be useful. The protocol does not limit the specific scenarios in which read-on-hit pre-PoPA requests are used. Therefore, it is possible that read-on-hit pre-PoPA requests can be used in other scenarios, not just those where a security check is awaited to complete to check if the target PA can be accessed within the target PAS.

[0032] When a pre-PoPA read-on-hit request misses in at least one pre-PoPA cache, the allocation of data for the target physical address in the at least one pre-PoPA cache is prevented based on line-filling requests issued to at least one post-PoPA memory system component in response to a pre-PoPA read-on-hit request. Thus, if a pre-PoPA read-on-hit request hits in at least one pre-PoPA cache (and therefore the required data is already in the PoPA pre-cache), the request is allowed to return the data; if a miss occurs, data is not allowed to be brought from post-PoPA memory system components to the pre-PoPA cache. Similarly, for access request types other than pre-PoPA read-on-hit requests, the allocation of data to the pre-PoPA cache based on line-filling requests to post-PoPA memory system components is prohibited until a security check to determine whether the target physical address can be safely accessed within the target PAS has been successfully completed. Using this method, it can be ensured that a pre-PoPA request read upon a hit can safely access data in the hit entry of the pre-PoPA cache, even if the request was made before the security check was completed, because the presence of data in the pre-PoPA cache can be an indication of the success of the previous security check of the target PA and the target PAS.

[0033] Various options exist to prevent data for the target physical address from being allocated to the pre-PoPA cache when a pre-PoPA read request that was read on a hit misses. In one example, the pre-PoPA memory system component processing a pre-PoPA read request that was read on a hit can simply prevent a line-fill request from being issued to a subsequent PoPA memory system component in response to a pre-PoPA read request that was read on a hit when a miss occurs in at least one of the pre-PoPA caches. Alternatively, some implementations may provide a buffer associated with the pre-PoPA memory system component that buffers data returned in response to a line-fill request issued in response to a pre-PoPA read request that was read on a hit when a miss occurs. Subsequent requests may not be able to access this buffer until confirmation has been received that the security check of the target physical address and the target PAS identifier has been successfully completed. Thus, there are multiple microarchitectural options to ensure that data for the target physical address is not allocated to the pre-PoPA cache when a pre-PoPA read request that was read on a hit misses in the pre-PoPA cache. When using this buffer option, subsequent requests after security checks have been completed can be executed faster than if a row fill request could only be issued at that point in time, but the buffer method can be more complex to implement. Both options are possible.

[0034] In some examples, there may be two or more levels of the pre-PoPA cache preceding the physical alias point. In this case, if any of the two or more levels of the pre-PoPA cache includes a valid entry corresponding to the target PA and target PAS identifiers, the pre-PoPA request read at the time of the hit is considered to have hit the pre-PoPA cache. In contrast, if none of the two or more levels of the pre-PoPA cache include a valid entry corresponding to the target PA and target PAS identifiers, a miss can be detected. It should be noted that in the case where the pre-PoPA request read at the time of the hit hits one of the pre-PoPA cache levels, there is a possibility of some data transfer between the corresponding levels of the pre-PoPA cache. This is because a hit in a pre-PoPA cache lower than that level (e.g., level 2 or 3) can cause the data to be upgraded to a pre-PoPA cache higher than that level (e.g., level 1 or 2), where the requester circuitry can access the data faster, and this upgrade of the target PA data can also trigger other addresses to downgrade their data to a cache lower than that level. Therefore, while a pre-PoPA request that reads on hit is not allowed to cause data to be promoted from a position after PoPA to the pre-PoPA cache, this does not preclude the movement of data between various levels of the pre-PoPA cache.

[0035] In some specific implementations, the only response action to the previous PoPA request read when a hit occurs in a miss scenario can be to return a no-data response.

[0036] However, for some examples, in addition to returning a no-data response, when a pre-PoPA request read on hit misses in at least one pre-PoPA cache, the pre-PoPA response action may also include requesting at least one of the one or more post-PoPA memory system components to perform at least one preparation operation to prepare for processing of a subsequent access request to the specified target PA. It is recognized that, in the case of a pre-PoPA request read on hit missing in at least one pre-PoPA cache, although it is not yet safe to return the corresponding data for the target PA, it is relatively likely that a subsequent request to the same target PA will be made once any security checks have been completed (because successful security checks are more common than unsuccessful ones). Therefore, by performing a preparation operation in response to a pre-PoPA request read on hit when the request misses in one or more pre-PoPA caches, this speeds up the processing of a subsequent access request to the specified target PA, since part of the operations required to access the data associated with the target PA from the post-PoPA memory system component may have already been performed while processing the pre-PoPA request read on hit.

[0037] Preparation operations can include any operation that can be predicted to cause subsequent access requests for the same target PA to be processed faster than if no preparation operation has been performed.

[0038] For example, in some implementations, the at least one preparation operation may include prefetching data associated with the target PA into the post-PoPA cache. Unlike the pre-PoPA cache, which treats alias PAs from different PASs as belonging to different memory system sources even if they actually correspond to the same memory system resource, the post-PoPA cache treats alias PAs as belonging to the same memory system resource. Since data is not moved from the post-PoPA cache to the pre-PoPA cache unless relevant security checks are performed, and the post-PoPA caches data for a given PA in the same way regardless of which PAS can access the given PA, data can be safely prefetched into the post-PoPA cache in response to a pre-PoPA request that is read upon a hit if the request misses in the pre-PoPA cache. By performing the prefetch operation, it means that when subsequent read requests are made to the same target PA, data can be returned from the post-PoPA cache faster than data must be retrieved from main memory. This contributes to improved performance.

[0039] Another example of a preparation operation could be performing a precharge or activation operation to prepare a memory location associated with a target PA for access. Some forms of memory technology, such as dynamic random access memory (DRAM), require precharging or activating a memory location before it can be read, and this precharge or activation operation takes a certain amount of time to perform. By speculatively performing a precharge or activation operation in response to a previous PoPA request that is read upon a hit, in the event of a cache miss in the previous PoPA cache, it is less likely that a precharge or activation operation will be required again when a subsequent access request attempts to access the same target PA. Therefore, access to the memory location associated with the target PA can be faster than in the case where a precharge or activation operation has not been performed.

[0040] These two examples of preparation operations (prefetching or performing a precharge or activation operation) can be implemented within the same system so that a prefetch request and a precharge or activation operation are initiated when a pre-PoPA request read on a hit misses in the pre-PoPA cache. Alternatively, some systems may not support both operations. For example, a system without any post-PoPA cache may perform only a precharge or activation operation as a preparation operation. It should be understood that other types of preparation operations may also be performed.

[0041] Additionally, preparation operations may be unnecessary, and some implementations may choose not to initiate any preparation operations if the previous PoPA request read on hit is not hit.

[0042] Protection checking circuitry can be provided to perform protection checks, thereby determining whether access to a given access request of the target PA is permitted within the target PAS based on a lookup of protection information corresponding to the target PA. The protection information may indicate at least one permitted PAS from which access to the target PA is allowed. Multiple methods may exist to indicate which PAS is permitted. For example, some implementations may support protection information that identifies only a single PAS from which access to the target PA is permitted, while other methods may support multiple PAS mapping to the same target PA, such that access to the target PA is permitted from any of these multiple PASs. The specific format of the protection information can vary considerably, but generally, the protection information can be any information used to determine whether a memory access request of a PA in a specified selected PAS is permitted to access corresponding data.

[0043] The protection check circuit can be a requester-side check circuit, which is provided to perform a protection check before issuing a corresponding request to the cache or other memory system components (at least for some types of requests - as mentioned above, a pre-PoPA request read on hit can be an exception because it does not need to wait for the protection check to complete).

[0044] In some examples, protection information used for protection checks can be derived from a protection table stored in memory. This allows for finer-grained control over which portions of the accessible physical address range in each PAS compared to a specific implementation where the protection information is stored in registers local to the protection check circuitry. However, when the protection information can be stored within the memory system, this could mean that when a protection check is required on a target PA, the protection check circuitry can issue one or more protection table roaming access requests to memory system components to perform a protection table roaming, thereby obtaining the protection information from the protection table indexed by the target PA.

[0045] For example, a protection table can be a multi-level protection table, where information derived from one level of the protection table can be used to identify the location of subsequent levels of the protection table. Such a multi-level table can help allow the protection table to be represented using multiple smaller, discontinuous portions of memory, thus avoiding the need to allocate a single contiguous block of address space for a table proportional to the size of the address range to be protected, which is necessary if a table with a single linear access is used. Particularly in the multi-level case, roaming such protection tables can be relatively slow because they may require a continuous chain of accesses to memory, where later accesses in the chain depend on values ​​returned from previous accesses. However, even with a single-level table, accesses to memory to obtain protection information still introduce some additional latency after the target PA of the memory access is established.

[0046] While it's possible to provide a protection information cache so that protection check circuitry can access information recently accessed from the protection table locally faster than it can access memory, one or more protection table roaming access requests may be required if the protection information for the target PA is not yet available in such a cache (either because such a cache is not provided at all, or because a request to the protection information cache for the target PA misses). If a protection information cache is provided, it can be a dedicated cache structure that stores protection information from the protection table, or it can be a shared cache structure that can be shared with other types of data, such as a translation lookahead buffer (TLB) / protection information cache that caches both the protection information and the address translation information used to translate VA to PA.

[0047] Therefore, protection checks can be relatively slow operations because they rely on obtaining protection information from memory. If a read request cannot be safely issued to the memory system until after the protection check is complete, this can introduce additional latency in accessing memory between when the PA for memory access becomes available after any address translation and when the actual read request is issued to the memory system components.

[0048] The aforementioned pre-PoPA request read on hit solves this problem because the requester circuit can issue the pre-PoPA request read on hit before completing the protection checks of the target PA and target PAS.

[0049] If a protection check determines that access to the target physical address within the target PAS is permitted when a no-data response is received in response to a pre-PoPA request read on a hit, the requester circuit may issue a read request specifying the same target PA and target PAS as the pre-PoPA request read on a previous hit. In response to the read request, at least one of the memory system components may provide a data response to return the data associated with the target PA to the requester circuit (even if the read request misses in the at least one pre-PoPA cache). In response to the read request, if the protection check has been successfully performed, the memory system component is permitted to cause the data associated with the target PA to be allocated into the at least one pre-PoPA cache based on a line-fill request issued to at least one of the at least one post-PoPA memory system component.

[0050] The read request issued after a successful protection check can be a type of read request that is prohibited from being issued by the requester circuit if the protection check determines that access to the target physical address within the target PAS is not permitted. This type of read request may be disallowed from the requester circuit until the protection check has been completed and determined that access to the target PA within the target PAS is permitted.

[0051] When a protection check determines that access to the target physical address within the target PAS is not permitted, the protection circuit can signal a fault. This fault can trigger the execution of an exception handler, which determines (in software) how to handle the fault.

[0052] As described, even if the alias PA address actually corresponds to the same memory system resource, the at least one pre-PoPA cache can treat the alias PA addresses as corresponding to different memory system resources. In one example, this can be achieved by tagging each cache entry in the at least one pre-PoPA cache with a corresponding PAS identifier. When a lookup request for an access request specifying a target PA and a target PAS is made in the pre-PoPA cache, the target PA address tag derived from the target PA can be compared with the address tag stored in one or more lookup cache entries, and the target PAS of the access request can be compared with the PAS tag stored in the one or more lookup cache entries. A cache hit can be detected when a valid entry in the lookup cache entries has both the stored address tag and the stored PAS tag corresponding to the target PA address tag and the target PAS identifier. A cache miss can be detected even if there are valid entries in the at least one pre-PoPA cache that correspond to the target PA but a different PAS identifier than the target PAS identifier in the access request, when none of the valid entries in the at least one pre-PoPA cache correspond to both the target PA and the target PAS identifier in the access request. (That is, entries whose stored address labels correspond to the target PA address labels but whose stored PAS labels do not correspond to the target PA address labels will return a cache miss.) Thus, it is possible for a given pre-PoPA to maintain separate valid cache entries corresponding to the same PA but different PAS at a given time. This method ensures that data from a single PAS is not returned in response to a request specifying a different PAS. This can be used to isolate data used by different software processes to ensure that sensitive data associated with one software process is not accessed by another software process that is not allowed to access that data, without relying on page table access permissions set by the operating system or hypervisor to implement this separation. In most use cases, only the PA used for sensitive data can be assigned to a single PAS. In some specific implementations, limited sharing of data from different PASs can be supported, where a given physical address can be mapped to two or more different PASs, and in this case, it is possible for the front PoPA cache to simultaneously cache multiple entries corresponding to the same target physical address but with different PAS identifiers.

[0053] The device may include address translation circuitry that translates a target virtual address (VA) based on address translation information corresponding to that VA, thereby obtaining the target physical address for a given access request. This can be a single-stage address translation directly from the target VA to the target PA, or a two-stage address translation based on first translation information mapping the target VA to a target intermediate address and second address translation information mapping the target intermediate address to the target PA. The address translation circuitry may also perform access permission checks based on permissions set in the address translation information. These protection checks may be additional layers of checks provided as a supplement to the access permission checks performed by the address translation circuitry.

[0054] The apparatus may have a PAS selection circuit that selects a target PAS identifier corresponding to a target PAS based on at least one of the following: the current operating domain of the requester circuit, and information specified in the address translation information corresponding to the target virtual address. The PAS selection circuit may, in some cases, be the address translation circuit itself, or the protection check circuit that performs protection checks based on protection information of the target physical address as described previously, or a completely separate circuit segment that selects which PAS as the target PAS. In some implementations, multiple operating domains may be supported for the requester circuit, which may support the execution of different software processes in different domains. The selection of the target PAS may depend on the current domain, such that some domains may be limited to accessing PAS selected from a finite subset of PASs. In some domains (e.g., the root domain described later), the selected PAS may be any of the PASs. In some cases, at least one operating domain may be allowed to select two or more different PAS identifiers as the target PAS identifier, and in this case, the selection of the target PAS identifier may depend on information corresponding to the target virtual address defined in the address translation information. This approach provides an architecture that offers the flexibility to support different software use cases where different processes running on the same hardware platform require some security guarantees.

[0055] The requester circuitry mentioned above can be any component in a data processing system that can act as the source of an access request for requesting access to memory system resources. For example, the requester circuitry can be a processor core capable of executing program instructions, such as a central processing unit (CPU) or graphics processing unit (GPU), or it can be another type of host device in the data processing system that may not necessarily be capable of executing program instructions but can act as the source of a request for memory (e.g., a display controller, network controller, or direct memory access (DMA) unit). Additionally, in some cases, internal memory system components may be able to initiate (or forward) requests. For example, interconnect components or cache controllers may be able to generate access requests and sometimes act as the requester circuitry for a given memory access. Thus, a pre-PoPA request for read-on-hit can be issued by any of these exemplary requester circuitries. In one example, if the Level 1 cache preceding PoPA receives a pre-PoPA request for a hit-read from the processor core and detects a miss, a line fill request issued to the Level 2 cache, which also precedes PoPA, can be designated as a pre-PoPA request for a hit-read to ensure that if the request misses in the Level 2 cache, the Level 2 cache will not pull data from the post-PoPA memory system component.

[0056] A memory system component may be provided, comprising a requester interface for receiving access requests from requester circuitry, and control circuitry for detecting whether a hit or miss is detected for the access request in a lookup of at least one pre-PoPA cache prior to the physical alias point. When the access request is a pre-PoPA request for a hit-on-read as discussed above, the control circuitry may provide a pre-PoPA response action for a hit-on-read (with a data response if the request hits in the pre-PoPA cache and a no-data response if it misses, as previously explained). Such a memory system component may be included at a point prior to the PoPA within a processing system to support the processing of pre-PoPA requests for hit-on-read as previously described. Detailed Implementation

[0058] Figure 1 An example of a data processing system 2 having at least one requester device 4 and at least one completer device 6 is schematically illustrated. Interconnect 8 provides communication between the requester device 4 and the completer device 6. The requester device is capable of issuing a memory access request for memory access to a specific addressable memory system location. The completer device 6 is the device responsible for servicing memory access requests directed to it. Although... Figure 1Not shown, but some devices may be able to act as both requester and completer devices. Requester device 4 may include, for example, processing elements such as a central processing unit (CPU) or a graphics processing unit (GPU), or other host devices such as a bus master, network interface controller, display controller, etc. Completer devices may include memory controllers responsible for controlling access to corresponding memory storage units, peripheral controllers for controlling access to peripheral devices, etc. Figure 1 An exemplary configuration of one of the requester devices 4 is shown in more detail, but it should be understood that other requester devices 4 may have similar configurations. Alternatively, other requester devices may have the same... Figure 1 The requester device 4 shown on the left has a different configuration. In this example, one of the completer devices 6 is a memory controller that controls access to memory 9. In this example, the completer device 6 has a memory controller cache 7, which can be used to cache some data read from memory 9 so that requests from interconnect 8 can be processed faster if they hit the memory controller cache 7 than if they miss and require access to memory 9.

[0059] The requester device 4 has processing circuitry 10 that performs data processing in response to instructions, referencing data stored in register 12. Register 12 may include a general-purpose register for storing operands and the results of the processed instructions, and a control register for storing control data to configure how the processing circuitry performs processing. For example, the control data may include a current domain 14 indicator for selecting which operation domain is the current domain, and a current exception level (EL) 15 indicator for indicating which exception level is the current exception level that the processing circuitry 10 is operating at.

[0060] Processing circuitry 10 may issue a memory access request specifying a virtual address (VA) identifying the addressable location to be accessed and a domain identifier (domain ID or "security state") identifying the current domain. Address translation circuitry 16 (e.g., a memory management unit (MMU)) translates the virtual address into a physical address (PA) through one or more stages of address translation based on page table data defined in a page table structure stored in the memory system. Translation lookup buffer (TLB) 18 acts as a lookup cache to cache some information in the page table information, thus enabling faster access than if the page table information had to be retrieved from memory every time an address translation is needed. In this example, in addition to generating the physical address, address translation circuitry 16 selects one of several physical address spaces associated with the physical address and outputs a Physical Address Space (PAS) identifier identifying the selected physical address space. The selection of PAS will be discussed in more detail below.

[0061] PAS filter 20 acts as a requester-side filter circuit to check whether access to the physical address within the specified physical address space identified by the PAS identifier is permitted based on the translated physical address and PAS identifier. This lookup is based on particle protection information stored in a particle protection table structure within the memory system. Similar to the cache of page table data in TLB 18, particle protection information can be cached within particle protection information cache 22. Although particle protection information cache 22 is in... Figure 1 In the example shown, it is a separate structure from TLB 18, but in other examples, these types of lookup caches can be combined into a single lookup cache structure so that a single lookup of an entry in the combined structure provides both page table information and granular protection information. Granular protection information defines information that restricts access to the physical address space of a given physical address, and based on this lookup, PAS filter 20 determines whether to allow memory access requests to continue being issued to one or more caches 24 and / or interconnects 8. If the specified PAS access to the specified physical address is not permitted for the memory access request, PAS filter 20 blocks the transaction and may signal a fault.

[0062] Although Figure 1 An example of a system including multiple requester devices 4 is shown, but for Figure 1 The feature shown by a requester device on the left-hand side can also be included in systems where only one requester device exists (such as a single-core processor).

[0063] Although Figure 1 An example of the address translation circuitry 16 and PAS filter 20 provided within the requester device 4 is shown, but other types of requesters may use the address translation functionality provided by a separate System Memory Management Unit (SMMU), which is a component separate from the requester device 4 itself. In this case, the SMMU can be coupled to the interconnect and can function as the previously mentioned address translation circuitry and protection check circuitry, similar to... Figure 1 The address translation circuit 16 and PAS filter 20 are shown.

[0064] Although Figure 1 An example is shown where address translation circuit 16 performs the selection of a PAS for a given request. However, in other examples, address translation circuit 16 may output information for determining which PAS to select, along with the PA, to PAS filter 20, and PAS filter 20 may select the PAS and check whether access to the PA is permitted within the selected PAS.

[0065] The provision of PAS filter 20 helps support systems that can operate in multiple operational domains, each associated with its own isolated physical address space. This means that, for at least a portion of the memory system (e.g., for some cache or coherence implementation such as a snooping filter), even if addresses within these address spaces actually relate to the same physical location in the memory system, each individual physical address space is treated as a completely separate set of addresses that identify the individual memory system location. This can be useful for security purposes.

[0066] Figure 2 Examples of different operating states and domains in which the processing circuit 10 can operate are shown, as well as examples of the types of software that can be executed in different exception levels and domains (of course, it should be understood that the specific software installed on the system is selected by the parties managing the system and is therefore not a fundamental feature of the hardware architecture).

[0067] The processing circuitry 10 can operate at multiple different exception levels 80 (in this example, four exception levels labeled EL0, EL1, EL2, and EL3), where EL3 refers to the exception level with the highest privilege level, and EL0 refers to the exception level with the lowest privilege level. It should be understood that other architectures may choose the reverse numbering so that the exception level with the highest number can be considered to have the lowest privilege. In this example, the lowest privilege exception level EL0 is used for application-level code, the next highest privilege exception level EL1 is used for operating system-level code, the next highest privilege exception level EL2 is used for hypervisor-level code managing the switching between multiple virtualized operating systems, and the highest privilege exception level EL3 is used for monitoring code managing the switching between corresponding domains and the allocation of physical addresses to the physical address space, as described later.

[0068] When an exception occurs while the processing software is at a specific exception level, for some types of exceptions, an exception of a higher (higher privilege) level is generated, where the specific exception level to generate the exception is selected based on the attributes of the specific exception that occurred. However, in some cases, it is possible for other types of exceptions to be generated at the same exception level as the exception level associated with the code that was processed when the exception occurred. When an exception occurs, information characterizing the state of the processor at the time the exception occurred can be saved, including, for example, the current exception level at the time the exception occurred. Therefore, once an exception handler has been processed to handle the exception, processing can return to the previous processing, and the saved information can be used to identify the exception level to which the processing should return.

[0069] In addition to different exception levels, the processing circuitry supports multiple operating domains, including the root domain 82, the safe (S) domain 84, the less safe domain 86, and the domain 88. For ease of reference, the less safe domain will be described hereinafter as the “unsafe” (NS) domain, but it should be understood that this is not intended to imply any particular level of safety (or lack thereof). Rather, “unsafe” simply indicates that the unsafe domain is intended for code that is not as safe as code operating in a safe domain. The root domain 82 is selected when the processing circuitry 10 is at the highest exception level EL3. When the processing circuitry is at one of the other exception levels EL0 through EL2, the current domain is selected based on the current domain indicator 14, which indicates which of the other domains 84, 86, and 88 is active. For each of the other domains 84, 86, and 88, the processing circuitry can be at any exception level of EL0, EL1, or EL2.

[0070] At boot time, multiple fragments of boot code (e.g., BL1, BL2, OEM boot) may be executed, for example, within a higher privilege exception level EL3 or EL2. Boot code BL1 and BL2 may be associated with, for example, a root domain, and the OEM boot code may operate within a security domain. However, once the system is booted, during runtime, processing circuitry 10 can be considered to operate at one time in one of domains 82, 84, 86, and 88. Each of domains 82 through 88 is associated with its own associated physical address space (PAS), which achieves the isolation of data from different domains within at least a portion of the memory system. This will be described in more detail below.

[0071] Non-security domain 86 can be used for regular application-level processing and for operating system and hypervisor activities to manage such applications. Thus, within non-security domain 86, there can be application code 30 operating at EL0, operating system (OS) 32 operating at EL1, and hypervisor 34 operating at EL2.

[0072] Security domain 84 enables the isolation of certain on-chip security, media, or system services into a separate physical address space from the physical address space used for non-secure processing. Non-secure domain code cannot access resources associated with security domain 84, while secure domain code can access both secure and non-secure resources; in this sense, secure and non-secure domains are not equivalent. An example of a system supporting this partitioning of security domain 84 and non-secure domain 86 is based on Arm... ® TrustZone provided by Limited ®The system architecture includes a trusted application 36 at EL0, a trusted operating system 38 at EL1, and optionally a secure partition manager 40 at EL2. If secure partitioning is supported, the secure partition manager uses stage 2 page tables to support isolation between different trusted operating systems 38 running in the secure domain 84, in a manner similar to how the hypervisor 34 manages isolation between virtual machines or guest operating systems 32 running in the non-secure domain 86.

[0073] Extending this system to support security domain 84 has become common in recent years because it enables a single hardware processor to support isolated secure processing, thus avoiding the need to execute that processing on a separate hardware processor. However, with the increasing prevalence of security domains, many real-world systems with such security domains now support relatively complex hybrid service environments offered by a wide variety of different software vendors within the security domain. For example, code operating in security domain 84 may include different pieces of software provided by, among other things: silicon wafer providers that manufacture integrated circuits; original equipment manufacturers (OEMs) that assemble the integrated circuits provided by the silicon wafer providers into electronic devices such as mobile phones; operating system vendors (OSVs) that provide the operating system 32 for such devices; and / or cloud platform providers that manage cloud servers that support services for multiple different clients via the cloud.

[0074] However, there is a growing desire to provide secure computing environments for parties providing user-level code (which may typically be expected to execute as application code 30 within a non-secure domain 86), environments that can be trusted not to leak information to other parties operating the code on the same physical platform. It may be desirable for such secure computing environments to be dynamically allocated at runtime and to be certified and provable, allowing users to verify adequate security guarantees on the physical platform before trusting the device to process potentially sensitive code or data. Users of such software may not want to trust a party providing a rich operating system 32 or hypervisor 34 that may typically operate in a non-secure domain 86 (or even if these providers are trusted, users may want to protect themselves from attackers who could compromise the operating system 32 or hypervisor 34). Furthermore, while a secure domain 84 can be used for such user-provided applications requiring secure processing, this can actually create problems for both users providing code that requires a secure computing environment and providers of existing code operating within a secure domain 84. For providers of existing code operating within security domain 84, adding arbitrary user-supplied code within the security domain would increase the attack surface of their code, which is likely undesirable. Therefore, it is strongly recommended that users not be allowed to add code to security domain 84. On the other hand, users providing code that requires a secure computing environment may be reluctant to trust that all providers of different fragments of code operating within security domain 84 have access to their data or code. If authentication or certification of code operating within a specific domain is required as a prerequisite for user-supplied code to perform its processing, it may be difficult to audit and authenticate all different fragments of code operating within security domain 84 provided by different software providers. This could limit opportunities for third parties to provide more secure services.

[0075] Therefore, as Figure 2As shown, an additional domain 88 (referred to as a domain domain) is provided, which can be used by code introduced by such users to provide a secure computing environment orthogonal to any secure computing environment associated with components operating in the secure domain. Within a domain domain, the software executed may include multiple domains, each of which can be isolated from other domains by a Domain Management Module (RMM) 46 operating at exception level EL2. The RMM 46 can control the isolation between corresponding domains 42, 44 executing domain domain 88, for example, by defining access permissions and address mappings in a page table structure, in a manner similar to how the hypervisor 34 manages the isolation between different components operating in non-secure domain 86. In this example, the domains include an application-level domain 42 executing at EL0 and a packaged application / operating system domain 44 executing across exception levels EL0 and EL1. It should be understood that it is not necessary to support both EL0 and EL0 / EL1 type domains simultaneously, and multiple domains of the same type can be created by the RMM 46.

[0076] Similar to security domain 84, domain 88 has its own physical address space allocated to it. However, while domain 88 and security domain 84 can each access the non-secure PAS associated with non-secure domain 86, they cannot access each other's physical address spaces. In this sense, the domain is orthogonal to security domain 84. This means that the code executing in domain 88 and security domain 84 is independent of each other. The code in the domain only needs to trust the switching code between the hardware RMM 46 and the management domain operating in root domain 82, which makes proof and authentication more feasible. Proof enables a given piece of software to request verification that the code installed on the device matches certain expected characteristics. This can be achieved by checking whether the hash of the program code installed on the device matches an expected value signed by a trusted party using a cryptographic protocol. RMM 46 and monitoring code 29 can be verified, for example, by checking whether the hash of the software matches the expected value signed by a trusted party, such as a silicon supplier manufacturing integrated circuits including data processing system 2 or an architecture provider designing processor architectures that support domain-based memory access control. This allows the user-provided code to verify the integrity of the domain-based architecture before executing any security or sensitive functions.

[0077] Thus, it can be seen that the code associated with domains 42 and 44 (which was previously executed in non-secure domain 86, as shown by the dashed lines indicating gaps in the non-secure domain where these processes were previously executed) can now be moved to the domain domain, where they can have stronger security guarantees because their data and code will not be accessed by other code operating in non-secure domain 86. However, the fact that domain domain 88 and secure domain 84 are orthogonal and therefore cannot see each other's physical address spaces means that the provider of the code in the domain domain does not need to trust the provider of the code in the secure domain, and vice versa. The code in the domain domain can simply trust the trusted firmware that provides the monitoring code 29 for root domain 82 and the RMM 46 provided by the silicon provider or the provider of the instruction set architecture supported by the processor (which may already be inherently needed to be trusted when the code executes on its device), so that a secure computing environment can be provided to users without further trust relationships with other operating system vendors, OEMs, or cloud hosts.

[0078] This can be used in a range of applications and use cases, including, for example, mobile wallets and payment applications, game anti-cheating and anti-piracy mechanisms, operating system platform security enhancements, secure virtual machine hosting, confidential computing, and gateway processing for networked or IoT devices. It should be understood that users can find many other applications with useful support for this domain.

[0079] To support security assurances provided to the domain, the processing system may support a proof reporting function, in which firmware images and configurations (e.g., monitoring code images and configurations or RMM code images and configurations) are measured at boot time or runtime, and domain content and configurations are measured at runtime, so that the domain owner can trace back the relevant proof reports to known implementations and certifications, thereby making a trust decision on whether to operate on the system.

[0080] like Figure 2As shown, a separate root domain 82 is provided for managing domain switching, and this root domain has its own isolated root physical address space. Even for systems with only non-secure domains 86 and secure domains 84 but no domain 88, the creation of the root domain and the isolation of its resources from the secure domains allows for a more robust implementation, but it can also be used for implementations that do not support domain 88. Root domain 82 can be implemented using monitoring code 29 provided (or certified) by the silicon provider or architect, and can be used to provide secure boot functionality, trusted boot measurement, on-chip system configuration, debug control, and management of firmware updates for firmware components provided by other parties (such as OEMs). Root domain code can be developed, certified, and deployed by the silicon provider or architect without relying on the final device. In contrast, secure domain 84 can be managed by the OEM to implement certain platform and security services. The management of non-secure domain 86 can be controlled by operating system 32 to provide operating system services, while domain 88 allows for the development of new forms of trusted execution environments that can be dedicated to user or third-party applications while being isolated from the existing security software environment in secure domain 84.

[0081] Figure 3 Another example of a data processing system 2 used to support these technologies is schematically illustrated. (The same reference numerals are used to indicate...) Figure 1 The same components. Figure 3 The address translation circuit 16 is shown in more detail, comprising a Stage 1 memory management unit 50 and a Stage 2 memory management unit 52. The Stage 1 MMU 50 is responsible for translating virtual addresses to physical addresses (when the translation is triggered by EL2 or EL3 code) or to intermediate addresses (when the translation is triggered by EL0 or EL1 code in a certain operating state, in which further Stage 2 translation by the Stage 2 MMU 52 is required). The Stage 2 MMU can translate intermediate addresses to physical addresses. The Stage 1 MMU can be based on a page table controlled by the operating system for translations initiated from EL0 or EL1, a page table controlled by the hypervisor for translations from EL2, or a page table controlled by the monitoring code 29 for translations from EL3. On the other hand, the Stage 2 MMU 52 can be based on a page table structure defined by the hypervisor 34, RMM46, or a security partition manager, depending on which domain is used. Dividing these translations into two phases in this way allows the operating system to manage address translation for itself and applications (assuming they are the only operating system running on the system), while RMM 46, Hypervisor 34, or SPM40 can manage isolation between different operating systems running in the same domain.

[0082] like Figure 3As shown, the address translation process using address translation circuit 16 returns a security attribute 54, which, combined with the current exception level 15 and current domain 14 (or security state), allows access to a specific physical address space (identified by the PAS identifier or "PAS TAG") in response to a given memory access request. The physical address and PAS identifier can be found in the granular protection table 56, which provides the previously described granular protection information. In this example, PAS filter 20 is shown as a granular memory protection unit (GMPU) that verifies whether the physical address of the selected PAS access request is allowed; if so, the transaction is allowed to proceed to any cache 24 or interconnect 8 that is part of the system architecture of the memory system.

[0083] The GMPU 20 allows memory to be allocated to separate address spaces while providing strong hardware-based isolation guarantees and offering spatial and temporal flexibility as well as efficient sharing schemes in terms of how physical memory is allocated to these address spaces. As previously described, execution units in the system are logically partitioned into virtual execution states (domains or "worlds"), with one execution state (root world) located at the highest exception level (EL3), which is called the "root world" and manages the allocation of physical memory to these worlds.

[0084] A single system physical address space is virtualized into multiple "logical" or "architectural" physical address spaces (PASs), each of which is an orthogonal address space with independent consistency properties. The system physical address is mapped to a single "logical" physical address space by extending the system physical address with a PAS label.

[0085] Allowing a subset of the logical-physical address space to be accessed in a given world. This is implemented by hardware filter 20, which can be attached to the output of address translation circuit 16.

[0086] The world uses fields in the translation table descriptor of the page table used for address translation to define the security attributes (PAS tags) for that access. Hardware filter 20 has access rights to a table (granular protection table 56 or GPT) that defines granular protection information (GPI) for each page in the system's physical address space, which indicates the associated PAS tag and (optionally) other granular protection attributes.

[0087] Hardware filter 20 checks the world ID and security attributes against the particle's GPI and determines whether access can be granted, thereby forming a particle-shaped memory protection unit (GMPU).

[0088] For example, GPT 56 can reside in on-chip SRAM or off-chip DRAM. If stored off-chip, GPT 56 can be protected for integrity by an on-chip memory protection engine that uses encryption, integrity, and freshness mechanisms to maintain the security of GPT 56.

[0089] Positioning the GMPU 20 on the requester side of the system (e.g., on the MMU output) instead of the completer side allows access permissions to be assigned at the page level, while allowing the interconnect 8 to continue hashing / stripping the page across multiple DRAM ports.

[0090] Transactions remain tagged with the PAS tag because they propagate throughout the system architecture until they reach a location defined as physical alias point 60. This allows filters to be positioned on the master side without compromising security guarantees compared to slave-side filtering. Because the transaction propagates throughout the system, the PAS tag can be used as a security-in-depth mechanism for address isolation: for example, a cache can add a PAS tag to an address label in the cache, preventing accesses to the same PA with an incorrect PAS tag from hitting the cache and thus improving resistance to side-channel attacks. The PAS tag can also be used as a context selector for a protection engine attached to the memory controller, which encrypts data before writing it to external DRAM.

[0091] A Physical Alias ​​Point (PoPA) is the location in the system where the PAS tag is stripped and the address is translated from a logical physical address back to a system physical address. The PoPA can be located on the system's completer side below the cache, where physical DRAM is accessed using the cryptographic context resolved via the PASTAG. Alternatively, it can be located above the cache to simplify system implementation at the cost of reduced security.

[0092] At any point in time, the world can request a page transition from one PAS to another. This request is made to monitoring code 29 at EL3, which checks the current state of the GPI. EL3 may allow only specific groups of transitions to occur (e.g., from a non-secure PAS to a secure PAS, but not from a domain PAS to a secure PAS). To provide a clean transition, the system supports a new instruction – “Data Cleanup and Invalidate to Physical Alias ​​Point” – which EL3 can submit before transitioning a page to a new PAS. This ensures that any residual state associated with the previous PAS is flushed from any cache upstream of PoPA 60 (closer to the requester side than PoPA 60).

[0093] Another feature achievable by attaching the GMPU 20 to the host side is efficient memory sharing between worlds. It might be desirable to grant shared access to a physical particle to a subset of N worlds, while preventing other worlds from accessing that physical particle. This can be achieved by adding "restricted sharing" semantics to the particle's protection information, while forcing it to use a specific PAS tag. As an example, the GPI could indicate that a physical particle can only be accessed by "Domain World" 88 and "Secure World" 84, while being tagged with the PAS tag of "Secure PAS" 84.

[0094] An example of the aforementioned characteristics is the ability to rapidly change the visibility properties of specific physical particles. Consider the case where each world is allocated a dedicated PAS that is only accessible to that world. For a specific particle, that world can request to make it visible to the insecure world at any point in time by changing its GPI from "exclusive" to "restricted sharing with the insecure world," without altering the PAS association. This increases the particle's visibility without requiring costly cache maintenance or data copying operations.

[0095] Figure 4 The concept of aliases on physical memory provided to the hardware by the corresponding physical address space is illustrated. As previously described, each of the domains 82, 84, 86, and 88 has its own corresponding physical address space 61.

[0096] At the point when the address translation circuit 16 generates the physical address, the physical address has a value within a certain range 62 supported by the system, which is the same regardless of which physical address space is selected. However, in addition to generating the physical address, the address translation circuit 16 can also select a specific physical address space (PAS) based on the current domain 14 and / or information in the page table entries used to derive the physical address. Alternatively, instead of the address translation circuit 16 performing the PAS selection, an address translation circuit (e.g., MMU) can output the physical address and information derived from the page table entries (PTE) for selecting the PAS, which the PAS filter or GMPU 20 can then use to select the PAS.

[0097] The selection of PAS for a given memory access request can be restricted according to the rules defined in the table below, based on the current domain that the processing circuit 10 is operating on when issuing the memory access request:

[0098]

[0099] For those domains where multiple physical address spaces are available, information from page table entries used to provide access to physical addresses is used to select among the available PAS options.

[0100] Thus, at the point in time when the PAS filter 20 outputs the memory access request to the system architecture (assuming it passes any filtering checks), the memory access request is associated with the physical address (PA) and the selected physical address space (PAS).

[0101] From the perspective of memory system components (such as caches, interconnects, snooping filters, etc.) operating prior to the Physical Alias ​​Point (PoPA) 60, the corresponding physical address space 61 is viewed as a completely separate address range corresponding to different system locations within memory. This means that, from the perspective of the pre-PoPA memory system components, the address range identified by a memory access request is actually four times the size of the range 62 that can be output in address translation, because the PAS identifier is actually treated as an additional address bit next to the physical address itself, so that the same physical address PAx can be mapped to multiple alias physical addresses 63 in different physical address spaces 61 depending on which PAS is selected. These alias physical addresses 63 all actually correspond to the same memory system location implemented in the physical hardware, but the pre-PoPA memory system components treat the alias address 63 as a separate address. Thus, if any pre-PoPA cache or snooping filter exists to allocate entries for such addresses, the alias address 63 will be mapped to a different entry with separate cache hit / miss decisions and separate consistency management. This reduces the likelihood or effectiveness of an attacker using a cache or consistency side channel as a mechanism to probe operations in other domains.

[0102] The system may include more than one PoPA 60 (e.g., as discussed below). Figure 14 (As shown in the diagram). At each PoPA 60, the aliased physical address is shrunk to a single de-aliased address 65 in the system physical address space 64. The de-aliased address 65 is provided downstream of any subsequent PoPA component so that the system physical address space 64, which actually identifies the memory system location, is once again the same size as the range of physical addresses that can be output in the address translation performed on the requester side. For example, at PoPA 60, the PAS identifier can be stripped from these addresses, and for downstream components, these addresses can simply be identified using the physical address values ​​without specifying the PAS. Alternatively, for some cases where some completer-side filtering is expected for memory access requests, the PAS identifier may still be provided downstream of PoPA 60, but may not be interpreted as part of the address, so that the same physical address appearing in different physical address spaces 61 will be interpreted downstream of the PoPA as involving the same memory system location, but the supplied PAS identifier can still be used to perform any completer-side security checks.

[0103] Figure 5This illustrates how the granular protection table 56 can be used to divide the system physical address space 64 into blocks of access allocations within a specific architecture physical address space 61. The granular protection table (GPT) 56 defines which portions of the system physical address space 65 are allowed to be accessed from each architecture physical address space 61. For example, the GPT 56 may include multiple entries, each corresponding to a physical address granule of a certain size (e.g., 4K pages), and may be defined as the PAS allocated to that granule, which can be selected from non-secure domains, secure domains, domains, and root domains. By design, if a particular granule or group of granules is allocated to a PAS associated with one of these domains, it can only be accessed within the PAS associated with that domain and not within PASs of other domains. However, it should be noted that while granules allocated to secure PASs (for example) cannot be accessed from within the root PAS, the root domain 82 can access the physical address granule by specifying PAS selection information in its page table to ensure that the virtual address associated with a page in that region of physically addressed memory is translated into a physical address in the secure PAS (not the root PAS). Thus, the sharing of data across domains can be controlled at the point in time when a PAS is selected for a given memory access request (to the extent permitted by the accessibility / inaccessibility rules defined in the previously described table).

[0104] However, in some specific implementations, in addition to allowing access to physical address particles within the allocated PAS defined by the GPT, the GPT can also use other GPT attributes to mark certain regions of the address space as shared with another address space (e.g., an address space associated with a domain of lower or orthogonal privileges (which typically does not allow it to select an allocated PAS for access requests to that domain)). This facilitates temporary sharing of data without requiring changes to the allocated PAS of a given particle. For example, in Figure 5 In the GPT, region 70 of the domain PAS is defined as being allocated to the domain domain and therefore is normally inaccessible from the non-secure domain 86 because the non-secure domain 86 cannot select the domain PAS for its access requests. Since the non-secure domain 26 cannot access the domain PAS, non-secure code typically cannot see the data in region 70. However, if a domain temporarily wishes to share some data in its allocated region of memory with the non-secure domain, it can request the monitoring code 29 operating in the root domain 82 to update GPT 56 to indicate that region 70 will be shared with the non-secure domain 86, and this allows region 70 to also be accessed from the non-secure domain 86. Figure 5The non-secure PAS access shown on the left-hand side does not require changing which domain is the domain allocated to region 70. If a domain has designated a region of its address space to be shared with a non-secure domain, then although a memory access request originating from the non-secure domain targeting that region may initially specify a non-secure PAS, PAS filter 20 can remap the PAS identifier of the request to instead specify a domain PAS, so that downstream memory system components treat the request as always originating from the domain domain. This sharing can improve performance because the operations used to allocate different domains to specific memory regions can be more performance-intensive, involving a greater degree of cache / TLB invalidation and / or data zeroing in memory or data copying between memory regions (which can be undesirable if the sharing is expected to be only temporary).

[0105] Figure 6 This is a flowchart illustrating how the current operating domain is determined, which can be performed by processing circuitry 10 or by address translation circuitry 16 or PAS filter 20. At step 100, it is determined whether the current exception level 15 is EL3. If so, at step 102, the current domain is determined to be root domain 82. If the current exception level is not EL3, at step 104, the current domain is determined to be one of non-secure domain 86, secure domain 84, and neighborhood domain 88, as indicated by at least two domain indicator bits in the processor's EL3 control register (since the root domain is indicated by the current exception level as EL3, it may not be necessary to have an encoding for the domain indicator bits corresponding to the root domain; therefore, at least one encoding of the domain indicator bits can be reserved for other purposes). The EL3 control register can be written to when operating at EL3 and cannot be written to from other exception levels EL2-EL0.

[0106] Figure 7An example of a page table entry (PTE) format is shown that can be used for page table entries in a page table structure. These page table entries are used by address translation circuitry 16 to map virtual addresses to physical addresses, virtual addresses to intermediate addresses, or intermediate addresses to physical addresses (depending on whether the translation is performed in an operational state where a stage 2 translation is fully required, and whether the translation is a stage 1 translation or a stage 2 translation if a stage 2 translation is required). In general, a given page table structure can be defined as a multi-level table structure implemented as a page table tree, where the first level of the page table is identified based on a base address stored in the processor's translation table base address register, and an index for selecting a specific level 1 page table entry within the page table is derived from a subset of bits of the input address to which a translation lookup is performed (the input address can be a virtual address for stage 1 translation or an intermediate address for stage 2 translation). A level 1 page table entry can be a "table descriptor" 110 that provides a pointer 112 to the next level page table, from which further page table entries can be selected based on a further subset of bits of the input address. Finally, after one or more lookups at consecutive levels of the page table, block or page descriptors PTE 114, 116, and 118 can be identified, which provide an output address 120 corresponding to the input address. The output address can be an intermediate address (for a stage 1 transition performed while the operation is still in the state of performing a further stage 2 transition) or a physical address (for a stage 2 transition or a stage 1 transition when stage 2 is not required).

[0107] To support the different physical address spaces mentioned above, in addition to the next-level page table pointer 112 or output address 120 and any attributes 122 used to control access to the corresponding block of memory, the page table entry format also specifies some additional states for physical address space selection.

[0108] For page table descriptor 110, the PTE used by any domain other than the non-secure domain 86 includes a non-secure table indicator 124, which indicates whether the next-level page table will be accessed from the non-secure physical address space or from the physical address space of the current domain. This facilitates more efficient page table management. Typically, the page table structure used by the root domain, domain domain, or secure domain may only need to define special page table entries for a portion of the virtual address space, and for other portions, the same page table entries used by the non-secure domain 26 can be used. Therefore, by providing the non-secure table indicator 124, this allows higher levels of the page table structure to provide dedicated domain / secure table descriptors, while at certain points in the page table tree, the root domain domain or secure domain can switch to using page table entries from the non-secure domain for portions of the address space that do not require higher security. Other page table descriptors in other parts of the page table tree can still be obtained from the associated physical address space associated with the root domain, domain domain, or secure domain.

[0109] On the other hand, block / page descriptors 114, 116, and 118 may include physical address space selection information 126 depending on which domain they are associated with. The insecure block / page descriptor 118 used in the insecure domain 86 does not include any PAS selection information because the insecure domain can only access insecure PASs. However, for other domains, block / page descriptors 114 and 116 include PAS selection information 126, which is used to select which PAS to translate the input address to. For the root domain 82, EL3 page table entries may have PAS selection information 126, which includes at least two bits to indicate the selected PAS to which the corresponding physical address will be translated when the PAS associated with any of the four domains 82, 84, 86, and 88 is associated. In contrast, for both the domain and security domains, the corresponding block / page descriptor 116 only needs to include one bit of PAS selection information 126, which selects between the domain PAS and non-security PAS for the domain domain, and between the security PAS and non-security PAS for the security domain. To improve circuit implementation efficiency and avoid increasing the size of page table entries, the block / page descriptor 116 can encode the PAS selection information 126 at the same location within the PTE for both the domain and security domains, regardless of whether the current domain is the domain or the security domain, so that the PAS selection information 126 can be shared.

[0110] thereby, Figure 8 This is a flowchart illustrating a method for selecting a PAS based on the current domain and information from the block / page PTE used to generate a physical address for a given memory access request (non-security table indicator 124, PAS selection information 126). PAS selection can be performed by address translation circuit 16, or by a combination of address translation circuit 16 and PAS filter 20 if the address translation circuit forwards PAS selection information 126 to PAS filter 20.

[0111] exist Figure 8At step 130, processing circuit 10 issues a memory access request specifying a given virtual address (VA) as the target VA. At step 132, address translation circuit 16 searches its TLB 18 for any page table entries (or cached information derived from such page table entries). If any required page table information is unavailable, address translation circuit 16 initiates a page table roaming to memory to obtain the required PTE (potentially requiring a series of memory accesses to progressively traverse the corresponding levels of the page table structure, and / or potentially requiring multiple stages of address translation to obtain a mapping from VA to intermediate address (IPA) and then from IPA to PA). It should be noted that any memory access request issued by address translation circuit 16 during the page table roaming operation may itself undergo address translation and PAS filtering, so the request received at step 130 may be a memory access request issued to request a page table entry from memory. Once the relevant page table information has been identified, the virtual address is translated into a physical address (possibly via IPA in two stages). At step 134, address translation circuit 16 or PAS filter 20 uses Figure 6 The method shown determines which domain is the current domain.

[0112] If the current domain is a non-secure domain, then at step 136, the output PAS selected for the memory access request is a non-secure PAS.

[0113] If the current domain is a secure domain, then at step 138, an output PAS is selected based on PAS selection information 126, which provides the physical address and is included in the block / page descriptor PTE, wherein the output PAS will be selected as a secure PAS or a non-secure PAS.

[0114] If the current domain is a domain domain, then at step 140, an output PAS is selected based on PAS selection information 126 included in the block / page descriptor PTE from its derived physical address, and in this case, the output PAS is selected as a domain PAS or a non-secure PAS.

[0115] If the current domain is determined to be the root domain at step 134, then at step 142, the output PAS is selected based on the PAS selection information 126 from which the physical address is derived in the root block / page descriptor PTE 114. In this case, the output PAS is selected as any physical address space in the physical address space associated with the root domain, the domain, the security domain, and the non-security domain.

[0116] Figure 9An example of an entry for GPT 56 for a given physical address particle is shown. GPT entry 150 includes an assigned PAS identifier 152 that identifies the PAS allocated to the physical address particle, and optionally includes further attributes 154, which may include, for example, the previously described shared attribute information 156 that makes the physical address particle visible in one or more other PASs besides the assigned PAS. The root domain may perform the setting of the shared attribute information 156 when code running in the domain associated with the assigned PAS issues a request. Additionally, the attributes may include a pass-through indicator 158, which indicates whether a GPT check (for determining whether the selected PAS of the memory access request is allowed to access the physical address particle) should be performed by the PAS filter 20 on the requester side or by the completer-side filtering circuitry on the completer device side of the interconnect, as will be discussed further below. If the pass-through indicator 158 has a first value, a requester-side filtering check may be required at the PAS filter 20 on the requester side, and if these checks fail, the memory access request can be blocked and a fault signal can be sent. However, if the pass-through indicator 158 has a second value, a requester-side filtering check based on GPT 56 may not be required for a memory access request specifying a physical address in the physical address particle corresponding to the GPT entry 150, and in this case, the memory access request can be passed all the way to cache 24 or interconnect 8 without checking whether the selected PAS is one of the allowed PAS that allows access to the physical address particle, and any such PAS filtering check is performed later on the completer side.

[0117] Figure 10 This is a flowchart illustrating a requester-side PAS filter check performed by PAS filter 20 at the requester side of interconnect 8. At step 170, PAS filter 20 receives a memory access request, which includes a physical address and other parameters as previously described. Figure 8 The output PAS selected as shown is associated with this. As further mentioned below, while waiting for subsequent steps 172-180 to be executed, it is possible to issue a pre-PoPA request for read-on-hit, which can help to access the data faster if the GPT check will succeed.

[0118] At step 172, the PAS filter 20 either obtains the GPT entry corresponding to the specified PA from the particle protection information cache 22 (if available), or obtains the GPT entry by issuing a request to the memory to obtain the required GPT entry from a table structure stored in memory. Once the required GPT entry has been obtained, at step 174, the PAS filter determines whether the output PAS selected for the memory access request is the same as the allocated PAS 152 defined in the GPT entry obtained at step 172. If so, at step 176, the memory access request (specifying the PA and the output PAS) can be allowed to pass to the cache 24 or the interconnect 8.

[0119] If the output PAS is not an assigned PAS, then at step 178, the PAS filter determines whether the output PAS is indicated in the shared attribute information 156 from the acquired GPT entry as a permitted PAS allowing access to the address granular corresponding to the specified PA. If so, then again at step 176, the memory access request is allowed to be passed to cache 24 or interconnect 8. The shared attribute information may be encoded as unique bits (or sets of bits) within GPT entry 150, or may be encoded as one or more codes of a field of GPT entry 150, in which case other codes of the same field may indicate other information. If step 178 determines that the shared attribute indicates that an output PAS other than an assigned PAS is allowed to access the PA, then at step 176, the PAS specified in the memory access request passed to cache 24 or interconnect 8 is an assigned PAS, not an output PAS. The PAS filter 20 transforms the PAS specified by the memory access request to match the assigned PAS, so that downstream memory system components treat it as the same as the request for the assigned PAS.

[0120] If the output PAS does not indicate permission to access the specified physical address in the shared attribute information 156 (or alternatively, in implementations that do not support shared attribute information 156, step 178 is skipped), then at step 180, it is determined whether the pass-through indicator 158 in the obtained GPT entry for the target physical address recognizes that the memory access request can be passed all the way to cache 24 or interconnect 8, regardless of the check performed at the requester-side PAS filter 20, and if a pass-through indicator is specified, then at step 176, the memory access request is again permitted (the output PAS is specified as the PAS associated with the memory access request). Alternatively, if any check at steps 174, 178, and 180 does not recognize permission to access the memory access request, then at step 182, the memory access request is blocked (for requests other than the type of pre-PoPA request read on hit described below, a pre-PoPA request read on hit is allowed without waiting for a successful GPT check). Therefore, memory access requests other than the pre-PoPA request read upon hit are not passed to cache 24 or interconnect 8, and faults can be signaled, which can trigger exception handling to deal with the fault.

[0121] Although steps 174, 178, and 180 are in Figure 10 These steps are shown sequentially, but may be implemented in parallel or in a different order if needed. It should also be understood that steps 178 and 180 are not necessary and some specific implementations may not support the use of shared attribute information 156 and / or pass-through indicator 158.

[0122] Figure 11 The operation of the address translation circuit 16 and the PAS filter is summarized. The PAS filter 20 can be viewed as an additional stage 3 check performed after the stage 1 (and optionally stage 2) address translation performed by the address translation circuit. It should also be noted that the EL3 translation is based on page table entries, which provide two bits of address-based selection information (in... Figure 11 In the example, they are labeled NS and NSE, while the selection information "NS" is used to select PAS in other states. Figure 11 The safety status indicated by the input for particle protection check refers to the domain ID of the current domain of the identification processing element.

[0123] Reading the previous PoPA request upon hit

[0124] The particle protection check performed by the PAS filter (protection check circuit) 20 can be relatively slow because it may require protection table roaming to retrieve protection information from memory. If a memory access request for a given target PA and target PAS cannot be initiated until after the protection check is complete, this can introduce additional latency for all accesses to memory. Figure 12 and Figure 13 The example illustrates the processing of a pre-PoPA request for a hit-on-hit read, which can be supported by the memory interface protocol used by the requester device 4, the interconnect 8, and the completer device 6, so that the latency can be reduced.

[0125] A pre-PoPA request read on hit causes one or more memory system components to provide a pre-PoPA response read on hit, which involves different results depending on whether the target PA and target PAS of the request are hit or miss in any pre-PoPA cache 24 located upstream of PoPA 60 (closer to the requester device 4 than PoPA 60).

[0126] Figure 12 This illustrates the processing of a pre-PoPA request read at the time of a cache hit when a cache hit is detected in the pre-PoPA cache 24. This cache may reside within the requester device 4 (e.g., ...). Figure 1 (As shown), but it can also be located in another part of the memory system upstream of PoPA 60, such as within interconnect 8. For simplicity, Figure 12 This pertains to a single hit in the pre-PoPA cache, but it should be understood that multiple levels of the pre-PoPA cache may exist and the hit may occur at any of these pre-PoPA cache levels.

[0127] like Figure 12 As shown, when the address translation circuit 16 of the given requester device 4 prepares the target PA for the required read memory access, the requester circuit issues a hit-read pre-PoPA request 190. This hit-read pre-PoPA request specifies the target PA and target PAS obtained by the address translation circuit 16 or the PAS filter 20 (whichever acts as the PAS selection circuit), without waiting for the PAS filter 20 (protection check circuit) to complete the aforementioned granular protection table check (GPT check) for the target PA and target PAS. The hit-read pre-PoPA request 190 is sent to at least one pre-PoPA memory system component, such as cache 24 or interconnect 8, which checks whether a cache hit is detected in the pre-PoPA cache, which treats alias PAs corresponding to the same memory system resources in different PASs as if they were all actually separate addresses. For example, a cache lookup in the previous PoPA cache 24 can be based on an address comparison of the target PA of the request against the address label of the corresponding cache entry, and a PAS identifier comparison between the target PAS identifier specified in the request and the corresponding PAS identifier label associated with the corresponding cache entry.

[0128] If a cache hit is detected in the pre-PoPA cache 24 in response to a pre-PoPA read request 190, a data response 191 is returned to the requester, providing the data read from the hit cache entry in the pre-PoPA cache 24. In this case, the requester device 4 can use the returned data to process subsequent instructions without waiting for the GPT check to complete; for example, the read data can be returned and stored in register 12 of the requester device 4. Once the GPT check is complete and... Figure 12 Success was confirmed at time point 192, so no further read requests were needed because the required data had already been returned. Although Figure 12 Not shown, but another option would be to stop any incomplete parts of the GPT check once the data response 191 has been received after the previous PoPA request 190 read in response to a hit has been read. For example, if there are still some granular protection table roaming access requests to be issued at the time the data response 191 is received, any remaining granular protection table roaming requests can be suppressed (alternatively, these roaming requests can still be issued to allow the corresponding protection information to be cached in the granular protection information cache 22).

[0129] In addition, although Figure 12 Not shown, but if multiple levels of pre-PoPA cache exist and the request misses in level 1 cache, some data transfer may occur between pre-PoPA cache levels to promote data that hits in level 2 or subsequent levels of caches prior to PoPA to higher cache levels, and, if necessary, evict data from higher cache levels to make room for the promoted data. Line-fill requests issued from higher-level pre-PoPA caches to lower-level pre-PoPA caches may also be issued as pre-PoPA requests for hit-on-hit reads to ensure that if a line-fill request misses in the final pre-PoPA cache (the ultimate cache prior to PoPA), data is not carried from memory system components after PoPA 60 to the pre-PoPA cache level.

[0130] Figure 13The processing of a prePoPA request 190 read at a hit is illustrated in the event that a miss is detected in any level of the provided prePoPA cache 24 by the prePoPA memory system components. When a cache miss is detected in the prePoPA cache, or if multiple levels of the prePoPA cache exist, when a miss is detected in all of these prePoPA caches accessible to the requester device 4 (which excludes any dedicated caches of other requesters that are not accessible to the requester device 4 that issued the prePoPA request read at the time of the initial hit), a no-data response 193 is returned to the requester device 4, indicating that the requested PA and PAS data cannot yet be safely returned.

[0131] Optionally, in addition to returning a no-data response 193, when a front-PoPA cache miss is detected, the front-PoPA memory system component may also issue a request 194 to the rear-PoPA memory system component (such as a memory controller, peripheral controller, or another element) to request a preparatory action to prepare for a later read request for the same PA. This request 194 for the preparatory action may trigger the rear-PoPA memory system to perform preparatory operations (such as prefetching data of the target PA into the rear-PoPA cache 7, and / or performing activation or precharge operations in memory 9), which makes it more likely that subsequent read requests for the same PA will be processed faster than in the case where no preparatory action has been performed. However, the request 194 for the preparatory action is not allowed to cause data of the target PA to be promoted from the rear-PoPA location to one of the front-PoPA caches 24.

[0132] When the requester device 4 is Figure 13 When the GPT check of the target PA is successful at time 192, the requester device 4 then issues a standard read request 195 specifying the same target PA and target PAS as the previous PoPA request 190 read when it was hit. This request will again miss in the previous PoPA cache 24 and may subsequently trigger a row fill request for the post-PoPA memory system components, which can return the data. Therefore, the requester device 4 then receives a data response 191 (as in...). Figure 12 The type of data response 191 received in the case of a cache hit is the same as that shown. In addition to returning the data to the requester, the read request 195 may also cause the data of the target PA to be allocated to at least one pre-PoPA cache 24, because the operation is now safe since the GPT check has been determined to be successful. Figure 13The read request 195 issued in the process can be a type of request that is not allowed to be issued until the GPT check for the specified target PA and target PAS has been verified as successful. Since the preparatory action request 194 has already been issued while processing the read request 190 on hit, the processing speed of read request 195 can be faster than in the case where no request is issued until the GPT check is determined to be successful, thereby speeding up the processing of read requests and thus improving performance.

[0133] However, other implementations may omit the request 194 for the preparation action, and in this case, performance can still be improved by supporting the pre-PoPA request for read-on-hit, since at least in cases such as Figure 12 When the request shown hits the pre-PoPA cache, the data can be returned earlier than in the case of a pre-PoPA request 190 that reads without any hit, and all accesses to memory must wait until the GPT check is successful before the read request 195 is issued.

[0134] In summary, in the context of a system with a cache marked with domain information before a Physical Alias ​​Point (PoPA) and a memory controller (and potentially further caches) after the PoPA, it is desirable to prefetch memory locations before fully performing granular protection checks in order to hide the cost of additional protection phases on TLB roaming. However, allocating such entries to the cache, where granular protection checks have been omitted, can undermine the security guarantees provided by the protection check circuitry 20. A pre-PoPA request read on hit provides a type of communication interface / coherence protocol request that acts as a cacheable read when it hits in the cache before the PoPA and returns a no-data response or acts as a no-data prefetch request when it misses before the PoPA. It is recognized that if the line is already in a cache before the PoPA with the same domain information as the new request, it is safe to move the line to another domain-marked cache closer to the CPU or return the line to the requester. However, it is not safe to bring the line from after the PoPA, as this could result in the incorrect domain information being marked. Latency can be minimized without compromising any inherent security guarantees provided by the protection check when using new requests with semantics roughly between cacheable reads and memory controller / cached prefetch requests. The requester can send this new request type immediately after completing the transition but before completing the granular protection check. The request is expected to look up the cache hierarchy prior to the PoPA, and if it hits a line with the same domain information as the new request, it is expected to return the data to the cache closer to the requester, where it can then be safely allocated. However, if the lookup misses in the cache prior to the PoPA, the request can be transformed into a prefetch request or other request for a prefetch action, which prefetches the line into the nearest cache after the PoPA or performs a precharge or activation operation without returning any data to any cache prior to the PoPA. The requester will receive a no-data response, indicating that no data is safe to return. Once the requester completes the particle protection check, a normal read transaction can be issued, which will then be able to retrieve the data with reduced latency because it is expected to hit in a cache after the PoPA (e.g., the memory controller's prefetch buffer) or not cause latency associated with precharge or activation operations. Alternatively, additional buffers / caches can be provided closer to the PoPA.

[0135] Figure 14This is a flowchart illustrating the operations performed by the requester device 4 when initiating a read access to memory. At step 200, the address translation circuit 16 of the requester device 4 performs address translation to convert the target VA of the access to the target PA. Additionally, the PAS selection circuit (address translation circuit 16 or PAS filter 20) selects the target PAS to be used for the access based on information from the current domain 14 currently performing the processing operation, or information from the address translation entry corresponding to the target VA, or both information from the current domain and the address translation entry.

[0136] At step 202, the PAS filter 20 determines whether the protection information for the specified target PA is already available in the protection information cache 22 (which may be a dedicated cache used only for caching GPT entries or may be combined with TLB 18 as previously discussed). If the protection information is already available in the protection information cache 22, a protection check is performed using the cached protection information at step 204, and it is determined at step 206 whether the protection check was successful. If so, at step 208, a read request 195 is issued by the requester circuit to the downstream memory system components, wherein the read request specifies the target PA obtained in address translation and the target PAS obtained in PAS selection. At step 210, the requester circuit receives a read data response providing the requested data. The read data may be stored in register 12 and may be used as an operand for subsequent instructions. On the other hand, if the protection check fails at step 206 and the PAS filter 20 determines that access to the target PA within the target PAS is not permitted (i.e., the target PAS is not within the allowed address space of the target PA), then at step 212, the PAS filter 20 signals a fault and blocks the read request 195 from being sent to downstream memory system components. This fault may cause an exception handler executing on the processing system to determine how to handle the fault, such as causing a system reset or causing a software routine to investigate why the request was made to the wrong address space of the given PA, disabling the execution of the thread attempting to access the wrong address space, or taking another form of response action (the specific response taken may depend on the specific requirements of the application running or the platform operator's choice).

[0137] On the other hand, if it is determined at step 202 that the protection information for the target PA is not yet available in the protection information cache 22 (because the protection information cache 22 is not provided at all, or because the request for the protection information for the target PA is a miss in the protection information cache 22), then at step 220, the requester device 4 issues a pre-PoPA request 190 for read-on-hit at the downstream memory system component, wherein the request specifies the target PA and target PAS obtained at step 200. The pre-PoPA request for read-on-hit at the hit is issued without waiting for the PAS filter 20 to perform a protection check. At step 222, a response to the pre-PoPA request for read-on-hit at the hit is received from the downstream memory system component. At step 224, the type of the response is determined by the requester device 4. If the type of the received response is a data response, there is no need to issue a subsequent read request for the target PA and target PAS, and therefore such a read request can be suppressed at step 226. If the received response is a no-data response 193, the requester device 4 waits for the result of the protection check.

[0138] Simultaneously, in parallel with the issuance of the pre-PoPA request read upon hit, if the protection information of the target PA is unavailable in the protection information cache 22, the requester device 4 also issues one or more protection table roaming access requests at step 230. These one or more protection table roaming access requests are access requests to obtain the protection information of the target PA from memory. Such protection table roaming access requests specify a PA different from the target PA used in the pre-PoPA request read upon hit. The address used for the protection table roaming access request may depend on the table base address maintained by the PAS filter 20 and a portion of the target PA in the pre-PoPA request read upon hit. The number of protection table roaming accesses required can vary depending on the extent to which relevant information exists in the protection information cache, how the protection table is defined, and how many levels of the protection table need to be traversed to find the protection information of the target PA. For example, the protection table can be a multi-level table, and different parts of the target PA can be indexed into each level of the table according to a base address maintained by the PAS filter 20, which identifies the base address of the first level of the protection table and then obtains the base address of the subsequent levels of the protection table based on the contents of the table entries read at the previous levels of the table. Different PAs may require different numbers of levels of traversal of the protection table.

[0139] Finally, protection information corresponding to the target PA is received at step 232, and at step 234, the PAS filter 20 uses the received protection information to perform a protection check. At this time, the PAS filter 20 may also allocate entries to the granular protection information cache 22 (if provided) associated with the protection information of the target PA, so that future accesses to the same address can be performed more quickly. The protection check determines whether access to the target PAS selected at step 200 is permitted to access the target PA (i.e., determines whether the target PAS is a permitted PAS of the target PA). At step 236, the PAS filter 20 determines whether the protection check was successful (the target PAS is a permitted PAS of the target PA), and if so, at step 240, the requester device 4 issues a read request specifying the target PA and the target PAS. This read request is... Figure 13 Type 195, as shown, will trigger a data response regardless of whether it is a miss or a hit in the pre-PoPA cache 24. This will occur if the protection check result has been determined to be successful before any response is received from the memory system in response to a pre-PoPA request read on a hit, or (more likely) if no data response is received in response to a pre-PoPA request read on a hit (e.g., ...). Figure 14 (As shown by the arrows from step 224 to step 236), then step 240 is executed. On the other hand, if the protection check is determined to be unsuccessful at step 236, then a fault is signaled again at step 212 and the issuance of read requests for the specified target PA and target PAS is blocked.

[0140] Figure 15 An example of a pre-PoPA memory system component 300 is shown in more detail. This pre-PoPA memory system component may be, for example, control circuitry logic associated with the pre-PoPA cache 24 or control logic within interconnect 8. The pre-PoPA memory system component 300 has a requester interface 302 for receiving requests from and providing responses to requester devices 4, and a completer interface 308 for providing requests to the post-PoPA memory system component and receiving responses to the requests in response. Control circuitry 304 processes incoming requests at requester interface 302, issues requests via completer interface 308, processes responses received at completer interface 308, and issues responses via requester interface 302. Additionally, control circuitry 304 controls the lookups of the pre-PoPA cache 24.

[0141] Figure 16A flowchart is shown illustrating the operations performed by the pre-PoPA memory system component 300. At step 310, the pre-PoPA memory system component receives a pre-PoPA request 190 upon receiving a hit specifying a target PA and a target PAS. At step 312, control circuitry 304 triggers a lookup of the target PA and target PAS in the pre-PoPA cache 24. If the lookup detects a cache hit in any pre-PoPA cache, resulting in a valid entry in the pre-PoPA cache 24 corresponding to both the target PA and target PAS, then at step 314, control circuitry 304 controls the requester interface 302 of the pre-PoPA memory system component 300 to return a data response 191, which provides the data cached in the hit entry of the pre-PoPA cache 24. Although Figure 16 As not shown in the diagram, it is also possible that at step 314, there may be movement of data between different levels of the pre-PoPA cache (e.g., when there is a cache miss in level 1 but a cache hit in level 2, data may be promoted from level 2 cache to level 1 cache, and this may cause data to be evicted from level 1 cache to level 2 cache, where both level 1 and level 2 caches are pre-PoPA caches).

[0142] On the other hand, if the lookup fails in the front PoPA cache (when there are no valid entries for the target PA and target PAS that correspond to cache entries with corresponding PA and PAS tags), then at step 316, a no-data response 191 is returned to the requester device 4 via the requester interface 302. At step 318, the control circuit 304 prevents any data for the target PA from being allocated to the front PoPA cache 24 based on any line-filling requests issued to the back PoPA memory system components. For example, this can be implemented by preventing the completer interface 308 from issuing any such line-filling requests to the back PoPA memory system components. Additionally, in a cache miss scenario, at step 320, the control circuit 304 controls the completer interface 308 to issue a request to the back PoPA memory system components to perform at least one preparatory action, thereby preparing for the processing of a later access request for the specified target PA. For example, the request may cause data for the target PA to be prefetched into the post-PoPA cache 7, or it may control the memory 9 to perform a precharge or activation operation to prepare a row of memory locations (including the location corresponding to the target PA) for future access, thereby allowing future requests specifying the same target PA to be processed with reduced latency. Step 320 is optional and may be omitted in some examples.

[0143] In this application, the phrase "configured as..." is used to mean that the elements of the device have a configuration capable of performing the defined operation. In this context, "configuration" means the arrangement or manner of interconnection of hardware or software. For example, the device may have dedicated hardware that provides the defined operation, or a processor or other processing device may be programmed to perform the function. "Configured as" does not mean that the elements of the device need to be changed in any way to provide the defined operation.

[0144] While exemplary embodiments of the invention have been described in detail herein with reference to the accompanying drawings, it should be understood that the invention is not limited to those precise embodiments, and various changes and modifications can be made therein by those skilled in the art without departing from the scope and spirit of the invention as defined in the appended claims.

Claims

1. A data processing apparatus, the apparatus comprising: A requester circuit is configured to issue an access request to the memory system, the access request specifying a target physical address and a target physical address space identifier, the target physical address space identifier identifying a target physical address space selected from a plurality of physical address spaces. as well as A memory system component, configured to respond to an access request issued by the requester circuit, the memory system component comprising: At least one pre-Physical Alias ​​Point (PoPA) memory system component, prior to PoPA, is configured to treat aliased physical addresses from different physical address spaces that actually correspond to the same memory system resource as aliased physical addresses corresponding to different memory system resources; and At least one post-PoPA memory system component, after the PoPA, is configured to treat the alias physical address as involving the same memory system resource; in: In response to a pre-PoPA request issued by the requester circuit specifying the target physical address and the target physical address space identifier, at least one memory system component of the memory system components is configured to provide a pre-PoPA response action for a hit-on-hit read, including: When a pre-PoPA request read upon hit hits at least one pre-PoPA cache prior to the PoPA, a data response is provided to the requester circuit to return cached data from the hit entry of the at least one pre-PoPA cache corresponding to the target physical address and the target physical address space identifier; and When the pre-PoPA request read on hit fails to hit in at least one pre-PoPA cache, a no-data response is provided to the requester circuit, the no-data response indicating that data for the target physical address will not be returned to the requester circuit in response to the pre-PoPA request read on hit.

2. The apparatus of claim 1, wherein the hit-time read pre-PoPA response action comprises: When a pre-PoPA request read on a hit misses in at least one pre-PoPA cache, data for the target physical address is prevented from being allocated to the at least one pre-PoPA cache based on a line-filling request issued to at least one of the at least one post-PoPA memory system components in response to the pre-PoPA request read on a hit.

3. The apparatus of claim 1 or 2, wherein the hit-time read pre-PoPA response action comprises: When a pre-PoPA request read upon hit fails to be hit in at least one pre-PoPA cache, a request is issued to at least one of the at least one post-PoPA memory system components to perform at least one preparation operation in order to prepare for processing of a later access request specifying the target physical address.

4. The apparatus of claim 3, wherein the at least one preparation operation comprises at least one of the following: The data associated with the target physical address is prefetched into the post-PoPA cache; Perform a precharge or activation operation to make the memory location associated with the target physical address ready for access.

5. The apparatus of claim 1, the apparatus further comprising a protection check circuit that performs a protection check to determine, based on a lookup of protection information corresponding to the target physical address, whether access to the target physical address for a given access request is permitted within the target physical address space, the protection information indicating at least one permitted physical address space from which access to the target physical address is permitted.

6. The apparatus of claim 5, wherein when the protection information corresponding to the target physical address is not yet available in the protection information cache, the protection check circuit is configured to issue one or more protection table roaming access requests to the memory system component to perform a protection table roaming, thereby obtaining the protection information from the protection table indexed by the target physical address.

7. The apparatus of claim 5 or 6, wherein the requester circuitry is configured to issue a pre-PoPA request for hit read prior to the completion of the protection check of the target physical address and the target physical address space.

8. The apparatus of claim 7, wherein responding to the protection check that access to the target physical address within the target physical address space is permitted upon receiving the no-data response in response to a pre-PoPA request read upon the hit, the requester circuitry is configured to issue a read request specifying the target physical address and the target physical address space; and In response to the read request, at least one of the memory system components is configured to provide a data response so as to return data associated with the target physical address to the requester circuit even if the read request misses in the at least one pre-PoPA cache.

9. The apparatus of claim 8, wherein in response to the read request, the memory system component is allowed to cause data associated with the target physical address to be allocated to the at least one pre-PoPA cache based on a line fill request issued to at least one of the at least one post-PoPA memory system component.

10. The apparatus of claim 8, wherein the requester circuitry is configured to prevent the issuance of the read request when the protection check determines that access to the target physical address within the target physical address space is not permitted.

11. The apparatus of claim 5 or 6, wherein the protection check circuit is configured to signal a fault when the protection check determines that access to the target physical address within the target physical address space is not permitted.

12. The apparatus according to claim 1 or 2, wherein: The at least one pre-PoPA cache is configured to detect a cache miss even if a valid entry in the at least one pre-PoPA cache corresponds to both the target physical address and the target physical address space identifier, when no valid entry in the at least one pre-PoPA cache corresponds to both the target physical address and the target physical address space identifier.

13. The apparatus according to claim 1 or 2, the apparatus comprising an address translation circuit, the address translation circuit obtaining the target physical address by converting the target virtual address based on address translation information corresponding to the target virtual address.

14. The apparatus of claim 13, further comprising a physical address space selection circuit that selects the target physical address space identifier corresponding to the target physical address based on at least one of the following: The current operating domain of the requester circuit; and The information specified in the address translation information corresponding to the target virtual address.

15. A data processing method, the method comprising: An access request is issued from the requester circuit to access the memory system, which includes memory system components. The access request specifies a target physical address and a target physical address space identifier, which identifies the target physical address space selected from a plurality of physical address spaces. as well as The memory system component responds to the access request issued by the requester circuit using the memory system component, the memory system component comprising: At least one pre-PoPA physical alias point memory system component, prior to PoPA, is configured to treat aliased physical addresses from different physical address spaces that actually correspond to the same memory system resource as aliased physical addresses corresponding to different memory system resources; and At least one post-PoPA memory system component, after the PoPA, is configured to treat the alias physical address as involving the same memory system resource; in: In response to a pre-PoPA request issued by the requester circuit specifying the target physical address and the target physical address space identifier, at least one memory system component in the memory system components provides a pre-PoPA response action for a hit-on-hit read, including: When a pre-PoPA request read upon hit hits at least one pre-PoPA cache prior to the PoPA, a data response is provided to the requester circuit to return cached data from the hit entry of the at least one pre-PoPA cache corresponding to the target physical address and the target physical address space identifier; and When the pre-PoPA request read on hit fails to hit in at least one pre-PoPA cache, a no-data response is provided to the requester circuit, the no-data response indicating that data for the target physical address will not be returned to the requester circuit in response to the pre-PoPA request read on hit.

16. A memory system component, the memory system component comprising: A requester interface is provided for receiving an access request from a requester circuit, the access request specifying a target physical address and a target physical address space identifier, the target physical address space identifier identifying a target physical address space selected from a plurality of physical address spaces. and A control circuit is configured to detect whether a hit or miss is detected in at least one pre-PoPA cache lookup prior to a Physical Alias ​​Point (PoPA), wherein the at least one pre-PoPA cache is configured to treat aliased physical addresses from different physical address spaces that actually correspond to the same memory system resource as aliased physical addresses corresponding to different memory system resources, and the PoPA is a point after which at least one post-PoPA memory system component is configured to treat the aliased physical address as relating to the same memory system resource; wherein: In response to a pre-PoPA read request issued by the requester circuit specifying the target physical address and the target physical address space identifier, the control circuit is configured to provide a pre-PoPA read-on-hit response action, including: When a lookup of the at least one pre-PoPA cache detects a hit, the requester interface is controlled to provide a data response to the requester circuitry, thereby returning the cached data in the hit entry of the at least one pre-PoPA cache corresponding to the target physical address and the target physical address space identifier; and When a lookup in at least one pre-PoPA cache detects a miss, the requester interface is controlled to provide a no-data response to the requester circuitry, the no-data response indicating that data for the target physical address will not be returned to the requester circuitry in response to a pre-PoPA request read upon a hit.

17. The memory system component of claim 16, wherein the hit-time read pre- PoPA response action comprises: When a lookup in the at least one pre-PoPA cache detects a miss, the allocation of data for the target physical address to the at least one pre-PoPA cache is prevented from being made to at least one of the at least one post-PoPA memory system components based on a pre-PoPA request read in response to a hit.

18. The memory system component of claim 16 or 17, wherein the hit-while- read pre-PoPA response action comprises: When a lookup in the at least one pre-PoPA cache detects a miss and the at least one PoPA memory system component includes at least one post-PoPA cache, a request is made to at least one of the at least one post-PoPA memory system components to perform at least one preparation operation in order to prepare for processing of a later access request specifying the target physical address.

19. The memory system component of claim 18, wherein the at least one preparation operation comprises at least one of: prefetching data associated with the target physical address into a post-PoPA cache; performing a pre-charge or activation operation to prepare a memory location associated with the target physical address for access.