Data Integrity Check of Granule Protection Data
A data integrity check mechanism for granule protection information in vulnerable memory areas addresses security and performance challenges by using lightweight encryption and integrity checks, ensuring secure and efficient access control in data processing systems.
Patent Information
- Application Number
- JP2022559849
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-04-14
- Filing Date
- 2021-04-09
- Publication Date
- 2025-07-02
- Estimated Expiration
- 2041-04-09
AI Technical Summary
Existing data processing systems face challenges in providing secure and efficient access control to virtual memory spaces, particularly when granular protection information is stored in memory areas vulnerable to tampering, leading to potential security breaches and performance overheads.
Implementing a data integrity check mechanism for granule protection information stored in vulnerable memory areas using lightweight encryption and integrity checks, ensuring that even if one bit is tampered, the entire block is deemed corrupted, thereby reducing the need for extensive metadata and maintaining security without significant performance impact.
This approach enhances security by ensuring granule protection information integrity while minimizing performance overheads, allowing for fine-grained access control and reducing the risk of data corruption in memory systems.
Smart Images

Figure 0007701936000004 
Figure 0007701936000005 
Figure 0007701936000006
Abstract
Description
Technical Field
[0001] This technique relates to the field of data processing.
Background Art
[0002] A data processing system can have an address translation circuit for converting a virtual address of a memory access request into a physical address corresponding to the location to be accessed within the memory system.
[0003] At least some examples are apparatuses having an address translation circuit that converts a target virtual address specified by a memory access request into a target physical address associated with a selected physical address space selected from among a plurality of physical address spaces, a granularity protection entry load circuit for loading a granularity protection data block including at least one granularity protection entry from memory, each granularity protection entry corresponding to a respective granularity of a physical address and specifying granularity protection information indicating which of the plurality of physical address spaces is the permitted physical address space in which the granularity of the physical address is permitted to be accessed, a filtering circuit for determining whether a memory access request should be permitted to access the target physical address based on whether the selected physical address space is indicated as a permitted physical address space by the granularity protection information in a target granularity protection entry corresponding to a target granularity of the physical address including the target physical address, and a integrity check circuit for performing a data integrity check on the granularity protection data block loaded from memory and notifying a failure when the data integrity check fails.
[0004] At least some examples are data processing methods, which include converting a target virtual address specified by a memory access request into a target physical address associated with a selected physical address space selected from among a plurality of physical address spaces, and loading from memory a granularity protection data block including at least one granularity protection entry, where each granularity protection entry corresponds to a respective granularity of physical addresses and specifies granularity protection information indicating which of the plurality of physical address spaces is the permitted physical address space that is permitted to access the granularity of physical addresses, performing a data integrity check on the granularity protection data block loaded from memory, and notifying of a failure when the data integrity check fails, and determining whether the memory access request is to be permitted to access the target physical address based on whether the selected physical address space is indicated as a permitted physical address space by the granularity protection information in a target granularity protection entry corresponding to the target granularity of the physical address including the target physical address.
[0005] At least some examples are computer programs that, when executed on a host data processing device, contain instructions for controlling the host data processing device to provide an instruction execution environment for executing target code, the computer program comprising: address translation program logic for converting a target virtual address specified by a memory access request into a target simulated physical address associated with a selected simulated physical address space selected from among a plurality of simulated physical address spaces; granule protection entry load program logic for loading from memory a granule protection data block including at least one granule protection entry, each granule protection entry corresponding to a respective granule of the simulated physical address and specifying granule protection information indicating which of the plurality of simulated physical address spaces is the permitted simulated physical address space in which the granule of the simulated physical address is permitted to be accessed; filtering program logic for determining whether a memory access request should be permitted to access the target simulated physical address based on whether the selected simulated physical address space is indicated as a permitted simulated physical address space by the granule protection information in a target granule protection entry corresponding to the target granule of the simulated physical address including the target simulated physical address; and a integrity check circuit for performing a data integrity check on the granule protection data block loaded from memory and notifying of a failure when the data integrity check fails.
[0006] At least some examples provide a computer-readable storage medium storing the computer program described above. The computer-readable storage medium can be a non-transitory storage medium or a transitory storage medium.
Brief Description of the Drawings
[0007] Further aspects, features, and advantages of the present technology will become apparent from the following description of examples read in conjunction with the accompanying drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
DETAILED DESCRIPTION OF THE INVENTION
[0008] The data processing system can support the use of virtual memory and is provided with an address translation circuit that converts a virtual address specified by a memory access request into a physical address associated with a location within the accessed memory system. The mapping between the virtual address and the physical address can be defined by one or more page table structures. A page table entry within the page table structure can also define some access permission information that can control whether a given software process running on the processing circuit can access a particular virtual address.
[0009] In some processing systems, all virtual addresses can be mapped by the address translation circuit onto a single physical address space used by the memory system to identify a location within the accessed memory. In such a system, control over whether a particular software process can access a particular address is provided only based on the page table structure used to provide the virtual-to-physical address translation mapping. However, such a page table structure can typically be defined by the operating system and / or hypervisor. If the operating system or hypervisor is accessed illegally, it may cause a confidentiality leak where an attacker may be able to access sensitive information.
[0010] Thus, in some systems where a particular process needs to be securely executed in isolation from other processes, the system can support operations within several regions and can support several distinct physical address spaces. For at least some components of the memory system, memory access requests in which virtual addresses are translated to physical addresses in different physical address spaces are treated as if accessing completely different addresses in memory, even if the physical addresses within each physical address space actually correspond to the same location in memory. As can be seen for some memory system components, by isolating accesses to different physical address spaces from different operating regions of the processing circuit, stronger security guarantees can be provided that do not depend on page table permission information set by the operating system or hypervisor.
[0011] In a system that can map the virtual address of a memory access request to a physical address in one of two or more distinct physical address spaces, granularity protection information can be used to limit which physical addresses are accessible within a particular physical address space. This can be useful to ensure that access to a particular location of physical memory, implemented either on-chip or off-chip, can be restricted to a particular physical address space, or a particular subset of physical address spaces if desired.
[0012] The granule protection information can be loaded from a table structure stored in the memory. The table structure provides several granule protection entries, and each granule protection entry corresponds to a respective granule of the physical address. The granule protection information specifies, for the granule of the corresponding physical address, which physical address space is the permitted physical address space to which the granule of the corresponding physical address is permitted to access. It is possible to specify that none of the physical address spaces are permitted for the granule protection entry of a given granule, or one of the physical address spaces is the permitted physical address space, or two or more of the physical address spaces are the permitted physical address spaces. The filtering circuit can check whether a memory access request is permitted to continue using the granule protection information from the target granule protection entry corresponding to the accessed target physical address. This lookup can be an additional step of the lookup performed in addition to any access permission check performed by the address translation circuit based on the page table structure.
[0013] Since the granule protection information controls which physical address spaces can access a given set of physical addresses, damage or tampering with the granule protection information can lead to a complete loss of the security guarantees provided by the separation of the respective physical address spaces (for example, if the granule protection information is damaged to indicate that a physical address reserved for access from one physical address space should be accessible from another physical address space, this can lead to a leakage of confidential information). In a typical system involving the separation of multiple physical address spaces, the information that controls the partitioning between the physical address spaces is stored in a hardware storage device that is relatively secure against tampering by unauthorized attackers. For example, the partitioning data can be stored in a set of on-chip registers that are not exposed to external access. In such a system, the physical storage device for the partitioning information can be designed to provide sufficient security guarantees, and it is surprising that there is a need for data integrity checks for such information.
[0014] However, the inventor has recently recognized an increasing demand for having a relatively large number of different software processes operating on the same system that may desire to be provided with a secure "sandbox" that is isolated from access by other processes. The relationships between such software processes are becoming more complex, and thus, in such a model, it may not be possible to define access rights using only registers or other on-chip storage. To provide a more fine-grained control over access rights, a more complex table that defines granularity protection information at the granularity of individual granules of the physical address space and defines information for each granule of the physical address may be desirable. The granules of the physical address space (for each item of granularity protection information is defined) may be of a particular size that may be the same as or different from the size of the pages used in the page table structure used by the address translation circuit. In some cases, the granule may be of a size larger than the page that defines the address translation mapping of the address translation circuit. Alternatively, the granularity protection information may be defined at the same page-level granularity as the address translation information within the page table structure. Defining the granularity protection information at the page-level granularity may be convenient as it allows for a more fine-grained control over which regions of the memory storage hardware are accessible from a particular physical address space and thus from a particular operating region of the processing circuit. However, other sizes of granules can also be used.
[0015] In such a model, when there may be a larger number of physical address spaces, a larger number of different access permission types, and / or a larger number of separate regions of physical addresses for which different access rights are to be defined, the hardware circuit area / power cost may be too high. Therefore, it may no longer be practical to define granular protection information in the area of the hardware storage device that is immune to attacks. Thus, the granular protection information may instead be stored in normal memory, and the normal memory may expose some of the granular protection information to external access, and there is a risk that the granular protection information may be changed (intentionally or unintentionally by an attacker), thereby losing the security of the processes that depend on the granular protection information.
[0016] In the method described below, the device has an integrity check circuit for performing a data integrity check on the granular protection data block loaded from memory by the granular protection entry load circuit. If the data integrity check passes, the granular protection data block can be used for filtering memory access requests by the filtering circuit. If the data integrity check fails, the integrity check circuit notifies a failure. This makes it possible to detect the possibility of forgery or damage of the granular protection information while the granular protection information is stored in the vulnerable part of the memory, makes it possible to use more complex forms of granular protection information, and due to its size, it becomes impossible to store everything in the hardware protection storage that is immune to attacks. Therefore, security can be improved and the flexibility of the use cases supported by the granular protection information can be enhanced.
[0017] The timing of data integrity checks can vary. In some cases, a granularity protection data block may be loaded for each memory access, and in that case, each non-GPT related memory access (e.g., a normal memory access issued by the processing circuit or a page table walk memory access generated by the address translation circuit) may perform a data integrity check against the corresponding granularity protection data used to check whether the target address of the non-GPT related memory access is permitted access. However, if the filtering circuit can access a granularity protection data cache for caching the granularity protection data (this cache may be a stand-alone cache dedicated to caching the granularity protection data or may be shared with the caching of address translation data), the performance can be faster. In this case, for a memory access request that hits in the granularity protection data cache, since it is not necessary to load the granularity protection data from memory, a data integrity check is not required (it is assumed that the cache is robust against errors and additional integrity checks may not be necessary), but for a request that misses in the cache, a data integrity check may be performed when the requested granularity protection data is returned from memory.
[0018] Data integrity protection techniques are known to protect normal data stored in vulnerable memory (such as off-chip memory), but there is a problem of incurring a relatively large cost in terms of processing performance, power consumption, and memory capacity. Encryption of data stored in a vulnerable memory area may be sufficient to ensure confidentiality, but by itself it does not protect against the possibility that the data is (intentionally or unintentionally) corrupted while it is stored in the vulnerable area of the memory, so when the data is read, it has a value different from the stored value. In systems that require protection against tampering, integrity metadata (such as a hash or authentication code) can be derived from the data stored at the time it is written, and then compared with the corresponding metadata recomputed from the data stored at the time the data is read to determine whether the data has changed. However, maintaining such metadata for a reasonably sized protected address area may require a relatively large metadata structure, which may be too large to store in hardware-protected memory, so some of that metadata may need to be stored in memory vulnerable to tampering, and as a result, further integrity metadata is derived from it and needs to be stored later to check the metadata against tampering. Since that further integrity metadata itself may require protection, etc., a tree structure can be maintained to track the metadata at each level, and ultimately, all the metadata (and the underlying protected data) can be verified using a root of trust that is small enough to store in hardware-protected memory that is not vulnerable to tampering. Maintaining this metadata (such as a Merkle tree) can be costly in terms of increasing the required memory capacity and in terms of performance cost because each actual read / write of data to the memory area to be protected by the metadata tree requires several other reads / writes just to maintain or check the multiple integrity metadata within the tree.In some implementations, there is only a relatively small portion of the address space that requires integrity protection. Therefore, the performance cost of such integrity protection technology is justified because most memory accesses do not require it for accessing addresses outside the integrity protection area.
[0019] However, when a data integrity check is applied to the granularity protection information used to filter whether a memory access is permitted based on whether the selected physical address space is the permitted physical address space, (when that granularity protection data is not already available and must be loaded from memory), this can mean that the need for a data integrity check no longer restricts access to read / write data in a relatively small portion of the physical address space, and thus the execution of the data integrity check can cause many more memory accesses to be delayed (even if the granularity protection data itself is stored in a relatively small integrity protection portion of the physical address space, access to that granularity protection data can be caused by access to any part of the physical address space). Therefore, this is counterintuitive to those skilled in the art, that it is desirable to store the granularity protection data in a memory area vulnerable to corruption / alteration, or that it is desirable to require a data integrity check on the granularity protection data block loaded from memory before it is used to filter memory access requests, and is another reason why the implementation cost of integrity protection is considered too high for data that may potentially be required for all memory accesses.
[0020] However, the inventor has recognized several aspects related to granularity protection data that can be used to enable checking of granularity protection data using a much lighter form of integrity check compared to conventional integrity check techniques.
[0021] Since the granule protection information for a given granule can be of a size smaller than the size of the data block requested from memory in a given transaction, when it is necessary to load a target granule protection entry to handle a particular memory access request, the granule protection entry load circuit may actually load a granule protection data block that includes a plurality of granule protection entries, each of those granule protection entries corresponding to a different granule of physical addresses. Even if other granule protection entries are not currently needed, they may be needed to handle other memory access requests promptly, so it may be useful to cache those other granule protection entries. Since the integrity check circuit has the entire block of available granule protection entries, it can perform a data integrity check on the block of granule protection entries as a whole.
[0022] Also, when the granule protection information is stored in a vulnerable memory area accessible by an attacker, a memory encryption circuit may be provided. A protected memory area can be defined that corresponds to the memory areas vulnerable when not protected by encryption and is designated to store secure data protected from attacks (since not all data is confidential, the protected area does not necessarily have to cover the entire address range of the vulnerable memory storage). When data is written to the protected area of the memory, the memory encryption circuit encrypts the data specified in the write access request and provides the encrypted data to be stored to the protected area of the memory. When data is read from the protected area of the memory, the memory encryption circuit decrypts the encrypted data read from the protected area and provides the decrypted data as a response to the corresponding read access request. Therefore, when granule protection data is stored in a protected area protected by encryption using a memory encryption circuit, when the granule protection entry load circuit loads the granule protection data block from the protected area of the memory, the memory encryption circuit can decrypt the granule protection data block read from the protected area of the memory before providing it to the filtering circuit or the integrity check circuit.
[0023] Therefore, in a system that discloses granule protection information in a memory area (such as off-chip memory) that may be vulnerable to tampering or damage, memory encryption can be used for the granule protection information when it is stored in that vulnerable memory area. Encryption / decryption functions used for the mapping between plaintext and ciphertext can be more secure if they are designed to have the property of "diffusion", where changing one bit of either the plaintext / ciphertext causes the other to change by a much larger number of bits.
[0024] The memory encryption circuit can perform encryption / decryption according to a specific cryptographic primitive that operates on blocks of data of a given size. The size of the data block operated on by the encryption primitive may be larger than the size of one granule protection entry so that multiple granule protection data blocks of granule protection entries can be encrypted as a common block of data by the memory encryption circuit.
[0025] Accordingly, when a granule protection data block is stored in the protected memory area to be encrypted, there is a correlation between the integrity of one granule protection entry in the granule protection data block and the integrity of another granule protection entry in the same granule protection data block. This is because even if only a single bit is changed in the ciphertext while stored in the memory, it may be possible to destroy multiple granule protection entries in the block. Thus, if the target granule protection entry is loaded because it is not yet accessible to the filtering circuit and the data integrity check fails for one of the multiple granule protection entries of the granule protection data block other than the target granule protection entry, the integrity check circuit can be configured to notify of a failure even if the data integrity check passes for the target granule protection entry.
[0026] Thus, even if the target granularity protection entry itself that is currently required to filter the current memory access request passes the check, a check is performed on a block of multiple entries, and if any one of the granularity protection entries within the block fails the integrity check, by notifying a failure, the probability that the data integrity check fails to detect corruption (false negative) of any of the granularity protection entries within the block can be extremely low, even if the probability of a false negative for an individual granularity protection entry is relatively high. Therefore, a lighter-weight data integrity check can be used (the probability of detecting forgery in an individual granularity protection entry is lower).
[0027] For example, the granularity protection information can be encoded using a specific number of bits (e.g., N bits) of the granularity protection entry. All of the 2 N possible encodings of those N bits are not necessarily required to represent valid options for the granularity protection information. For example, the number of valid options required may not be an exact power of two, and thus there may be some spare encodings left over. Also, the designer of the instruction set architecture can choose to leave some encodings available to provide room for future architecture extensions in case additional options need to be added later.
[0028] Therefore, the granule protection information can be encoded using N bits with several encodings. The first subset of the N-bit encoding is a valid encoding indicating valid options for the granule protection information. The second subset of the N-bit encoding is an invalid encoding, which does not indicate valid options for the granule protection information. In one example, the data integrity check may include checking for each granule protection entry in the granule protection data block whether the granule protection entry has one of the first subset of encodings or the second subset of encodings. If any one of the granule protection entries in the granule protection data block has one of the second subset of encodings, a failure is notified (even if the granule protection entry with the invalid encoding is not the target granule protection entry required for filtering the current memory access request).
[0029] It may be surprising that simply checking whether the encoding of the granule protection information is a valid encoding is sufficient to function as a data integrity check. This check is considered insufficient because it cannot detect tampering that changes the granule protection information from one valid encoding to a different valid encoding.
[0030] However, as described above, by using encryption, if one granule protection entry within a loaded block has an invalid encoding, it can be notified that due to the diffusion characteristics of the encryption function, there is a risk that all other granule protection entries within the same encryption primitive block may also be changed. Therefore, in order to allow an attacker to modify the encrypted block of the stored granule protection entry so that the data integrity check does not detect the tampering, this requires that the encrypted version of the stored data be changed such that each granule protection entry is changed to one of the valid encodings of the first subset and no granule protection entry is changed to one of the invalid encodings of the second subset.
[0031] Often, the combination of the number of granule protection entries within a data block and the number of encodings within the second subset of encodings can be sufficient such that the probability that all granule protection entries have one of the invalid encodings of the second subset due to tampering becomes acceptably small. If this is still not the case, the architecture designer can ensure that this is achieved either by (i) increasing the number of granule protection entries within the data block or (ii) providing one or more additional bits of redundant granule protection information to increase the number of invalid encodings within the second subset. Thus, another possible reason for having an invalid encoding of granule protection information can simply be to enhance protection against tampering. The second option of adding redundant granule protection information bits that are not required to encode valid information may seem wasteful, but it may require fewer additional bits than are required to store an authentication code or other metadata for comparison with the stored data, avoiding the significant additional read / write performance overhead for maintaining integrity metadata as described above.
[0032] Therefore, in this approach, the data integrity check may be a simple check of whether the encoding is valid, and does not require additional reading / writing of stored metadata or authentication code calculation to derive an authentication code from the stored data. The data integrity check can be implemented with a relatively simple (low-power and low-latency) set of logic gates to check whether the encoding of each granule protection entry is one of the first subset of valid encodings. This significantly reduces the performance and circuit area cost of performing the integrity check. When this approach to data integrity checking is used, the software that manages the granule protection entries can ensure that other granule protection entries within the same granule protection data block are set to valid encodings when setting the granule protection entries for the granules of interest (even if the granules corresponding to those granule protection entries are not currently in use).
[0033] From an architectural perspective, the choice of which memory region should store the granularity protection information is made by the system designer implementing a particular processing system, and thus is not an architectural feature defined by the instruction set architecture used to define the interaction between program code and processing hardware. Therefore, in any given processor system implementing an instruction set architecture, it is not essential that the granularity protection data be stored in any particular location or that a memory encryption circuit be provided. The hardware system designer (or the party operating the hardware system, such as a server operator) can choose, for a particular implementation, to store the granularity protection information in a region of on-chip memory that is protected from tampering by its hardware design (and thus does not require encryption), or to store it in a memory region that is potentially vulnerable to attack because it is outside the trust boundary where memory encryption can be used (e.g., off-chip memory on a different integrated circuit such that the memory interface between integrated circuits can be probed by an attacker).
[0034] In either case, the architecture definition of the granule protection entry load circuit, filtering circuit, and integrity check circuit can be such that the circuit performs a data integrity check (by verifying that the encoding is valid as described above) when the granule protection block is loaded from memory before the granule protection information is used for filtering memory accesses. If the system designer chooses to implement the granule protection information in hardware protected memory that is not vulnerable to tampering, that data integrity check may be redundant, but in either case it is relatively lightweight and thus does not have a major performance impact. However, if the system designer chooses to store the granule protection information in memory that is vulnerable to tampering, the system designer may be expected to also provide a memory encryption circuit, in which case the data integrity check for verifying the valid encoding of the granule protection information can help protect against the risk that an attacker may modify the encrypted granule protection information while it is stored in the vulnerable memory.
[0035] Data integrity checks that check for valid encoding of granule protection information are particularly lightweight and effective. In other examples, however, it may be preferable to implement a data integrity check that includes determining whether at least one integrity signature value associated with a granule protection data block matches at least one signature check value derived from the granule protection data block. In the case of a system that only needs to protect a relatively small physical address space using granule protection data, current integrity protection techniques may be viable and may provide a stronger probability of detecting tampering than the above check for valid encoding. Also, currently expensive in terms of performance for protecting a larger physical address space, but may become more viable in the future, and as integrity protection techniques themselves are further developed, it may become possible to provide a larger storage capacity for integrity metadata in hardware-protected memory, and it is expected that access to memory will become faster. Thus, in other implementations, the data integrity check can be based on integrity signature values stored in relation to the granule protection data block, which can be checked when the granule protection data block is read to confirm whether they match the signature check values, and the signature check values include a newly calculated signature derived from the read granule protection data block using the same signature calculation function used to derive the integrity signature value when the granule protection data block was written. For example, the signature can be a hash or message authentication code calculated from the granule protection data block based on a private key and, optionally, a freshness value such as a counter or other value that changes each time the block is written (this can protect against replay attacks based on replacing the granule protection data block and the latest value of the signature with old values that were previously correct but are no longer correct).
[0036] In response to a memory access request, if there is no defined granule protection entry valid for the target granule of the physical address, or if the filtering circuit determines that the memory access request is not permitted to access the target physical address based on the target granule protection entry, the filtering circuit may notify the fault notified by the integrity check circuit of a fault associated with a different fault type or different fault syndrome information in the case where the data integrity check fails. Thus, whether via different fault types (e.g., different fault identifiers) or by recording information in the syndrome register to indicate the cause of the fault, the exception handler that processes the fault can identify from the fault type or syndrome information whether the fault was (i) caused by a loss of data integrity of the granule protection information, which does not necessarily mean that the memory access request should have been rejected if the granule protection information was not corrupted, or (ii) caused by the granule protection information being invalid or the granule protection information indicating that the selected physical address space required to be accessed by the memory access request is not a permitted physical address, in which case the memory access request is rejected. Thereby, the handler can take different response actions.
[0037] The filtering circuit may comprise a requester-side filtering circuit for determining whether to pass a memory access request to the cache or to pass it to an interconnect for communicating with a completer device for servicing the memory access request, based on the granularity protection information within the target granularity protection entry. (After the memory access request is routed via the interconnect to a completer device such as a memory controller or a peripheral device controller to service the memory access request) Comparing with performing filtering of the memory access request based on the granularity protection information on the completer side, the advantage of performing the granularity protection lookup on the requester side of the interconnect is that this enables finer-grained control over which physical addresses are accessible from a given physical address space rather than which are practical on the completer side. This is because the completer side may typically have a relatively limited ability to access the entire memory system. For example, the memory controller of a given memory unit may only be allowed to access locations within that memory unit and may not be able to access other regions of the address space. Providing finer-grained control may rely on a more complex table of granularity protection information that may be stored in the memory system, and it may be more practical to access such a table from the requester side, which has more flexibility in issuing memory access requests to a wider subset of the memory system. Also, performing the granularity protection lookup on the requester side can help enable the ability to dynamically update the granularity protection information during execution, which may not be practical for a completer-side filtering circuit that may be limited to accessing a relatively small amount of statically defined data defined at startup.Another advantage of the requester-side filtering circuit is that it allows the interconnect to distribute different addresses within the same granule to different computer ports that communicate with different computer devices (e.g., different DRAM (Dynamic Random Access Memory) units), which is performance-efficient. However, it would be impractical if the entire granule had to be directed to the same computer unit so that the computer side could perform a granule protection lookup to verify whether a memory access is possible. Therefore, there can be many advantages to performing the granule protection lookup, which determines whether access can be made from a particular physical address space selected for a given memory access request to a particular physical address, on the requester side rather than on the computer side.
[0038] The integrity check circuit may include a requester - side integrity check circuit for performing a data integrity check on the granule protection data block loaded from the memory after the granule protection data block has been received from the interconnect. When a memory protection engine with a memory encryption circuit is provided to apply encryption to the data stored in the protected area of the memory, the memory encryption circuit may typically be implemented within the memory system at the interface with the memory storage hardware that is vulnerable to eavesdropping / tampering. This may typically be on the completer side of the interconnect rather than on the requester side. In a system that provides a function for integrity verification applied to general data stored in the protected area of the memory (e.g., based on an authentication code or other integrity metadata) to the memory protection engine, this can usually also be managed from the completer - side memory protection engine. In contrast, in some embodiments, the above - mentioned integrity check circuit may be on the requester side, for example, as part of a memory management unit with a filtering circuit or a system memory management unit. In particular, in an implementation that uses a data integrity check based on verification of whether the granule protection information has one of a first subset of valid encodings or a second subset of invalid encodings, this may be a custom - type verification specific to the granule protection information rather than a more general - form integrity verification applicable to other data. Therefore, it may be more efficient to implement this on the requester side. Thus, by incorporating the integrity check logic into a requester - side system component that also has a filtering circuit (and possibly address translation as well), system components that use granule protection information can include a circuit for guaranteeing the integrity of the information before use, which can simplify the system design. This eliminates the need for the system designer to consider data integrity verification when designing the completer - side components of the memory system.
[0039] A memory access request may be issued by a requester circuit (e.g., a processing circuit) that can operate in any of two or more operating regions.
[0040] The selected physical address space associated with the memory access can be selected in various ways. In some examples, the selected physical address space can be selected (by an address translation circuit or a filtering circuit) based at least on the current region of operation of the requester circuit in which the memory access request was issued. The selection of the selected physical address space may also depend on physical address space selection information specified in at least one page table entry used for translating the target virtual address to the target physical address.
[0041] The address translation circuit may limit which physical address spaces are accessible depending on the current region. If a particular physical address space is accessible in the current region, this means that the address translation circuit can translate a virtual address specified for a memory access issued from the current region to a physical address within that particular physical address space. This does not necessarily mean that the memory access is permitted, because even if a particular memory access can translate its virtual address to a physical address within a particular physical address space, the filtering circuit can perform a further check based on the granularity protection information to determine whether that physical address is actually permitted to be accessed within that particular physical address space. Nevertheless, stronger security guarantees can be provided by limiting which subset of the physical address space is accessible in the current region.
[0042] For example, the processing circuit can support processing within a root area related to managing the switching between other areas where the processing circuit can operate. The root area can have an associated root physical address space. By providing a dedicated root area for controlling the switching, it can help maintain security by limiting the extent to which code running in one area can trigger a switch to another area. For example, the root area can perform various security checks when a switch between areas is requested. Thus, the processing circuit can support processing being executed in one of at least three areas, namely the root area and at least two other areas.
[0043] If the data integrity check of the granule protection data block fails, the integrity check circuit can notify that the failure is to be processed by the program code executed in the root area. Thus, the problem of data integrity loss can be managed by the root area responsible for the separation of the physical address space and areas and can be regarded as having the highest level of security.
[0044] In some examples, the processing circuit can support two additional regions in addition to the root region. For example, the other regions can include a secure region associated with a secure physical address space and a less secure region associated with a less secure physical address space. The less secure physical address space can be accessible from each of the less secure region, the secure region, and the root region. The secure physical address space can be accessible from the secure region and the root region, but may not be accessible from the less secure region. The root physical address space can be accessible to the root region, but may not be accessible to the less secure region and the secure region. Thus, this allows code running in the secure region to protect its code or data from access by code operating in the less secure region with stronger security guarantees than when the page table is used as the sole security control mechanism. For example, portions of the code that require stronger security can be run within a secure region managed by a trusted operating system separate from the non-secure operating system operating in the less secure region. An example of a system that supports such secure and less secure regions can be a processing system operating according to a processing architecture that supports the TrustZone trademark architecture feature provided by Arm Limited (Cambridge, UK). In conventional TrustZone implementations, the monitor code for managing the switch between the secure region and the less secure region uses the same secure physical address space used by the secure region. In contrast, providing a root region for managing the switch between other regions and allocating a dedicated root physical address space for use by the root region helps to improve security and simplify system development.
[0045] However, in other examples, other regions may include at least three other regions in addition to a further region, such as a root region. These regions can include the secure regions and less secure regions described above, but may also include at least one further region associated with a further physical address space. The less secure physical address space may also be accessible from the further region, but the further physical address space may be accessible from the further region and the root region, and may not be accessible from the less secure region. Thus, similar to the secure region, the further region can be considered more secure than the less secure region, allowing for further partitioning of code into separate worlds associated with separate physical address spaces and restricting their interactions.
[0046] In some examples, each region is hierarchically associated with a privilege level that increases as the system ascends from a less secure region to a root region via secure regions and further regions, with the further region being considered to have higher privilege than the secure region and thus may have access to the secure physical address space.
[0047] However, there is an increasing desire for a secure computing environment that limits the need for a software provider to trust other software providers associated with other software running on the same hardware platform. For example, there can be uses in several fields where the provider of software code may be negative about trusting the provider of the operating system or hypervisor (components that may have been considered trustworthy in the past), such as in mobile payment and banking, implementing anti-cheat or anti-piracy mechanisms in computer games, security extensions for operating system platforms, secure virtual machine hosting in cloud systems, and confidential computing. In systems such as those based on the above TrustZone (registered trademark) architecture that support secure and less secure regions with each having its own physical address space, as secure components operating in the secure region become more prevalent, the set of software that normally operates in the secure region has come to expand to include several software provided by different software providers, including, for example, the following parties. An original equipment manufacturer (OEM) that assembles a processing device (such as a mobile phone) from components including a silicon integrated circuit chip provided by a specific silicon provider, an operating system vendor (OSV) that provides an operating system to be executed on the device, and a cloud platform operator (or cloud host) that maintains a server farm providing server space for hosting virtual machines on the cloud.Thus, when regions are implemented in a strict order that increases privilege, an application provider that provides application-level code that desires to provide a secure computing environment may not want to trust parties (such as OSVs, OEMs, cloud hosts, etc.) that may have previously provided software to execute secure regions. Similarly, parties that provide code that operates in a secure region are unlikely to want to trust an application provider that provides code that operates in a higher-privilege region that is given access to data related to a less-privileged region, so problems can occur. Thus, it is recognized that a strict hierarchy of regions that continuously increases privilege may not be appropriate.
[0048] Thus, in the more detailed examples below, additional regions can be considered orthogonal to the secure region. The additional region and the secure region can each access a less-secure physical address space, but the additional physical address space associated with the additional region is inaccessible from the secure region, and at the same time, the secure physical address space associated with the secure region is inaccessible from the additional region. The root region can still access the physical address space associated with both the secure region and the additional region.
[0049] Thus, in this model, an additional region (an example of which is the realm region described in the examples below) and the secure region have no dependency on each other and thus do not need to trust each other. The secure region and the additional region only need to trust the root region, which is inherently trusted because it manages entry to the other regions.
[0050] The following example describes a single instance of a further region (realm region), but it is understood that each of the secure region and at least two further regions can access a less secure physical address space, cannot access the root physical address space, and cannot access physical address spaces associated with each other, so that the principle of further regions orthogonal to the secure region can be extended to provide a plurality of further regions.
[0051] The less secure physical address space may be accessible from all regions supported by the processing circuit. This is useful for facilitating the sharing of data or program code between software running in different regions. If a particular item of data or code is to be made accessible in different regions, it can be allocated to the less secure physical address space so that it can be accessed from any region.
[0052] The memory system may include a physical alias point (PoPA) which is the point at which an aliased physical address from a different physical address space corresponding to the same memory system resource is mapped (de-aliased) to a single physical address that uniquely identifies that memory system resource. The memory system may include at least one pre-PoPA memory system component provided upstream of the PoPA, which treats the aliased physical addresses as if they corresponded to different memory system resources.
[0053] For example, at least one PoPA pre-memory system component can include a cache or translation lookaside buffer, which can cache data, program code, or address translation information for aliased physical addresses in separate entries, so that if the same memory system resources are required to be accessed from different physical address spaces, the access causes a different cache or TLB entry to be allocated. Also, the PoPA pre-memory system component can include a coherence control circuit such as a coherent interconnect, snoop filter, or other mechanism for maintaining coherence between cache information at each master device. The coherence control circuit can assign separate coherence states to each aliased physical address within different physical address spaces. Thus, aliased physical addresses are treated as separate addresses for the purpose of maintaining coherence even if they actually correspond to the same underlying memory system resources. At first glance, tracking coherence separately for aliased physical addresses might seem to cause a problem of coherence loss, but in fact, this is not a problem because if processes operating in different regions are intended to actually share access to a particular memory system resource, they can access that resource using a less secure physical address (or use the limited sharing feature described below to access the resource using one of the other physical address spaces). Another example of a PoPA pre-memory system component can be a memory protection engine provided to protect data stored in off-chip memory from loss of confidentiality and / or tampering. Such a memory protection engine can, for example, separately encrypt data associated with a particular memory system resource using different encryption keys depending on the physical address space from which the resource was accessed, treating the aliased physical address as if it corresponded to a different memory system resource (e.g., an encryption scheme that depends on the address can be used, and the physical address space identifier can be considered part of the address for this purpose).
[0054] Regardless of the form of the PoPA pre-memory system components, it may be useful to treat such PoPA memory system components as if the aliased physical addresses corresponded to different memory system resources. This is because it provides hardware-implemented isolation between accesses issued to different physical address spaces, and as a result, information associated with one region is not leaked to another region due to characteristics such as a cache timing side channel or a side channel being triggered by the coherence control circuit and involving a change in coherence.
[0055] In some implementations, the aliased physical addresses within different physical address spaces may be represented using different numerical physical address values for each of the different physical address spaces. This approach may require a mapping table in PoPA to determine which different physical address values correspond to the same memory system resource. However, this overhead of maintaining the mapping table may be considered unnecessary, and in some implementations, it may be simpler if the aliased physical addresses include physical addresses represented using the same numerical physical address value in each of the different physical address spaces. When this approach is taken, at the physical aliasing point, it may be sufficient to simply discard the physical address space identifier that identifies which physical address space is being accessed using the memory access, and then provide the remaining physical address bits downstream as the unaliased physical address.
[0056] Accordingly, in addition to the PoPA front memory system components, the memory system may also include PoPA memory system components configured to de-alias a plurality of aliased physical addresses to obtain de-aliased physical addresses provided to at least one downstream memory system component. The PoPA memory system component can be a device that accesses a mapping table to find the de-aliased address corresponding to the aliased address within a specific address space, as described above. However, the PoPA component can simply be a location within the memory system where the physical address tag associated with a given memory access is discarded, such that the physical address provided downstream uniquely identifies the corresponding memory system resource regardless of which physical address space it was provided from. Alternatively, in some cases, the PoPA memory system component may still provide a physical address space tag to at least one downstream memory system component (e.g., for the purpose of enabling completer-side filtering as further discussed below). However, PoPA marks a point within the memory system such that downstream memory system components no longer treat the aliased physical addresses as different resources and can map the same memory system resources considering each of the aliased physical addresses. For example, if a memory controller or a hardware memory storage device downstream of PoPA receives the physical address tag and physical address of a given memory access request, if the physical address corresponds to the same physical address as a previously seen transaction, any hazard checks or performance improvements that are executed for each transaction accessing the same physical address (such as merging accesses to the same address) can be applied, even if each transaction specified a different physical address space tag. In contrast, for memory system components upstream of PoPA, such hazard checks or performance improvement steps taken for transactions accessing the same physical address may not be triggered if these transactions specify the same physical address within different physical address spaces.
[0057] The above-described technology can be implemented in a hardware device having hardware circuit logic for implementing the functions as described above. Accordingly, the processing circuit and the address conversion circuit may include hardware circuit logic. However, in other examples, a computer program that controls a host data processing device to provide an instruction execution environment for executing target code may provide processing program logic and address conversion program logic that execute, in software, functions equivalent to those of the above-described processing circuit and address conversion circuit. This can be useful, for example, to enable target code written for a particular instruction set architecture to be executed on a host computer that may not support that instruction set architecture. Thus, simulation software can emulate the functions expected by an instruction set architecture not provided by the host computer by providing an equivalent instruction execution environment for the target code such that the functions expected when the target code is executed on a hardware device that actually supports the instruction set architecture are emulated. Accordingly, address conversion program logic, granularity protection entry load program logic, filtering program logic, and integrity check logic may be provided to emulate the functions of the aforementioned address conversion circuit, granularity protection entry load circuit, filtering circuit, and integrity check circuit. In the case of the approach in which simulation of the architecture is provided, each physical address space is a simulated physical address space because, although they do not correspond to the physical address spaces actually identified by the hardware components of the host computer, they are mapped to addresses within the virtual address space of the host.Providing such a simulation can be useful for various purposes, for example, making old code written for one instruction set architecture executable on different platforms that support different instruction set architectures, or assisting in the software development of new software that is to be executed for a version of an instruction set architecture when hardware devices that support the new version of the instruction set architecture are not yet available (thereby making it possible to start the development of software for the new version of the architecture in parallel with the development of hardware devices that support the new version of the architecture).
[0058] FIG. 1 schematically shows an example of a data processing system 2 having at least one requester device 4 and at least one completer device 6. An interconnect 8 provides communication between the requester device 4 and the completer device 6. The requester device can issue a memory access request that requests access to a specifically addressable memory system location. The completer device 6 is a device having the responsibility of processing the memory access request directed thereto. Although not shown in FIG. 1, some devices may be able to function as both a requester device and a completer device. The requester device 4 can include, for example, a processing element such as a central processing unit (CPU) or a graphics processing unit (GPU), or other master devices such as a bus master device, a network interface controller, a display controller, etc. The completer device can include a memory controller that plays a role of controlling access to a corresponding memory storage device, a peripheral controller for controlling access to peripheral devices, etc. FIG. 1 shows a more detailed exemplary configuration of one of the requester devices 4, but it will be understood that other requester devices 4 can have a similar configuration. Alternatively, other requester devices may have a configuration different from that of the requester device 4 shown on the left side of FIG. 1.
[0059] The requester device 4 has a processing circuit 10 for executing data processing in response to an instruction by referring to the data stored in the register 12. The register 12 may include not only general-purpose registers for storing operands and the results of processed instructions, but also control registers for storing control data for configuring how the processing is executed by the processing circuit. For example, the control data may include a current region indication 14 used to select which operating region is the current region, and a current exception level indication 15 indicating which exception level is the current exception level at which the processing circuit 10 is operating.
[0060] The processing circuit 10 may be capable of issuing a memory access request that specifies a virtual address (VA) identifying an addressable location to be accessed and a region identifier (region ID or "security state") identifying the current region. The address translation circuit 16 (e.g., a memory management unit (MMU)) converts the virtual address to a physical address (PA) via one of more stages of address translation based on page table data defined in a page table structure stored in the memory system. The translation lookaside buffer (TLB) 18 functions as a lookup cache for caching a portion of that page table information for faster access than would be required if the page table information had to be fetched from memory each time an address translation is needed. In this example, in addition to generating the physical address, the address translation circuit 16 also selects one of several physical address spaces associated with the physical address and outputs a physical address space (PAS) identifier identifying the selected physical address space. The selection of the PAS will be discussed in more detail below.
[0061] The PAS filter 20 functions as a requester-side filtering circuit for verifying whether the physical address is permitted to be accessed within the physical address space identified and specified by the PAS identifier based on the converted physical address and the PAS identifier. This lookup is based on the granule protection information stored in the granule protection table structure stored in the memory system. The granule protection entry load circuit (also known as the granule protection table load or GPT load circuit) 21 loads a block of data including the granule protection information from the memory. The integrity of the loaded block of granule protection information is checked by the integrity check circuit 23, as further explained below. The granule protection information can be cached in the granule protection information cache 22, similar to the cache of the page table data in the TLB 18. In the example of FIG. 1, the granule protection information cache 22 is shown as a separate structure from the TLB 18, but in other examples, these types of lookup caches can be combined as a single lookup cache structure, such that a single lookup of an entry in the combined structure provides both the page table information and the granule protection information. The granule protection information defines the physical address space within which a given physical address can be accessed, and based on this lookup, the PAS filter 20 determines whether to proceed with issuing a memory access request to one or more caches 24 and / or the interconnect 8. If the specified PAS for the memory access request is not permitted to access the specified physical address, the PAS filter 20 can block the transaction and signal a fault.
[0062] FIG. 1 shows an example of a state where the system has multiple requester devices 4, but the features shown for one of the request devices on the left side of FIG. 1 can also be included in a system with only one requester device, such as a single-core processor.
[0063] FIG. 1 shows an example in which the selection of a PAS for a given request is performed by the address translation circuit 16. In other examples, however, information for determining which PAS to select can be output by the address translation circuit 16 to the PAS filter 20 together with the PA, and the PAS filter 20 can select a PAS and check whether the PA can be accessed within the selected PAS.
[0064] The provision of the PAS filter 20 helps to support a system that can operate in several operating regions, each associated with its own isolated physical address space. Here, for at least a part of the memory system (for example, for a coherence implementation mechanism such as some caches or snooping filters), separate physical address spaces are treated as if they were a different set of addresses identifying a different memory system location, even if the addresses within those address spaces actually point to the same physical location within the memory system. This can be useful for security purposes.
[0065] FIG. 2 shows examples of different operating states and regions in which the processing circuit 10 can operate, as well as examples of the types of software that can be executed in different exception levels and regions (it will of course be understood that the specific software installed on the system is selected by the party managing the system and is therefore not an essential feature of the hardware architecture).
[0066] The processing circuit 10 is operable at several different exception levels 80, in this example, four exception levels labeled EL0, EL1, EL2, and EL3, where in this example, EL3 refers to the exception level with the highest level of privilege and EL0 refers to the exception level with the lowest privilege. In other architectures, it will be understood that the reverse numbering may be chosen and the exception level with the largest number may be considered to have the lowest privilege. In this embodiment, the lowest privilege exception level EL0 is for application level code, the next highest privilege exception level EL1 is used for operating system level code, the next highest privilege exception level EL2 is used for hypervisor level code that manages switching between several virtual operating systems, and the highest privilege exception level EL3 is used for monitor code that manages switching between respective regions and the allocation of physical addresses to the physical address space, as will be described later.
[0067] When an exception occurs at a particular exception level while software is being processed, for some types of exceptions, the exception is accepted for a higher (more privileged) exception level, and the particular exception level at which the exception is accepted is selected based on the attributes of the particular exception that occurred. However, in some situations, other types of exceptions may be accepted at the same exception level as the exception level associated with the code being processed when the exception was accepted. When an exception is accepted, information characterizing the state of the processor at the time the exception was accepted can be saved, including, for example, the current exception level at the time the exception was accepted. Thus, when an exception handler is processed to handle the exception, the processing can return to the previous processing and use the saved information to identify the exception level to which the processing should return.
[0068] In addition to the different exception levels, the processing circuit also supports several operating regions including a root region 82, a secure (S) region 84, a less secure region 86, and a realm region 88. For ease of reference, the less secure region is described below as the "non-secure" (NS) region, although it should be understood that this is not intended to imply a particular level (or lack) of security. Instead, "non-secure" simply indicates that the non-secure region targets code that is less secure than the code operating in the secure region. The root region 82 is selected when the processing circuit 10 is at the highest exception level EL3. When the processing circuit is at one of the other exception levels EL0 - EL2, the current region is selected based on the current region indicator 14, which indicates which of the other regions 84, 86, 88 is active. For each of the other regions 84, 86, 88, the processing circuit can be at any of the exception levels EL0, EL1, or EL2.
[0069] At startup, some boot code (e.g., BL1, BL2, OEM boot) can be executed, for example, within a higher privilege exception level EL3 or EL2. The boot codes BL1, BL2 can be associated with, for example, the root region, and the OEM boot code can operate in the secure region. However, once the system is booted, during execution, the processing circuit 10 can be considered to operate in one of the regions 82, 84, 86, and 88 at a time. Each of the regions 82 - 88 is associated with its own related physical address space (PAS). This enables the isolation of data from different regions within at least a part of the memory system. This will be described in more detail below.
[0070] The non-secure region 86 can be used for normal application-level processing and operating system and hypervisor activities for managing such applications. Thus, within the non-secure region 86, there can be application code 30 operating at EL0, operating system (OS) code 32 operating at EL1, and hypervisor code 34 operating at EL2.
[0071] The secure region 84 enables specific system-on-chip security, media, or system services to be isolated in a physical address space separate from the physical address space used for non-secure processing. The secure region and the non-secure region are not equivalent in the sense that non-secure region code cannot access resources associated with the secure region 84, while the secure region can access both secure and non-secure resources. An example of a system that supports such a partitioning of the secure and non-secure regions 84, 86 is a system based on the TrustZone (registered trademark) architecture provided by Arm (registered trademark) Limited. The secure region can execute a trusted application 36 at EL0, a trusted operating system 38 at EL1, and optionally a secure partition manager 40 at EL2. When secure partitioning is supported at EL2, EL2 can use a two-page table to support isolation between different trusted operating systems 38 running within the secure region 84 in a manner similar to how the hypervisor 34 can manage isolation between virtual machines or guest operating systems 32 running within the non-secure region 86.
[0072] Extending the system to support the secure region 84 has become common in recent years because it enables a single hardware processor to support isolated secure processing and avoids the need for the processing to be performed on a different hardware processor. However, as the popularity of using secure regions has grown, many practical systems with such secure regions now support, within the secure region, a relatively high degree of mixed environments of services provided by a wide range of different software providers. For example, the code operating within the secure region 84 can include different software, and its providers can include (among others) silicon providers that manufacture integrated circuits, original equipment manufacturers (OEMs) that assemble integrated circuits provided by silicon providers into electronic devices such as mobile phones, operating system vendors (OSVs) that provide the operating system 32 for the device, and / or cloud platform providers that manage cloud servers that support services for a number of different customers via the cloud.
[0073] However, there is an increasing demand for a secure computing environment in which the provider of user-level code (which can normally be expected to execute as application 30 within non-secure region 86) can be trusted not to leak information to other parties operating code on the same physical platform. Such a secure computing environment is preferably dynamically allocable during execution and provably guaranteed such that a user can verify whether sufficient security guarantees are provided on the physical platform before entrusting the device with the processing of potentially sensitive code or data. A user of such software may not wish to trust the provider of operating system 32 or hypervisor 34, which may normally operate in non-secure region 86 and be feature-rich (alternatively, even if those providers themselves are trustworthy, the user may wish to protect themselves from the operating system 32 or hypervisor 34 being illegally accessed by an attacker). Also, while secure region 84 can be used for such user-provided applications that require secure processing, in practice, this causes problems for both the user providing code that requires a secure computing environment and the provider of existing code operating within secure region 84. For the provider of existing code operating within secure region 84, the area of attack for potential attacks against their code increases by the addition of arbitrary user-provided code within the secure region. This may be undesirable and thus it may be strongly recommended that users not be able to add code to secure region 84. On the other hand, a user providing code that requires a secure computing environment may find it difficult to audit and prove all of the separate code provided by different software providers operating within secure region 84 if the guarantee and proof of code operating in a particular area are required as a prerequisite for the user-provided code to execute processing, and may be reluctant to entrust all of the different code providers operating within secure region 84 with access to their data or code.This may limit the opportunity for third parties to provide more secure services.
[0074] Accordingly, as shown in FIG. 2, an additional area 88 called a realm area is provided and can be used by such code introduced by a user to provide a secure computing environment orthogonal to any secure computing environment associated with components operating in the secure area 24. In the realm area, the software to be executed can include several realms, and each realm can be isolated from other realms by a realm management module (RMM) 46 operating at exception level EL2. The RMM 46 can control the isolation between respective realms 42, 44 executing the realm area 88, for example, by defining access permissions and address mappings within the page table structure, in a similar way as the hypervisor 34 manages the separation between different components operating in the non-secure area 86. In this example, the realms include an application-level realm 42 executed at EL0 and a capsule application / operating system realm 44 executed across exception levels EL0 and EL1. It will be understood that it is not essential to support both EL0 and EL0 / EL1 type realms, and multiple realms of the same type can be established by the RMM 46.
[0075] Similar to the secure region 84, the realm region 88 has its own physical address space assigned to it. However, the realm region and the secure regions 88 and 84 can each access the non-secure PAS associated with the non-secure region 86, but the realm region and the secure regions 88 and 84 are orthogonal to the secure region 84 in the sense that they cannot access each other's physical address spaces. This means that the code executed in the realm region 88 and the secure region 84 have no dependency on each other. The code within the realm region only needs to trust the hardware, the RMM 46, and the code operating in the root region 82 that manages the switching between regions, which means that proof and assurance are more feasible. Proof enables a given software to require verification that the code installed on the device matches certain expected characteristics. This can be done by checking whether the hash of the program code installed on the device matches the expected value signed using a cryptographic protocol by a trusted party. The RMM 46 and the monitor code 29 can be proven, for example, by checking whether the hash of this software matches the expected value signed by a trusted party such as a silicon provider that manufactured the integrated circuit including the processing system 2, or an architecture provider that designed the processor architecture that supports region-based memory access control. Thereby, the user-provided codes 42 and 44 can verify whether the integrity of the region-based architecture is trustworthy before executing any secure or sensitive functions.
[0076] Thus, as indicated by the dotted lines showing the gaps within the non-secure regions where these processes would have been previously executed, the code associated with realms 42, 44 that would have been previously executed within non-secure region 86 can now be moved to realm regions where they can have stronger security guarantees, as their data and code can no longer be accessed by other code operating in non-secure region 86. However, due to the fact that realm region 88 and secure region 84 are orthogonal and thus cannot see each other's physical address spaces, this means that the provider of the code within the realm region need not trust the provider of the code within the secure region, and vice versa. The code within the realm region can simply trust the firmware that provides the monitor code 29 of root region 82 and RMM 46, which can be provided by the silicon provider, or the provider of the instruction set architecture supported by the processor. These providers may need to be inherently trusted from the start when the code is being executed on their devices, such that no other additional trust relationships with other operating system vendors, OEMs, or cloud hosts are required by the user in order for the user to be provided with a secure computing environment.
[0077] This is useful for various purpose applications and use cases, for example, including mobile wallets and payment applications, fraud and copyright infringement prevention mechanisms in games, operating system platform security extensions, secure virtual machine hosting, confidential computing, networking, or gateway processing for the Internet of Things. It will be understood that the user can find many other applications for which realm support is useful.
[0078] To support the security guarantees provided to the realm, the processing system can support a proof report function, and during startup or execution, measurements of the firmware image and configuration, such as the image and configuration of the monitor code, or the image and configuration of the RMM code, are taken. During execution, the content and configuration of the realm are measured, enabling the realm owner to trace back relevant proof reports to known implementations and guarantees and make a trust decision on whether it operates on that system.
[0079] As shown in Figure 2, a separate root region 82 for managing region switching is provided, which has its own separate root physical address space. Creating a root region and isolating resources from the secure region enables a more robust implementation even in a system that has only non-secure and secure regions 86, 84 and does not have a realm region 88, but can also be used in an implementation that supports the realm region 88. The root region 82 can be implemented using monitor software 29 provided (or guaranteed) by the silicon provider or architecture designer, and can be used to provide secure boot functionality, trusted boot measurements, system-on-chip configuration, debug control, and firmware update management for firmware components provided by other parties such as OEMs. The code for the root region can be developed, guaranteed, and deployed by the silicon provider or architecture designer without dependencies on the final device. In contrast, the secure region 84 can be managed by the OEM to implement specific platforms and security services. The management of the non-secure region 86 can be controlled by an operating system 32 that provides operating system services, while the realm region 88 is isolated from the existing secure software environment within the secure region 84 and at the same time enables the development of a new form of a trusted execution environment that can be dedicated to user or third-party applications.
[0080] Figure 3 schematically shows another example of the processing system 2 for supporting these techniques. Elements that are the same as in Figure 1 are denoted by the same reference numerals. Figure 3 shows more details of the address translation circuit 16 and includes a stage 1 memory management unit 50 and a stage 2 memory management unit 52. The stage 1 MMU 50 can be involved in either the conversion from a virtual address to a physical address (when the conversion is triggered by EL2 or EL3 code) or to an intermediate address (in an operating state where further stage 2 conversion by the stage 2 MMU 52 is required and the conversion is triggered by EL0 or EL1 code). The stage 2 MMU can convert the intermediate address to a physical address. The stage 1 MMU can be based on a page table controlled by the operating system for conversions starting from EL0 or EL1, a page table controlled by the hypervisor for conversions from EL2, or a page table controlled by the monitor code 29 for conversions from EL3. On the other hand, the stage 2 MMU 52 can be based on a page table structure defined by the hypervisor 34, the RMM 46, or the secure partition manager 14 depending on which region is being used. By separating the conversion into two stages in this way, the operating system can manage the address translation for itself and for applications under the assumption that they are the only operating systems running on the system, while the RMM 46, the hypervisor 34, or the SPM 40 can manage the isolation between different operating systems running within the same region.
[0081] As shown in FIG. 3, the address translation process using the address translation circuit 16 can return a security attribute 54 that, in combination with the current exception level 15 and the current region 14 (or security state), enables a section of a particular physical address space (identified by a PAS identifier or "PAS TAG") to be accessed in response to a given memory access request. The physical address and the PAS identifier can be looked up in a granularity protection table 56 that provides the granularity protection information described above. In this example, the PAS filter 20 is shown as a granular memory protection unit (GMPU) that verifies whether the selected PAS can access the requested physical address and, if so, allows the transaction to be passed to any cache 24 or interconnect 8 that is part of the system fabric of the memory system.
[0082] The GMPU 20 enables memory to be allocated to different address spaces while simultaneously providing a strong hardware-based isolation guarantee, offering not only spatial and temporal flexibility in how physical memory is allocated to these address spaces but also an efficient sharing scheme. As described above, the execution units within the system are logically divided into virtual execution states (regions or "worlds") in which there is one execution state (the root world), called the "root world", located at the highest exception level (EL3), and the root world manages the physical memory allocation to these worlds.
[0083] A single system physical address space is virtualized into multiple "logical" or "architectural" physical address spaces (PASs), and each such PAS is an orthogonal address space with independent coherence attributes. The system physical address is mapped to a single "logical" physical address space by extending it with a PAS tag.
[0084] A given world is permitted access to a subset of the logical physical address space. This is implemented by a hardware filter 20 that can be attached to the output of the memory management unit 16.
[0085] The world defines the security attributes (PAS tags) of access using the fields of the translation table descriptors of the page table used for address translation. The hardware filter 20 has access to a table (granular protection table 56, or GPT) that defines granular protection information (GPI) for each page within the system physical address space, and the granular protection information indicates the PASTAG it is associated with, and (optionally) other granular protection attributes.
[0086] The hardware filter 20 checks the world ID and security attributes against the GPI of the granule to determine whether access can be permitted, thus forming a granular memory protection unit (GMPU).
[0087] The GPT 56 can be present, for example, in on-chip SRAM or off-chip DRAM. When stored off-chip, the GPT 56 can be integrity protected by an on-chip memory protection engine that can use mechanisms for encryption, integrity, and freshness to maintain the security of the GPT 56.
[0088] By positioning the GMPU 20 on the requester side of the system (e.g., on the MMU output) rather than on the completer side, the interconnect 8 is allowed to continuously hash / stripe pages across multiple DRAM ports while enabling the assignment of access permissions at the page granularity.
[0089] Since the transaction propagates through the entire system fabric 24, 8 until it reaches the location defined as the point of the physical aliasing point 60, the PASTAG remains tagged. This allows the filter to be placed on the master side without weakening the security guarantee compared to slave - side filtering. As the transaction propagates through the system, the PAS TAG can be used as a detailed security mechanism for address isolation. For example, the cache can add the PAS TAG to the address tags in the cache to prevent accesses made with an incorrect PAS TAG for the same PA from hitting the cache, thereby improving side - channel resistance. The PAS TAG can also be used as a context selector for a protection engine attached to a memory controller that encrypts data before writing to external DRAM.
[0090] The physical aliasing point (PoPA) is the location within the system where the PAS TAG is stripped and the address returns from the logical physical address to the system physical address. The PoPA can be located below the cache on the completer side of the system where accesses to physical DRAM are made (using the encrypted context resolved via the PAS TAG). Alternatively, it may be located above the cache to simplify the system implementation at the expense of weakening security.
[0091] At any point in time, the world can request a page transition from one PAS to another. The request is made at EL3 to the monitor code 29 and examines the current state of the GPI. EL3 may allow only a specific set of transitions (e.g., from a realm PAS to a secure PAS but not from a non-secure PAS to a secure PAS) to occur. To provide a clean transition, a new instruction "delete data and invalidate up to the physical aliasing point" is supported by the system and EL3 can issue this before transitioning the page to the new PAS. This ensures that any residual state associated with the previous PAS is flushed from any caches upstream of PoPA60 (closer to the requester side).
[0092] Another feature that can be achieved by attaching the GMPU20 on the master side is the efficient sharing of memory between worlds. It may be desirable to allow a subset of N worlds to have shared access to a physical granule while preventing other worlds from accessing it. This can be achieved by adding "limited sharing" semantics to the granule protection information and forcing it to use a specific PAS TAG. As an example, the GPI can indicate that a physical granule can be accessed only by the "realm world" 88 and the "secure world" 84 while the PAS TAG of the secure PAS 84 is tagged on the physical granule.
[0093] The above example of a feature results in a rapid change in the visibility characteristics of a particular physical granule. Consider the case where each world has a private PAS assigned to it that is accessible only to that world. For a particular granule, a world can request to make it visible to the non-secure world at any point in time without changing the association of the PAS by changing its GPI from "exclusive" to "limited sharing with the non-secure world". In this way, the visibility of that granule can be increased without the need for costly cache maintenance or data copy operations.
[0094] FIG. 4 shows the concept of aliasing each physical address space on a physical memory provided in hardware. As described above, each of regions 82, 84, 86, 88 has its own respective physical address space 61.
[0095] At the time when a physical address is generated by address translation circuit 16, the physical address has a value within a specific numerical range 62 supported by the system, which is the same regardless of which physical address space is selected. However, in addition to generating the physical address, address translation circuit 16 can also select a specific physical address space (PAS) based on the current region 14 and / or the information in the page table entry used to derive the physical address. Alternatively, instead of address translation circuit 16 that performs the selection of the PAS, an address translation circuit (e.g., MMU) can output a physical address and information derived from a page table entry (PTE) used for the selection of the PAS, and then this information can be used by a PAS filter or GMPU 20 to select the PAS.
[0096] The selection of the PAS for a given memory access request can be limited according to the rules defined in the following table, depending on the current region in which processing circuit 10 issues the memory access request.
[0097]
Table 1
[0098] Therefore, when the PAS filter 20 outputs a memory access request to the system fabric 24, 8 (assuming it has passed any filtering checks), the memory access request is associated with a physical address (PA) and a selected physical address space (PAS).
[0099] From the perspective of memory system components (caches, interconnects, snooping filters, etc.) that operate before the physical alias (PoPA) point 60, each physical address space 61 is considered a completely different address range corresponding to different system locations within the memory. This means that from the perspective of pre-PoPA memory system components, the address range identified by a memory access request is actually four times the size of the range 62 that can be output in address translation. This is because, in effect, the PAS identifier is treated as an additional address bit alongside the physical address itself, so that depending on the selected PAS, the same physical address PAx can be mapped to several alias physical addresses 63 within different physical address spaces 61. Although these alias physical addresses 63 all actually correspond to the same memory system location implemented in physical hardware, pre-PoPA memory system components treat the alias addresses 63 as different addresses. Thus, if there is a pre-PoPA cache or snooping filter that allocates an entry to such an address, the alias address 63 will be mapped to different entries with different cache hit / miss decisions and different coherence management. This reduces the likelihood or effectiveness of an attacker using a cache or coherence side channel as a mechanism to probe the operation of other areas.
[0100] The system may include two or more PoPAs 60 (e.g., as shown in FIG. 14 discussed below). In each PoPA 60, the aliased physical address is folded into a single non-aliased address 65 within the system physical address space 64. The non-aliased address 65 is provided downstream of any PoPA post-component, such that the system physical address space 64 that actually identifies the memory system location is again the same size as the range of physical addresses that can be output in address translation performed on the requester side. For example, in the PoPA 60, the PAS identifier may be stripped from the address, and for downstream components, the address can be identified using only the physical address value without specifying the PAS. Alternatively, if some completer-side filtering of memory access requests is desired, the PAS identifier may still be provided downstream of the PoPA 60 but may not need to be interpreted as part of the address. As a result, the same physical address appearing in different physical address spaces 60 will be interpreted to refer to the same memory system location downstream of the PoPA. However, the supplied PAS identifier may still be used to perform a completer-side security check.
[0101] FIG. 5 shows a way in which the system physical address space 64 can be divided into chunks allocated for access within the physical address space 61 on a particular architecture using the granularity protection table 56. The granularity protection table (GPT) 56 defines which portions of the system physical address space 65 are accessible from the physical address space 61 on each architecture. For example, the GPT 56 may include several entries each corresponding to a granularity of physical addresses of a particular size (e.g., 4K pages), and may define the allocated PAS for that granularity, which may be selected from among the non-secure, secure, realm, and root regions. By design, if a particular granularity or set of granularities is allocated to a PAS associated with one of the regions, it can only be accessed within the PAS associated with that region and cannot be accessed within the PAS of other regions. However, note that the root region 82 can still access that granularity of the physical address by specifying, in its page table, PAS selection information to ensure that the virtual address associated with the page mapped to that area of the physically addressed memory is translated to a physical address within the secure PAS instead of the root PAS, even though the granularity allocated to the secure PAS cannot be accessed from within the root PAS. Thus, sharing of data between regions (to the extent permitted by the access rules defined in the foregoing table) can be controlled at the time of selecting the PAS for a given memory access request.
[0102] However, in some implementations, in addition to enabling access to the granularity of physical addresses within the assigned PAS as defined by the GPT, the GPT can use other GPT attributes to mark a certain region of the address space (e.g., an address space associated with a lower or orthogonal privilege region where it is usually not permitted to select the PAS assigned to the access requests of that region) as being shared with another address space. This can facilitate temporarily sharing data without the need to change the PAS assigned to a given granularity. For example, in FIG. 5, the region 70 of the realm PAS is defined within the GPT to be assigned to the realm region, and the non-secure region 86 is normally inaccessible from the non-secure region 86 because it cannot select the realm PAS for its access requests. The non-secure region 26 cannot access the realm PAS, so non-secure code could not normally view the data in region 70. However, if the realm desires to temporarily share some of its data within the assigned region of memory with the non-secure region, it can request that the monitor code 29 operating in the root region 82 update the GPT 56 to indicate that region 70 is shared with the non-secure region 86, thereby making region 70 accessible from the non-secure PAS shown on the left side of FIG. 5 without the need to change which region is assigned to region 70. When the realm region indicates that a region of its address space is to be shared with the non-secure region, a memory access request issued from the non-secure region and targeting that region can initially specify the non-secure PAS, but the PAS filter 20 can remap the request's PAS identifier to specify the realm PAS instead. Thereby, downstream memory system components can handle the request as if it had been issued from the realm region from the beginning. This sharing can improve performance because the operations for assigning different regions to a particular memory region can be more performance-intensive, involving a higher degree of cache / TLB invalidation and / or zeroing of data in memory or copying of data between memory regions.This may not be justified if sharing is expected to be only temporary.
[0103] Figure 6 is a flowchart showing how the current operating region can be determined, which can be executed by the processing circuit 10, or by the address translation circuit 16 or the PAS filter 20. At step 100, it is determined whether the current exception level 15 is EL3, and if so, then at step 102, it is determined that the current region is the root region 82. If the current exception level is not EL3, then at step 104, it is determined that the current region is one of the non-secure, secure, and realm regions 86, 84, 88 as indicated by at least two region indication bits 14 in the EL3 control register of the processor (since the root region is indicated by the current exception level being EL3, it is not necessary to have the code of the region indication bit 14 corresponding to the root region, and thus the code of at least one region indication bit can be reserved for other purposes). The EL3 control register is writable when operating at EL3 and cannot be written from other exception levels EL2 to EL0.
[0104] FIG. 7 shows an example of a page table entry (PTE) in a page table structure used by the address translation circuit 16 for mapping from a virtual address to a physical address, mapping from a virtual address to an intermediate address, or mapping from an intermediate address to a physical address (depending on whether the conversion is being performed in an operating state where stage 2 conversion is necessary in the first place, and if stage 2 conversion is necessary, which of stage 1 and stage 2 the conversion is). In general, a given page table structure may be defined as a multi-level table structure where the first level of the page table is implemented as a page table tree identified based on a base address stored in the processor's translation table base address register, and the index for selecting a particular level 1 page table entry within the page table is derived from a subset of the bits of the input address for which the translation look-up is to be performed (the input address can be the virtual address for stage 1 conversion of the intermediate address for stage 2 conversion). The level 1 page table entry can be a "table descriptor" 110 that provides a pointer 112 to the next level of the page table, from which further page table entries can be selected based on a further subset of the bits of the input address. Finally, after one or more look-ups to successive levels of the page table, block or page descriptors PTEs 114, 116, 118 that provide an output address 120 corresponding to the input address can be identified. The output address can be an intermediate address (for stage 1 conversion in an operating state where further stage 2 conversion is also performed), or a physical address (for stage 2 conversion, or stage 1 conversion when stage 2 is not required).
[0105] To support the above-described separate physical address spaces, the page table entry format can specify some additional state for use in physical address space selection, in addition to the next level page table pointer 112 or output address 120, and any attributes 122 for controlling access to the corresponding block of memory.
[0106] For the table descriptor 110, the PTEs used by any region other than the non-secure region 86 include a non-secure table indicator 124 that indicates which level of page table should be accessed from either the non-secure physical address space or the physical address space of the current region. This helps to facilitate more efficient management of the page table. In many cases, the page table structure used by the root, realm, or secure region 24 may only need to define special page table entries for a portion of the virtual address space. For the other portion, the same page table entries used by the non-secure region 26 can be used. Thus, by providing the non-secure table indicator 124, a higher-level realm / secure dedicated table descriptor can be provided in the page table structure, while at some point in the page table tree, the root realm or secure region can switch to using page table entries from the non-secure region for portions of the address space where higher security is not required. Other page table descriptors in other parts of the page table tree can still be fetched from the associated physical address space related to the root, realm, or secure region.
[0107] On the one hand, the block / page descriptors 114, 116, 118 may include physical address space selection information 126 depending on which region they are associated with. The non-secure block / page descriptor 118 used within the non-secure region 86 does not include any PAS selection information because the non-secure region can only access the non-secure PAS. However, for other regions, the block / page descriptors 114, 116 include PAS selection information 126 that is used to select the PAS for translating the input address. For the root region 22, the EL3 page table entry may have PAS selection information 126 that includes at least 2 bits to indicate the PAS associated with any one of the four regions 82, 84, 86, 88 as the selected PAS for which the corresponding physical address is translated. In contrast, for the RELM and secure regions, the corresponding block / page descriptor 116 only needs to include 1-bit PAS selection information 126, whereby for the RELM region, a selection is made between the RELM and the non-secure PAS, and for the secure region, a selection is made between the secure PAS and the non-secure PAS. To improve the efficiency of circuit implementation and avoid increasing the size of the page table entry, for the RELM and secure regions, the block / page descriptor 116 can encode the PAS selection information 126 at the same position within the PTE regardless of whether the current region is RELM or secure, so that the PAS selection bits 126 can be shared.
[0108] Accordingly, FIG. 8 is a flowchart showing a method of selecting a PAS based on information 124, 126 from the block / page PTE that is used to generate a physical address for the current region and a given memory access request. The PAS selection can be performed by the address translation circuit 16, or in the case where the address translation circuit sends the PAS selection information 126 to the PAS filter 20, it can be performed by a combination of the address translation circuit 16 and the PAS filter 20.
[0109] In step 130 of FIG. 8, the processing circuit 10 issues a memory access request by designating a given virtual address (VA) as the target VA. In step 132, the address translation circuit 16 looks up any page table entry (or cache information derived from such a page table entry) in its TLB 18. If any required page table information is not available, the address translation circuit 16 starts a page table walk to memory to fetch the required PTE (potentially stepping through each level of the page table structure and / or multiple stages of address translation to obtain the mapping from VA to intermediate physical address (IPA) and from IPA to PA, requiring a series of memory accesses). Any memory access request issued by the address translation circuit 16 in the page table walk operation itself can be subject to address translation and PAS filtering. Thus, note that the request received in step 130 can be a memory access request issued to request a page table entry from memory. When relevant page table information is identified, the virtual address is translated to a physical address (possibly in two stages via the IPA). In step 134, the address translation circuit 16 or the PAS filter 20 determines which region is the current region using the technique shown in FIG. 6.
[0110] If the current region is a non-secure region, in step 136, the output PAS selected for this memory access request is a non-secure PAS.
[0111] If the current region is a secure region, in step 138, the output PAS is selected based on the PAS selection information 126 included in the block / page descriptor PTE that provided the physical address, and the output PAS is selected as either a secure PAS or a non-secure PAS.
[0112] If the current area is the Realm area, in step 140, the output PAS is selected based on the PAS selection information 126 included in the block / page descriptor PTE from which the physical address was derived. In this case, the output PAS is selected as either the Realm PAS or the non-secure PAS.
[0113] If it is determined in step 134 that the current area is the Root area, in step 142, the output PAS is selected based on the PAS selection information 126 within the root block / page descriptor PTE 114 from which the physical address was derived. In this case, the output PAS is selected as any of the physical address spaces associated with the Root, Realm, Secure, and non-secure areas.
[0114] FIG. 9 shows an example of a granularity protection data block 150 that provides granularity protection entries 152 corresponding to some of the granularities of the physical address space. In this example, the 64-bit granularity protection data block 150 includes 16 granularity protection entries (GPEs) 152, each of which provides 4-bit granularity protection information (GPI).
[0115] Within the 4-bit encoding space provided to each GPE (which provides 16 possible encodings), a specific number of encodings can be defined as valid encodings to indicate the valid options for the granularity protection information. For example, 7 of the 16 available encodings can be allocated to encode the following valid granularity protection information items. · Permit access to the Root PAS and prohibit access to other PASs. · Permit access to the Secure PAS and prohibit access to other PASs. · Permit access to the Realm PAS and prohibit access to other PASs. · Permit access to the non-secure PAS and prohibit access to other PASs. · If the current exception level is EL3, access is permitted to any PAS, and if the current exception level is EL0, EL1, or EL2, access is prohibited to any PAS. · Access to any PAS is permitted. · Access is prohibited for all PASs (this encoding may also be used if the GPE is invalid). Which specific 4-bit encoding of the GPI is used for each of these options is an arbitrary design choice for the instruction set architecture designer, and it should be understood that any mapping of each encoding to the 4-bit value of the GPI field can be used.
[0116] Also, these are only some of the possible options for valid granularity protection information, and it should be understood that in some architecture implementations, other options can be supported, or some of these options can be omitted. For example, if the "restricted sharing" semantics are implemented as described above, additional encoding can be assigned to indicate sharing. For example, the encoding can indicate that access to the region PAS and the non-secure PAS is permitted (the granularity is usually assigned to the realm PAS, but the realm region has previously requested to the root region to indicate that the granularity is shared with the non-secure PAS).
[0117] In general, there may be some preliminary invalid encodings that are not used to indicate valid options for the GPI. These preliminary invalid encodings may occur as an intentional design choice because the number of valid options is not an exact power of two, or because the instruction set architecture designer leaves room for future expansion, or to increase the probability of detecting a loss of data integrity based on an analysis similar to the quantitative analysis provided below.
[0118] In the above example, 9 out of the 16 available encodings are invalid. This can be used to provide a data integrity check as described later.
[0119] The data integrity check determines, for each GPE in the loaded block, whether the GPE has one of the valid GPI encodings or one of the invalid GPI encodings. If any of the GPEs in the block has an invalid GPI encoding (even if it is not the GPE required to process the current memory access request), a data integrity failure is notified. If all GPEs have valid GPI encodings, the data is assumed to be valid and trustworthy.
[0120] If a data block 150 that provides several granule protection entries is loaded from an area of memory that may be tampered with by an attacker, the granule protection data block can be assumed to be encrypted before being written to that area of memory. The encryption primitive used for encryption / decryption can operate on a block of data of a size corresponding to the granule protection data block 150. If an attacker flips a single bit in the encrypted data in memory (DRAM), for example by performing a row hammer type attack, the diffusion property of the cryptographic function normally used for memory encryption means that changing a single bit in the ciphertext stored in memory may result in a decrypted plaintext that is randomized by the decryption function when the modified encrypted data is later read from memory. Therefore, the decrypted data after the attacker has tampered with the ciphertext can be regarded as a random bit string.
[0121] Therefore, considering that the output of decryption can be regarded as virtually random when an attacker tampers with the ciphertext data stored in memory, the probability that all of the resulting plaintext GPEs after the attacker's modification will ultimately have a valid encoding of GPI is as follows (evaluated for 16 4-bit GPEs in a 64-bit GPT block):
[0122] [Table 2] Thus, in the above example using nine invalid encodings, if an attacker modifies even one bit of the data stored in memory, the probability of being able to avoid the data integrity check is only one in 555074.
[0123] This probability varies depending on the number of invalid encodings and the number of GPEs within one data block that is loaded from memory in one transaction (and that undergoes encryption as a common operation across the entire block).
[0124] More generally, for an N-bit GPI per GPE, X out of 2 encodings of the GPI being invalid, and Y GPEs per block, the probability of a random output that generates no faults is N Therefore, the architecture designer can choose N, X, and Y as needed to provide an acceptable level of risk. For example, if one in 555074 in the above example is considered too high a risk, this can be adjusted by providing one additional redundant bit per GPE to increase the number of invalid encodings X, or by grouping together a larger number Y of GPEs within each granularity protection data block. For example, if instead of dealing with 16 GPEs, encryption and data integrity checks are applied to an entire 64-byte cache line containing 128 4-bit GPEs, this increases the probability of establishing no faults from ((16 - X) / 16)^16 to ((16 - X) / 16)^128, so that even if X is 2 and only 2 out of 16 values are invalid, the probability of no fault occurring is reduced to one in 26 million.
[0125]
Equation
[0126] Obviously, the exact number of invalid encodings per data block and the number of GPEs are design choices, but this demonstrates the principle that a simple check of whether the encoding of each GPE within the loaded block is invalid may be sufficient to enforce data integrity. Thus, as shown in FIG. 1, the PAS filter 20 includes a granular protection entry load circuit (GPT load circuit) 21 that loads the granular protection data block 150 from memory when the GPEs required to check the memory access to the target PA are not yet available in the granular protection information cache 22, and an integrity check circuit 23 for performing a data integrity check on the loaded granular protection data block 150. For example, the integrity check circuit 23 can include a set of boolean logic gates that check whether the value of each GPE within the loaded block is one of the invalid encodings and trigger signaling of an error if any of the GPEs have an invalid encoding. This avoids the cost of calculating the authentication code and maintaining the integrity tree of the metadata for verifying integrity, resulting in a significant performance improvement.
[0127] In an example where an integrity scheme based on checking valid / invalid encodings is valid, the software that manages the GPT table (executed in the root region) may need to ensure that not only the GPEs corresponding to the granules of the physical addresses currently in use, but also all GPEs within the block have valid encodings. This requirement may not have to be enforced by circuitry within the hardware (since the root region code operates at EL3 and is trusted code proven using the proof mechanism described above, there may be no checks on writing GPT data to memory as to whether the GPE encoding is valid, and it can be trusted to set the GPT data in an appropriate manner).
[0128] Figure 10 shows an alternative technique for data integrity verification. In this example, block 150 of GPE152 is associated with a signature 154 that is calculated when written to memory by applying a hash function or other authentication code generation function to the contents 152 of the GPT data block 150. In this example, the signature 154 is stored with the GPE in memory (so that the signature is loaded from memory when the block is read), but the signature 154 could also be stored in a different storage location from the associated block of GPE152. When reading a block of GPE from memory, after decrypting GPE152 (and the signature if stored as part of the same block), a signature check value is calculated from the read value of GPE152 and compared to the previously stored integrity signature 154. If a mismatch is detected between the previously stored signature 154 and the signature check value, a fault is notified to indicate a loss of data integrity. If the signature 154 matches the signature check value, the GPT data block 150 can be cached by the filtering circuit 20 or used for filtering memory accesses. In this approach, additional integrity metadata may need to be stored to enable signature verification to protect against an attacker replacing GPE152 and signature 154 with different GPE values and different matching signatures if the signature 154 is not stored in memory within a trusted boundary (a trusted boundary representing a boundary beyond which data can be considered vulnerable to attack). For example, a signature tree may be used to combine the signatures 154 of different GPT data blocks 150, and each tree node protects the signature within the tree node at a lower level until traced to the root of the tree which can ultimately correspond to a signature stored within a trusted boundary. In this case, there may be a greater performance overhead when verifying the integrity of GPT block 150 because a sequence of verifications traversing the tree needs to be performed to confirm that each of the intervening signatures is valid before the integrity of the GPT block can be confirmed.
[0129] FIG. 11 is a flowchart showing the filtering of memory access requests by the filtering circuit 20. In step 170, a memory access request is received from the address translation circuit 16 (assuming that the permission check performed by the address translation circuit has already passed), a target PA is specified, and it is associated with the selected PAS. As described above, the selected PAS can be indicated by the address translation circuit that provides an explicit identifier indicating the selected PAS, or by the address translation circuit 16 that transfers the PAS selection information 126 read from the page table. The PAS selection information read from the page table can be used to determine the selected PAS in the filtering circuit 20 together with the indication of the current area of the processing circuit 10.
[0130] In step 172, filtering circuit 20 searches the granule protection information cache 22 to determine whether the GPE required for target PA is already accessible. If not, in step 174, granule protection entry load circuit 21 issues a load request to request from memory a granule protection data block including the target GPE corresponding to the granule of PA including target PA. When GPT is a table with a single-level linear index, the address of the granule protection data block can be determined from the GPT base address and the offset derived from target PA. When GPT is a multi-level structure similar to the page table structure used by address translation circuit 16, the address of the first-level GPT entry that functions as a table descriptor providing a pointer to a further level of granule protection table can be derived using the GPT base address and target PA, and by walking through one or more further levels of granule protection tables, the granule protection data block can finally be located in memory. Filtering circuit 20 (GMPU) can be trusted to issue correct GPT memory accesses, so the memory access requests issued during GPT walk do not themselves require a PAS check (in contrast to the page table walk accesses issued by address translation circuit 16, which are subject to PAS check and thus can cause a GPT lookup similar to a normal memory access request issued by the processing circuit).
[0131] If any block of granule protection information is stored in a protected area of memory that is subject to memory encryption, when reading information from memory decryption, as shown in FIG. 13 described later, it is applied by memory protection engine 330 (memory encryption circuit) before providing the decrypted data to filtering circuit 20.
[0132] In an implementation that does not support caching of granule protection entries, step 172 can be omitted and the method can proceed directly from step 170 to step 174 (in this case, all memory access requests may require granule protection information to be loaded).
[0133] In step 176, for any received granule protection data block obtained from the memory system as part of the GPT lookup, the integrity check circuit 23 performs an integrity check to determine whether there is a risk that the received granule protection data block has been tampered with. As described above, in some implementations, this check can simply involve checking whether each GPE within the obtained granule protection data block has a valid encoding, and if any GPE within the granule protection data block has an invalid encoding, the check can be considered to have failed. Alternatively, the data integrity check may check whether the integrity signature check value derived from the obtained block matches the previously stored integrity signature associated with the block, and may be considered to have failed if there is a mismatch between the check value and the previously stored integrity signature.
[0134] In step 178, it is determined whether the data integrity check has passed or failed. If the data integrity check fails, in step 180, the integrity check circuit 23 triggers the signaling of a granularity protection fault associated with a fault type or syndrome information indicating that the fault was caused by a data integrity check violation. The fault is signaled such that it is the type of fault that is processed by an exception handler code executed at EL3, and as a result, the root region code processes the fault. For example, the root region code can trigger other actions to prevent the system from continuing to function in the presence of a system reset or a potential attack. The root region code can also cause, for example, the system operator or provider of the software being executed to be notified of the fault. The exact actions taken in response to the signaling of the fault can vary depending on the requirements of a particular implementation.
[0135] On the other hand, if the data integrity check passes, in step 182, the obtained granularity protection data block can be cached in the granularity protection information cache 22, and in step 184, the GPE corresponding to the target PA is extracted from the obtained granularity protection data block and can be used to check whether the memory access request received in step 170 is permitted to continue. If the target PA hits in the granularity protection information cache 22 in step 172, steps 174 to 182 can be omitted, and instead, the method proceeds directly to step 184 using the target GPE obtained from the granularity protection information cache 22.
[0136] Regardless of whether the target GPE has already been cached or has been retrieved from memory and passed the data integrity check, in step 184, the filtering circuit 184 determines, based on the GPI of the target GPE, whether the target GPE is invalid or whether the selected PAS associated with the memory access request indicates that it is not a permitted PAS for the granularity of the corresponding physical address. If the target GPE is valid and the selected PAS indicates that it is a permitted PAS, in step 186, the filtering circuit 20 permits the memory access request to proceed and passes the memory access request, along with indications of the PA and PAS associated with the memory access request, to the cache 24 or the interconnect 8. If the target GPE is invalid or the selected PAS indicates that it is not a permitted PAS, in step 188, the filtering circuit 20 triggers the signaling of a granularity protection fault associated with fault type or syndrome information indicating that the memory access request has been blocked due to an access violation that does not meet the requirements indicated by the target GPE. Again, the granularity protection fault is signaled to be processed by program code operating in the root region at EL3, but is distinguished by the fault type or syndrome information from the fault signaled in step 180 due to a data integrity violation, such that the root region code can take different actions in response to the fault.
[0137] Figure 12 summarizes the operation of the address translation circuit 16 and the PAS filter. PAS filtering 20 can be regarded as an additional stage 3 check that is performed after the stage 1 (and optionally stage 2) address translation performed by the address translation circuit. Also, the EL3 translation is based on a page table entry that provides selection information (labeled NS, NSE in the example of Figure 12) based on a 2-bit address, while the single-bit selection information "NS" is also noted to be used to select the PAS in other states. The security state shown in Figure 12 as input to the granularity protection check refers to the region ID that identifies the current region of the processing element 4.
[0138] Figure 13 shows a more detailed example of a data processing system that can implement some of the techniques described above. Elements that are the same as in previous examples are indicated by the same reference numbers. In the example of Figure 13, in addition to the processing circuit 10, address translation circuit 16, TLB 18, and PAS filter 20, the cache 24 is shown in more detail, including a level 1 instruction cache, a level 1 data cache, a level 2 cache, and optionally a shared level 3 cache 24 shared between processing elements, showing the processing element 4 in more detail. The interrupt controller 300 can control the processing of interrupts by each processing element.
[0139] As shown in FIG. 13, the processing element 4 that can execute program instructions to trigger access to the memory is not the only type of requester device that can include the requester-side PAS filter 20. In other examples, a system MMU 310 (provided to provide an address translation function for requester devices 312, 314 that do not support their own address translation function, such as on-chip devices 312 like a network interface controller or a display controller, or off-chip devices 314 that can communicate with the system via a bus) can include a PAS filter 20 to perform a requester-side check of GPT entries, similar to that for the PAS filter 20 within the processing element 4. Other requester devices can include a debug access port 316 and a control processor 318, and these can also have an associated PAS filter 20, thereby checking whether memory accesses issued by the requester devices 316, 318 to a specific physical address space are permitted under the PAS assignment defined by the GPT 56. Although not explicitly shown in FIG. 13, any of the PAS filters 20 associated with these devices can also include the GPT load circuit 21 and the integrity check circuit 23 described above. If any of these devices issues a memory access request specifying a physical address (which can occur when the MMU is downstream in that environment), the PAS selected for these devices is, by default, considered to be the PAS associated with the current region (i.e., a non-secure PAS if the current region is non-secure, a secure PAS if the current region is secure, a realm PAS if the current region is a realm, and a root PAS if the current region is a root).
[0140] The interconnect 8 is shown in more detail in FIG. 13 as a coherent interconnect 8, which also, similar to the routing fabric 320, includes a snooping filter 322 for managing coherence between caches 24 within each processing element, and one or more system caches 324 capable of caching shared data shared among requesting devices. The snooping filter 322 and the system cache 324 can be located upstream of the PoPA 60, and thus, using the PAS identifiers selected by the MMUs 16, 310 for a particular master, their entries can be tagged. Requesting devices 316, 318 not associated with an MMU are by default always assumed to issue requests for a particular region such as a non-secure region (or, if they are trusted, the root region).
[0141] FIG. 13 shows, as if they were referring to different address locations, the interconnect 8 and a memory protection engine (MPE) 330 provided between the interconnect 8 and a given memory controller 6 for controlling access to off-chip memory 340, as another example of a PoPA pre-component that handles aliased physical addresses within each PAS. The MPE 330 can be responsible for encrypting data written to the off-chip memory 340 to maintain confidentiality and decrypting the data when read (this can include a GPT block including a GPE as described above). Also, the MPE can generate integrity metadata when writing data to memory and verify whether the data has changed using the metadata when the data is read from the off-chip memory, thereby preventing tampering of the data stored in the off-chip memory. When encrypting data or generating a hash of memory integrity, different keys can be used depending on which physical address space is being accessed, even if accessing an aliased physical address that actually corresponds to the same location within the off-chip memory 340. This improves security by further isolating data associated with different operating regions.
[0142] In this example, since PoPA60 is between the memory protection engine 330 and the memory controller 6, when a request reaches the memory controller 6, the physical address is no longer treated as being mapped to different physical locations within the memory 340 depending on the physical address space in which they are accessed.
[0143] FIG. 13 shows another example of the completer device 6 that can be a peripheral bus or a non - coherent interconnect used to communicate with areas of the peripheral device 350 or the on - chip memory 360 (e.g., implemented as a static random access memory (SRAM)). Also, the peripheral bus or non - coherent interconnect 6 can be used to communicate with secure elements 370 such as an encryption unit that executes encryption processing, a random number generator 372, or certain fuses 374 that store information statically embedded in hardware. Also, various power / reset / debug controllers 380 can be accessible via the peripheral bus or non - coherent interconnect 6.
[0144] For on-chip SRAM 360, it may be useful to provide a PAS filter 400 on the slave side (the completer side). The PAS filter 400 can perform completer-side filtering of memory accesses based on completer-side protection information that defines which physical address spaces are accessible to a given block of physical addresses. This completer-side information can be defined more coarsely than the GPT used by the requester-side PAS filter 20. For example, the slave-side information can simply indicate for other regions that an entire SRAM unit 361 may be dedicated for use by the realm area, another SRAM unit 362 may be dedicated for use by the root area, and so on. Thus, relatively coarsely defined blocks of physical addresses can be addressed to different SRAM units. This completer-side protection information can be statically defined by loading bootloader code in the information for the completer-side PAS filter at startup that cannot be changed during execution. Therefore, it is not as flexible as the GPT used by the requester-side PAS filter 20. However, for use cases where the partitioning of each region of the physical address into specific areas accessible is known at startup and does not require fine-grained partitioning without change, it may be more efficient to use the slave-side PAS filter 400 instead of the requester-side PAS filter 20. This is because it can eliminate the power and performance cost on the requester side of obtaining GPT entries and comparing the assigned PAS and shared attribute information with the information for the current memory access request. Also, if a pass-through indicator can be shown for the top-level GPT entry (or other table descriptor entry at a level other than the final level) in the multi-level GPT structure, access to further levels of the GPT structure (which can be performed to find more fine-grained information regarding the assigned PAS for requests that receive requester-side checks) can be avoided for requests targeting one of the regions of physical addresses mapped to the on-chip memory 360 monitored by the completer-side PAS filter 400.
[0145] Therefore, supporting a hybrid approach that enables both the requester side and the completer side of the protection information to be checked can be useful for performance and power efficiency. The system designer can define which approach to take for a particular memory region.
[0146] FIG. 14 shows a simulator implementation that can be used. The above-described embodiments implement the present invention in terms of an apparatus and method for operating specific processing hardware that supports the technology, but it is also possible to provide an instruction execution environment according to the embodiments described herein that is implemented using a computer program. Such a computer program is often referred to as a simulator as long as the computer program provides a software-based implementation of a hardware architecture. Various simulator computer programs include binary translators including emulators, virtual machines, models, and dynamic binary translators. Typically, an implementation form of a simulator can optionally execute a host operating system 420 that supports a simulator program 410 and be executed by a host processor 430. In some configurations, there may be multiple layers of simulation between the hardware and the provided instruction execution environment, and / or there may be multiple different instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations that execute at a reasonable speed, but such an approach may be justified in some situations, for example, when it is desired to execute native code for another processor for reasons of compatibility or reuse. For example, a simulator implementation may provide an instruction execution environment having additional features not supported by the host processor hardware, or may typically provide an instruction execution environment associated with a different hardware architecture. An overview of simulation is described in "Some Efficient Architecture Simulation Techniques", Robert Bedichek, 1990 Winter USENIX Conference, pages 53-63.
[0147] Previously, embodiments have been described with reference to specific hardware configurations or functions. However, in simulated embodiments, equivalent functions can be provided by appropriate software configurations or functions. For example, a specific circuit may be implemented as computer program logic in a simulated embodiment. Similarly, memory hardware such as registers or caches may be implemented as software data structures in simulated embodiments. In configurations where one or more of the hardware elements referred to in the foregoing embodiments are present in host hardware (e.g., host processor 430), some simulated embodiments may use the host hardware where appropriate.
[0148] The simulator program 410 may be stored in a computer-readable storage medium (which may be a non-transitory medium) and provides a program interface (instruction execution environment) for the target code 400 (which may include an application, an operating system, and a hypervisor), and is the same as the interface of the hardware architecture modeled by the simulator program 410. Accordingly, the program instructions of the target code 400 may be executed from within the instruction execution environment using the simulator program 410. Thus, a host computer 430 that does not actually have the hardware features of the foregoing apparatus 2 can emulate these features. This can be useful for allowing testing of the target code 400 developed for a new version of a processor architecture, for example, by being executed within a simulator that is executed on a host device that does not support the new version of the processor architecture, before a hardware device that actually supports that architecture becomes available.
[0149] The simulator code emulates the behavior of the processing circuit 10, including, for example, instruction decoding program logic that decodes the instructions of the target code 400, and maps the instructions to a sequence of corresponding instructions within the native instruction set supported by the host hardware 430, and includes processing program logic 412 that executes the functions corresponding to the decoded instructions. The processing program logic 412 also simulates the processing of code at different exception levels and regions as described above. The register emulation program logic 413 maintains a data structure within the host address space of the host processor that emulates the architectural register state defined according to the target instruction set architecture associated with the target code 400. Thus, instead of such architectural state being stored in the hardware registers 12 as in the embodiment of FIG. 1, it is instead stored in the memory of the host processor 430, and the register emulation program logic 413 maps the register references of the instructions of the target code 400 to the corresponding addresses for obtaining the emulated architectural state data from the host memory. This architectural state may include the aforementioned current region indication 14 and current exception level indication 15.
[0150] The simulation code includes address translation program logic 414 and filtering program logic 416 that respectively emulate the functions of the address translation circuit 16 and the PAS filter 20 with reference to the same page table structure and GPT56 as described above. Thus, the address translation program logic 414 converts the virtual address specified by the target code 400 into a simulated physical address within one of the PASs (referring to the physical location in memory from the perspective of the target code), but in fact these simulated physical addresses are mapped onto the (virtual) address space of the host processor by the address space mapping program logic 415. The filtering program logic 416 performs a lookup of the granularity protection information in order to determine whether to allow the memory access triggered by the target code to proceed, similar to the PAS filter described above.
[0151] The GPT load program logic 418 controls the loading of the GPT block 150 from memory as needed, similar to the GPT load circuit 21 shown in FIG. 1. However, in the case of the simulator, since the granularity protection information cache 22 may not be simulated, the simulator embodiment operates similar to a hardware device without the GPI cache 22. Thus, each memory access request is treated as if it were missing from the cache, so step 172 is omitted in FIG. 11, and as a result, the method proceeds directly from step 170 to step 174, and if the data integrity check at step 178 passes, the method proceeds directly from step 178 to step 184. The integrity check program logic 419 emulates the integrity check circuit 23 to perform a data integrity check on the block of granularity protection data loaded by the GPT load program logic 418.
[0152] In the present application, the term "configured to..." is used to mean that an element of a device has a configuration capable of performing a defined operation. In this context, "configuration" means the way in which hardware or software is arranged or interconnected. For example, a device may have dedicated hardware that provides a defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not mean that some change needs to be made to the device element in order to provide the defined operation.
[0153] Exemplary embodiments of the present invention are described in detail herein with reference to the accompanying drawings, but it is to be understood that the present invention is not limited to these exact embodiments and that various changes and modifications can be made by those skilled in the art without departing from the scope of the present invention as defined by the appended claims.
Claims
1. An apparatus comprising: an address translation circuit that translates a target virtual address specified by a memory access request into a target physical address associated with a selected physical address space selected from among a plurality of physical address spaces; a granularity protection entry load circuit for loading from memory a granularity protection data block including at least one granularity protection entry, each granularity protection entry corresponding to a respective granularity of a physical address and specifying granularity protection information indicating which of the plurality of physical address spaces is the permitted physical address space in which the granularity of the physical address is permitted to be accessed; a filtering circuit for determining whether the memory access request is to be permitted to access the target physical address based on whether the selected physical address space is indicated as a permitted physical address space by the granularity protection information in a target granularity protection entry corresponding to a target granularity of the physical address including the target physical address; a integrity check circuit for performing a data integrity check on the granularity protection data block loaded from memory and notifying of a failure when the data integrity check fails; a physical alias point (PoPA) memory system component configured to de-alias a plurality of aliased physical addresses from different physical address spaces corresponding to the same memory system resource to map any of the plurality of aliased physical addresses to an un-aliased physical address provided to at least one downstream memory system component; at least one PoPA pre-memory system component provided upstream of the PoPA memory system component, the at least one PoPA pre-memory system component being configured to handle the aliased physical addresses from different physical address spaces as if the aliased physical addresses corresponded to different memory system resources; and the apparatus comprising the same.
2. The granular protection data block includes a plurality of granular protection entries corresponding to granules with different physical addresses, The apparatus according to claim 1, wherein the integrity check circuit is configured to perform the data integrity check on the plurality of granular protection entries in the granular protection data block.
3. In response to the memory access request, When the target granular protection entry cannot yet be accessed by the filtering circuit, the granular protection entry load circuit is configured to load the granular protection data block including the target granular protection entry from the memory, The apparatus according to claim 2, wherein when the data integrity check fails for one of the plurality of granular protection entries of the granular protection data block other than the target granular protection entry, the integrity check circuit is configured to notify the failure even if the data integrity check passes for the target granular protection entry.
4. The apparatus according to any one of claims 1 to 3, comprising a memory encryption circuit for encrypting data stored in a protected area of the memory and decrypting data read from the protected area of the memory.
5. The apparatus according to claim 4, wherein when the granular protection entry load circuit loads a granular protection data block from a protected area of the memory, the memory encryption circuit is configured to decrypt the granular protection data block read from the protected area of the memory before providing the granular protection data block to the filtering circuit or the integrity check circuit.
6. An address conversion circuit that converts a target virtual address specified by a memory access request into a target physical address associated with a selected physical address space selected from a plurality of physical address spaces, A granule protection entry load circuit for loading a granule protection data block including at least one granule protection entry from a memory, each granule protection entry corresponding to a respective granule of a physical address, and specifying granule protection information indicating which of the plurality of physical address spaces is the permitted physical address space in which access to the granule of the physical address is permitted. A granule protection entry load circuit; A filtering circuit for determining whether the memory access request should be permitted to access the target physical address based on whether the selected physical address space is indicated as the permitted physical address space by the granule protection information in the target granule protection entry corresponding to the target granule of the physical address including the target physical address; A integrity check circuit for performing a data integrity check on the granule protection data block loaded from the memory and notifying a failure when the data integrity check fails; An apparatus comprising: The granule protection information is encoded using N bits of the granule protection entry, a first subset of the N-bit encoding is a valid encoding indicating a valid option of the granule protection information, and a second subset of the N-bit encoding is an invalid encoding; The data integrity check includes checking, for each granule protection entry in the granule protection data block, whether the granule protection information of the granule protection entry has one of the first subset of the encoding or the second subset of the encoding. An apparatus.
7. The apparatus according to claim 6, wherein the integrity check circuit is configured to notify the failure when any granule protection entry in the granule protection data block has one of the second subset of the encoding.
8. The apparatus according to any one of claims 1 to 5, wherein the data integrity check includes determining whether at least one integrity signature value associated with the granule protection data block matches at least one signature check value derived from the granule protection data block.
9. The apparatus according to any one of claims 1 to 8, further comprising a requester-side filtering circuit in the filtering circuit for determining whether to pass the memory access request to the cache or to pass it to an interconnect for communicating with a completer device for servicing the memory access request based on the granule protection information in the target granule protection entry.
10. The apparatus according to claim 9, wherein the integrity check circuit includes a requester-side integrity check circuit for performing the data integrity check on the granule protection data block loaded from the memory after the granule protection data block is received from the interconnect.
11. The apparatus according to any one of claims 1 to 10, wherein at least one of the address translation circuit and the filtering circuit is configured to select the selected physical address space based at least on a current operating region of a requester circuit that issued the memory access request, and the current region includes one of a plurality of regions of operation.
12. The address translation circuit is configured to convert the target virtual address to the target physical address based on at least one page table entry. The apparatus according to claim 11, wherein when at least the current region is one of a subset of the plurality of regions, at least one of the address translation circuit and the filtering circuit is configured to select the selected physical address space based on the current region and physical address space selection information specified by the at least one page table entry.
13. The plurality of regions includes a root region for managing switching between other regions. The apparatus according to claim 11 or 12, wherein when the data integrity check fails, the integrity check circuit is configured to notify the failure processed by the program code executed in the root area.
14. In response to the memory access request, if no valid granule protection entry is defined for the target granule of the physical address, or if the filtering circuit determines that the memory access request is not permitted to access the target physical address based on the target granule protection entry, the filtering circuit is configured to, when the data integrity check fails, notify the failure notified by the integrity check circuit with a failure associated with a different failure type or different failure syndrome information. The apparatus according to any one of claims 1 to 13.
15. The apparatus according to any one of claims 1 to 5 and 8, wherein the aliased physical address is represented using the same physical address value within the different physical address spaces.
16. A method comprising: converting a target virtual address specified by a memory access request into a target physical address associated with a selected physical address space selected from among a plurality of physical address spaces; loading from memory a granule protection data block including at least one granule protection entry, each granule protection entry corresponding to a respective granule of a physical address and specifying granule protection information indicating which of the plurality of physical address spaces is the permitted physical address space in which the granule of the physical address is permitted to be accessed; performing a data integrity check on the granule protection data block loaded from memory and notifying of a failure when the data integrity check fails. Determining whether the memory access request should be permitted to access the target physical address based on whether the selected physical address space is shown as a physical address space permitted by the granularity protection information in a target granularity protection entry for a target granularity of physical addresses including the target physical address; Demultiplexing the plurality of aliased physical addresses from different physical address spaces corresponding to the same memory system resource to map any of the plurality of aliased physical addresses to an unaliased physical address provided to at least one downstream memory system component; Treating the aliased physical addresses from different physical address spaces upstream of the demultiplexing as if the aliased physical addresses corresponded to different memory system resources; A method comprising.
17. A computer program including instructions for controlling the host data processing device to provide an instruction execution environment for executing target code when executed on the host data processing device, The computer program being Address translation program logic for translating a target virtual address specified by a memory access request into a target simulated physical address associated with a selected simulated physical address space selected from among a plurality of simulated physical address spaces; Granularity protection entry load program logic for loading from memory a granularity protection data block including at least one granularity protection entry, each granularity protection entry corresponding to a respective granularity of simulated physical addresses and specifying granularity protection information indicating which of the plurality of simulated physical address spaces is the permitted simulated physical address space in which access to the granularity of the simulated physical address is permitted; Filtering program logic for determining whether the memory access request should be permitted to access the target simulated physical address based on whether the selected simulated physical address space is shown as a simulated physical address space permitted by the granularity protection information in a target granularity protection entry that includes the target simulated physical address, Integrity check program logic for performing a data integrity check on the granularity protection data block loaded from memory and notifying of a failure when the data integrity check fails comprising The granularity protection information is encoded using N bits of the granularity protection entry, a first subset of the N-bit encoding is a valid encoding indicating valid options for the granularity protection information, and a second subset of the N-bit encoding is an invalid encoding, The data integrity check includes checking, for each granularity protection entry in the granularity protection data block, whether the granularity protection information of the granularity protection entry has one of the first subset of the encoding or the second subset of the encoding, a computer program. **Claim 18** A computer-readable storage medium storing the computer program according to claim 17.
Citation Information
Patent Citations
Memory control device and memory control method
JP2003242030A
Information processing device and access management program
WO2018158909A1
Invalidation of a target realm in a realm hierarchy
WO2019002810A1