Data processing apparatus and data processing method

By allowing access to partial translation information in the translation circuit even when permission information detection is incomplete, the problems of latency and security risks in memory address translation are solved, achieving more efficient and secure memory access.

CN115335815BActive Publication Date: 2026-07-31ARM LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ARM LTD
Filing Date
2021-03-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, obtaining permission information during memory address translation introduces significant overhead, especially in the context of a multi-stage MMU, leading to increased latency and security risks.

Method used

By allowing access to partial conversion information, particularly the conversion information address, in the conversion circuit even when permission information detection is incomplete, the reliance on permission information is reduced. This includes delaying or omitting permission information detection operations, especially during read access, and only allowing write access when permission information is permitted.

Benefits of technology

It reduces memory address translation latency, improves system security, and reduces security risks caused by unauthorized information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115335815B_ABST
    Figure CN115335815B_ABST
Patent Text Reader

Abstract

An apparatus includes: a conversion circuit for performing a conversion operation to generate a converted second memory address in a second memory address space as a conversion of a first memory address in a first memory address space, wherein the conversion circuit is configured to generate the converted second memory address based on conversion information stored at one or more conversion information addresses; an access control circuit for performing an operation to detect access control information to indicate whether memory access to the given second memory address is permitted; and an access circuit for allowing access to data stored at the given second memory address when the access control information indicates that memory access to the given second memory address is permitted; the access circuit is configured to selectively allow access to the conversion information address by the conversion circuit without the access control circuit having already performed the operation to detect access control information to indicate whether memory access to the conversion information address is permitted.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] This disclosure relates to apparatus and methods.

[0002] The data processing system may have address translation circuitry to translate the virtual address of a memory access request into the physical address corresponding to the location to be accessed in the memory system.

[0003] The process of generating this address translation may itself require multiple memory accesses. Summary of the Invention

[0004] In one exemplary arrangement, an apparatus is provided, the apparatus comprising:

[0005] A conversion circuit is configured to perform a conversion operation to generate a converted second memory address in a second memory address space as a conversion of a first memory address in a first memory address space, wherein the conversion circuit is configured to generate the converted second memory address based on conversion information stored at one or more conversion information addresses.

[0006] The permission circuit is used to perform the operation of detecting permission information to indicate whether memory access to a given second memory address is permitted; and

[0007] An access circuit that allows access to data stored at a given second memory address when an authorization message indicates that memory access to a given second memory address is permitted.

[0008] The access circuit is configured to selectively allow the translation circuit to access the translation information address without requiring the permission circuit to have completed the operation of detecting permission information to indicate whether memory access to the translation information address is permitted.

[0009] In another exemplary arrangement, a method is provided that includes:

[0010] Performing a conversion operation to generate a converted second memory address in the second memory address space as a conversion of a first memory address in the first memory address space includes generating the converted second memory address based on conversion information stored at one or more conversion information addresses;

[0011] Perform the operation of detecting permission information to indicate whether memory access to a given second memory address is permitted;

[0012] Accessing data stored at a given second memory address when permission information indicates that memory access to that address is permitted; and

[0013] Selectively access translation information addresses without requiring the permission circuitry to have completed the operation of detecting permission information to indicate whether memory access to the translation information address is permitted.

[0014] In another exemplary arrangement, a computer program is provided for controlling a host data processing device to provide an instruction execution environment for executing object code; the computer program includes:

[0015] The conversion logic is used to perform a conversion operation to generate a converted second memory address in a second memory address space as a conversion of a first memory address in a first memory address space, wherein the conversion logic is configured to generate the converted second memory address based on conversion information stored at one or more conversion information addresses.

[0016] The permission logic is used to perform operations that detect permission information to indicate whether memory access to a given second memory address is permitted; and

[0017] Access logic, which allows access to data stored at a given second memory address when permission information indicates that memory access to a given second memory address is permitted;

[0018] The access logic is configured to selectively allow the translation logic to access the translation information address without requiring the permission logic to have completed the operation of detecting permission information to indicate whether memory access to the translation information address is permitted.

[0019] Other relevant aspects and features are defined by the appended claims. Attached Figure Description

[0020] The present invention will be further described by way of example only, with reference to embodiments shown in the accompanying drawings, wherein:

[0021] Figure 1 An example of a data processing device is shown;

[0022] Figure 2 Multiple domains in which the processing circuitry can operate are shown;

[0023] Figure 3 An example of a processing system that supports particle protection lookup is shown;

[0024] Figure 4 This schematically illustrates multiple physical address spaces as aliases in the system physical address space that identify locations within the memory system;

[0025] Figure 5An example is shown of partitioning the effective hardware physical address space so that different architecture physical address spaces have access to the corresponding portions of the system physical address space;

[0026] Figure 6 and Figure 7 The diagram illustrates data encryption and decryption.

[0027] Figure 8 An aspect of the operation of an exemplary memory management unit (MMU) is illustrated schematically;

[0028] Figure 9 A single-stage MMU is schematically illustrated;

[0029] Figure 10 A two-stage MMU is schematically illustrated;

[0030] Figure 11 and Figure 12 The operation of a two-stage MMU and a single-stage MMU with particle protection operation is illustrated (respectively).

[0031] Figure 13 A single-stage MMU with at least partial omission of particle protection operation is schematically shown;

[0032] Figure 14 A two-stage MMU with at least partial omission of particle protection operation is schematically shown;

[0033] Figure 15 and Figure 16 The delayed MMU operation with particle protection is schematically illustrated;

[0034] Figure 17 It is a schematic flowchart illustrating the method; and

[0035] Figure 18 An example of a simulator that can be used is shown. Detailed Implementation

[0036] Before discussing the implementation scheme with reference to the accompanying drawings, the following description of the implementation scheme is provided.

[0037] An exemplary embodiment provides an apparatus comprising:

[0038] A conversion circuit is configured to perform a conversion operation to generate a converted second memory address in a second memory address space as a conversion of a first memory address in a first memory address space, wherein the conversion circuit is configured to generate the converted second memory address based on conversion information stored at one or more conversion information addresses.

[0039] The permission circuit is used to perform the operation of detecting permission information to indicate whether memory access to a given second memory address is permitted; and

[0040] An access circuit that is used to access data stored at a given second memory address when permission information indicates that memory access to a given second memory address is permitted;

[0041] The access circuit is configured to access the translation information address without requiring the permission circuit to have already performed the operation of detecting the permission information to indicate whether memory access to the translation information address is permitted.

[0042] This disclosure recognizes that operations (such as those performed by translation circuitry, such as a memory management unit or MMU) can themselves involve numerous memory accesses. Where access information needs to be obtained before each of those accesses, obtaining this access information can introduce significant overhead into translation generation, especially when the access information is also held in memory. This can be a specific problem in the context of a multi-stage MMU.

[0043] In an exemplary arrangement, in the absence of a completed process for obtaining authorization information, access is permitted to at least some transformation information, which is information (such as so-called page table entries) used by the transformation circuit to generate the transformation.

[0044] This arrangement can help reduce the latency associated with obtaining memory address translations.

[0045] While these arrangements can be applied to read and write operations performed by the conversion circuitry, in the exemplary embodiments, it should be noted that (a) a significant portion of the latency associated with obtaining a conversion typically involves read operations performed by the conversion circuitry, and (b) the security risk may be less in the absence of a prior situation where the process for obtaining authorization information has not been completed, provided the arrangement is limited to read operations performed by the conversion circuitry. Therefore, in the exemplary embodiments, when access to the conversion information address involves a read access, the access circuitry is configured to access the conversion information address before the authorization circuitry has completed its operation of detecting authorization information; and when access to the conversion information address involves a write access, the access circuitry is configured to access the conversion information address only if the authorization information indicates that memory access to the conversion information address is permitted.

[0046] In some examples, the permission circuitry is configured to perform another operation to detect the memory type applicable to a given second memory address, at least the first memory type or a different second memory type applicable to the given second memory address. For example, the access circuitry can be configured to access a translation information address without the permission circuitry having already performed the operation of detecting permission information to indicate whether memory access to the translation information address is permitted only if the memory type applicable to the translation information address is the first memory type. This is particularly relevant if the first memory type is a memory type in which the data stored at the given address has not been altered by a read operation from the given address. For example, the other memory type could be a memory type in which the data stored at the given address may be altered by a read operation from the given address, such as a memory type associated with input / output circuitry, where the given address is such as an address mapped to a register (such as a first-in-first-out (FIFO) register) in which a read operation alters the nature of the data to be read by a subsequent read operation.

[0047] In the exemplary arrangement, the operation of detecting permission information can be postponed, while in other examples, this operation can be omitted or omitted. As an example of at least partial omission, the permission circuit is configured not to perform the operation of detecting permission information with respect to at least some of the translation information addresses. As another measure to avoid security risks arising from continuing without having obtained permission information, the translation circuit can be configured not to provide the translation information retrieved from the translation information address as output to circuits outside the translation circuit (or actually to software running on the processor accessing the translation circuit) with respect to which the operation of detecting permission information has not yet been completed.

[0048] This disclosure is particularly applicable to conversion circuits, wherein conversion information applicable to a conversion of a given first memory address comprises a hierarchical structure of conversion information entries (e.g., so-called page table entries or PTEs), wherein data indicating the conversion information address of the next conversion information entry is indicated by the preceding conversion information entry. In such an arrangement, the data indicating the conversion information address of the next conversion information entry may indicate at least a portion of the first memory address applicable to the next conversion information entry; and the conversion circuit is configured to perform a conversion operation to generate a corresponding conversion information address.

[0049] This arrangement can be useful in contexts where obtaining permission information is delayed, such as in an arrangement where the permission circuit is configured to postpone initiating the operation of detecting permission information for the next transition information entry until after access to that next transition information entry is initiated.

[0050] When the conversion circuit is operable with respect to memory access transactions, each memory access transaction is associated with a first memory address for conversion, the conversion circuit associates a converted second memory address with each memory access transaction, the authorization circuit can be configured to perform an operation to detect authorization information relative to the converted second memory address of each memory access transaction, and the access circuit is configured to provide the memory access transaction with the result of access to the converted second memory address only if the authorization data permits access to the converted second memory address.

[0051] In an exemplary arrangement related to the operation of the conversion circuit, the first memory address may include one of a virtual memory address and an intermediate physical address; and the second memory address may include a physical memory address.

[0052] This invention is particularly suitable for use with memories having multiple memory partitions, each associated with a partition identifier and having a corresponding physical address range within a physical address space. Here, the access control circuitry can be configured to detect access control information by: detecting a region identifier associated with a second memory address, selected from a plurality of region identifiers, each region identifier indicating access control to a corresponding set of memory partitions, wherein for at least one of the region identifiers, the corresponding set of memory partitions includes a subset of one or more, but not all, memory partitions; and comparing the detected region identifier with the partition identifier associated with the second memory address.

[0053] As an additional security layer to prevent memory access with incorrect region identifiers, the device may include encryption and decryption circuitry for encrypting data stored in the memory and decrypting data retrieved from the memory; wherein the encryption and decryption circuitry is configured to apply a corresponding encryption and corresponding decryption from a set of encryption and corresponding decryption to each memory partition, the set of encryption and corresponding decryption such that data encrypted to a given memory partition by the corresponding encryption for that memory partition cannot be decrypted by applying decryption to another memory partition.

[0054] In an exemplary arrangement, the access control circuitry is configured to be associated with a translated second memory address, and the data indicates a region identifier associated with the translated second memory address.

[0055] Encryption and decryption operations can be arranged such that the encryption and decryption circuitry is configured to apply decryption by applying decryption selected according to data indicating a region identifier associated with the translated second memory address to decrypt data retrieved from memory at the translated second memory address.

[0056] As another measure to mitigate the security risks arising from memory access in the absence of an operation that has already completed the detection of authorization information, and in the context of means including one or more cache memories for holding data retrieved from and / or stored in memory, the cache memory can be configured to associate a corresponding region identifier with each data item held by the cache memory; and the cache memory can be configured to suppress access to the data item associated with a given region identifier in response to a memory access associated with data indicating a different region identifier.

[0057] As another measure to mitigate the security risks arising from memory access in the absence of an operation that has already completed the detection of access rights information, the translation circuit is configured to detect a translation failure with respect to a given translation operation when the use of translation information by the translation circuit does not provide a valid address translation; and in response to the detection of a translation failure, the translation circuit is configured to control the access rights circuit to perform the operation of detecting access rights information with respect to any translation information address accessed as part of a given translation operation.

[0058] In an exemplary embodiment, the apparatus includes a processor for executing program instructions at the current exception level of a hierarchical structure of exception levels, each exception level being associated with a security privilege such that instructions executed at a higher exception level can access resources that cannot be accessed by instructions executed at a lower exception level; wherein the processor is required to execute instructions at the highest exception level in the exception levels in order to set up permission circuitry to detect permission information data.

[0059] Another exemplary implementation provides a method comprising:

[0060] Performing a conversion operation to generate a converted second memory address in the second memory address space as a conversion of a first memory address in the first memory address space includes generating the converted second memory address based on conversion information stored at one or more conversion information addresses;

[0061] Perform the operation of detecting permission information to indicate whether memory access to a given second memory address is permitted;

[0062] Accessing data stored at a given second memory address when permission information indicates that memory access to that address is permitted; and

[0063] Access the translation information address without requiring the permission circuitry to have already performed the operation of detecting permission information to indicate whether memory access to the translation information address is permitted.

[0064] Another exemplary embodiment provides a computer program for controlling a host data processing device to provide an instruction execution environment for executing object code; the computer program includes:

[0065] The conversion logic is used to perform a conversion operation to generate a converted second memory address in a second memory address space as a conversion of a first memory address in a first memory address space, wherein the conversion logic is configured to generate the converted second memory address based on conversion information stored at one or more conversion information addresses.

[0066] The permission logic is used to perform operations that detect permission information to indicate whether memory access to a given second memory address is permitted; and

[0067] Access logic, which allows access to data stored at a given second memory address when permission information indicates that memory access to a given second memory address is permitted;

[0068] The access logic is configured to selectively allow the translation logic to access the translation information address without requiring the permission logic to have completed the operation of detecting permission information to indicate whether memory access to the translation information address is permitted.

[0069] Introducing and controlling access to the physical address space

[0070] The data processing system can support the use of virtual memory, providing address translation circuitry to translate the virtual address specified by a memory access request into a physical address associated with the location in the memory system to be accessed. The mapping between virtual and physical addresses can be defined in one or more page table structures. Page table entries within the page table structure can also define access permission information that controls whether a given software procedure executing on the processing circuitry is allowed to access a specific virtual address.

[0071] In some processing systems, address translation circuitry maps all virtual addresses to a single physical address space, which the memory system uses to identify locations in memory to be accessed. In such systems, control over whether a particular address is accessible to a specific software process is based solely on the page table structure used to provide the virtual-to-physical address translation mapping. However, this page table structure is typically defined by the operating system and / or hypervisor. If the operating system or hypervisor is compromised, this can introduce security vulnerabilities, allowing attackers to access sensitive information.

[0072] Therefore, for systems that require certain processes to be executed securely in isolation from other processes, such systems can support operations in multiple domains and multiple different physical address spaces. For at least some components of the memory system, even if physical addresses in corresponding physical address spaces actually correspond to the same location in memory, memory access requests whose virtual addresses are translated into physical addresses in different physical address spaces are treated as completely separate addresses accessing memory. By isolating accesses from different operational domains of the processing circuitry into corresponding different physical address spaces as perceived by some memory system components, this provides stronger security guarantees independent of page table permission information set by the operating system or hypervisor.

[0073] Processing circuitry can support processing within a root domain, which is responsible for managing switching between other domains in which the processing circuitry can operate. By providing a dedicated root domain for controlling switching, this helps maintain security by limiting the extent to which code executing in one domain can trigger a switch to another. For example, the root domain can perform various security checks when a domain switch is requested.

[0074] Therefore, the processing circuitry can support processing performed in one of at least three domains: the root domain and at least two other domains. The address translation circuitry can translate the virtual address of a memory access performed from the current domain into a physical address in at least one of a plurality of physical address spaces selected based on the current domain.

[0075] The root physical address space can be exclusively accessed from the root domain. Therefore, processing circuitry might be unable to access the root physical address space while operating in one of the other domains. This enhances security by ensuring that code executing in one of the other domains does not tamper with data or program code that the root domain relies on for managing switching between domains or controlling what permissions processing circuitry has in one of the other domains. On the other hand, in this example, all multiple physical address spaces can be accessed from the root domain. Since code executing in the root domain must be trusted by either party providing the code operating in one of the other domains, and since the root domain code is responsible for switching to the specific domain in which that party's code is executing, the root domain can inherently be trusted to access any physical address space. Making all physical address spaces accessible from the root domain allows functions such as translating memory regions into and out of a domain, copying code and data into a domain (e.g., during boot), and providing services to that domain.

[0076] Example description

[0077] Figure 1An example of a data processing system 2 having at least one requester device 4 and at least one completer device 6 is schematically illustrated. Interconnection 8 provides communication between the requester device 4 and the completer device 6. The requester device is capable of issuing a memory access request for memory access to a specific addressable memory system location. The completer device 6 is the device responsible for servicing memory access requests directed to it. Although... Figure 1 Not shown, but some devices may be able to act as both requester and completer devices. Requester device 4 may include, for example, processing elements such as a central processing unit (CPU) or a graphics processing unit (GPU), or other host devices such as a bus master, network interface controller, display controller, etc. Completer devices may include memory controllers responsible for controlling access to corresponding memory storage units, peripheral controllers for controlling access to peripheral devices, etc. Figure 1 An exemplary configuration of one of the requester devices 4 is shown in more detail, but it should be understood that other requester devices 4 may have similar configurations. Alternatively, other requester devices may have the same... Figure 1 The left side shows different configurations of the requester device 4.

[0078] The requester device 4 has processing circuitry 10 that performs data processing in response to instructions, referencing data stored in register 12. Register 12 may include general-purpose registers for storing operands and the results of processed instructions, and control registers for storing control data to configure how the processing circuitry performs processing. For example, the control data may include a current domain indicator 14 for selecting which operation domain is the current domain, and a current exception level indicator 15 for indicating which exception level is the current exception level that the processing circuitry 10 is operating at.

[0079] Processing circuitry 10 may issue a memory access request specifying a virtual address (VA) identifying the addressable location to be accessed and a domain identifier (domain ID or "security state") identifying the current domain. Address translation circuitry 16 (e.g., a memory management unit (MMU)) translates the virtual address into a physical address (PA) through one or more stages of address translation based on page table data defined in a page table structure stored in the memory system. Translation lookup buffer (TLB) 18 acts as a lookup cache to cache some information in the page table information, thus enabling faster access than if the page table information had to be retrieved from memory every time an address translation is needed. In this example, in addition to generating the physical address, address translation circuitry 16 selects one of several physical address spaces associated with the physical address and outputs a Physical Address Space (PAS) identifier identifying the selected physical address space. The selection of PAS will be discussed in more detail below.

[0080] PAS filter 20 acts as a requester-side filter circuit to check whether access to the specified physical address within the physical address space identified by the PAS identifier is permitted based on the translated physical address and the PAS identifier. This lookup is based on particle protection information stored in a particle protection table (GPT) structure within the memory system. Similar to page table data being cached in TLB 18, particle protection information can be cached in particle protection information cache 22. The particle protection information defines the information restricting access to the physical address space from which a given physical address can be accessed, and based on this lookup, PAS filter 20 determines whether a memory access request is allowed to continue being issued to one or more caches 24 and / or interconnects 8. If the specified PAS of the memory access request is not permitted to access the specified physical address, PAS filter 20 blocks the transaction and may signal a fault.

[0081] The PAS filter can function (in part) in response to a control signal (schematically shown as signal 21) from the address translation circuitry, which indicates that the omission or postponement of at least some checks or other operations performed by the PAS filter may or should occur. These operations will be discussed in more detail below.

[0082] Although Figure 1 An example of a system including multiple requester devices 4 is shown, but for Figure 1 The feature shown by a requester device on the left-hand side can also be included in systems where only one requester device exists (such as a single-core processor).

[0083] Although Figure 1 An example is shown where address translation circuit 16 performs the selection of a PAS for a given request. However, in other examples, address translation circuit 16 may output information for determining which PAS to select, along with the PA, to PAS filter 20, and PAS filter 20 may select the PAS and check whether access to the PA is permitted within the selected PAS.

[0084] The provision of PAS filter 20 helps support systems that can operate in multiple operational domains, each associated with its own isolated physical address space. This means that, for at least a portion of the memory system (e.g., for some cache or coherence implementation such as a snooping filter), even if addresses within these address spaces actually relate to the same physical location in the memory system, each individual physical address space is treated as a completely separate set of addresses that identify the individual memory system location. This can be useful for security purposes.

[0085] Figure 2Examples of different operating states and domains in which the processing circuit 10 can operate are shown, as well as examples of the types of software that can be executed in different exception levels and domains (of course, it should be understood that the specific software installed on the system is selected by the parties managing the system and is therefore not a fundamental feature of the hardware architecture).

[0086] The processing circuitry 10 can operate at multiple different exception levels 80 (in this example, four exception levels labeled EL0, EL1, EL2, and EL3), where EL3 refers to the exception level with the highest privilege level, and EL0 refers to the exception level with the lowest privilege level. [It should be understood that other architectures may choose the reverse numbering so that the exception level with the highest number can be considered to have the lowest privilege.] In this example, the lowest privilege exception level EL0 is used for application-level code, the next highest privilege exception level EL1 is used for operating system-level code, the next highest privilege exception level EL2 is used for hypervisor-level code managing the switching between multiple virtualized operating systems, and the highest privilege exception level EL3 is used for monitoring code managing the switching between corresponding domains and the allocation of physical addresses to the physical address space.

[0087] Therefore, the processing circuit 10 is configured to execute program instructions at the current exception level in a hierarchical structure of exception levels, each exception level being associated with a security privilege, such that instructions executed at a higher exception level can access resources that cannot be accessed by instructions executed at a lower exception level. As discussed below, the processing circuit needs to execute instructions at the highest exception level (e.g., EL3) in order to set the privilege circuit or PAS filter 20 to detect privilege information data.

[0088] When an exception occurs while the processing software is at a specific exception level, for some types of exceptions, an exception of a higher (higher privilege) level is generated, where the specific exception level to generate the exception is selected based on the attributes of the specific exception that occurred. However, in some cases, it is possible for other types of exceptions to be generated at the same exception level as the exception level associated with the code that was processed when the exception occurred. When an exception occurs, information characterizing the state of the processor at the time the exception occurred can be saved, including, for example, the current exception level at the time the exception occurred. Therefore, once an exception handler has been processed to handle the exception, processing can return to the previous processing, and the saved information can be used to identify the exception level to which the processing should return.

[0089] In addition to different exception levels, the processing circuitry supports multiple operating domains, including the root domain 82, the safe (S) domain 84, the less safe domain 86, and the domain 88. For ease of reference, the less safe domain will be described hereinafter as the “unsafe” (NS) domain, but it should be understood that this is not intended to imply any particular level of safety (or lack thereof). Rather, “unsafe” simply indicates that the unsafe domain is intended for code that is not as safe as code operating in a safe domain. The root domain 82 is selected when the processing circuitry 10 is at the highest exception level EL3. When the processing circuitry is at one of the other exception levels EL0 through EL2, the current domain is selected based on the current domain indicator 14, which indicates which of the other domains 84, 86, and 88 is active. For each of the other domains 84, 86, and 88, the processing circuitry can be at any exception level of EL0, EL1, or EL2.

[0090] At boot time, multiple fragments of boot code (e.g., BL1, BL2, OEM boot) can be executed, for example, within a higher privilege exception level EL3 or EL2. Boot code BL1, BL2 may be associated with, for example, a root domain, and OM boot code may operate within a security domain. However, once the system is booted, during runtime, processing circuitry 10 can be considered to operate at one time within one of domains 82, 84, 86, and 88. Each of domains 82 through 88 is associated with its own associated physical address space (PAS), which achieves the isolation of data from different domains within at least a portion of the memory system. This will be described in more detail below.

[0091] Non-security domain 86 can be used for regular application-level processing and for operating system and hypervisor activities to manage such applications. Thus, within non-security domain 86, there can be application code 30 operating at EL0, operating system (OS) code 32 operating at EL1, and hypervisor code 34 operating at EL2.

[0092] Security domain 84 enables the isolation of certain on-chip security, media, or system services into a separate physical address space from the physical address space used for non-secure processing. Non-secure domain code cannot access resources associated with security domain 84, while secure domain code can access both secure and non-secure resources; in this sense, secure and non-secure domains are not equivalent. An example of a system supporting this partitioning of security domain 84 and non-secure domain 86 is based on Arm... ® TrustZone provided by Limited ®The system architecture includes a trusted application 36 at EL0, a trusted operating system 38 at EL1, and optionally a secure partition manager 40 at EL2. If secure partitioning is supported, the secure partition manager uses stage 2 page tables to support isolation between different trusted operating systems 38 running in the secure domain 84, in a manner similar to how the hypervisor 34 manages isolation between virtual machines or guest operating systems 32 running in the non-secure domain 86.

[0093] Extending this system to support security domain 84 has become common in recent years because it enables a single hardware processor to support isolated secure processing, thus avoiding the need to execute that processing on a separate hardware processor. However, with the increasing prevalence of security domains, many real-world systems with such security domains now support relatively complex hybrid service environments offered by a wide variety of different software vendors within the security domain. For example, code operating in security domain 84 may include different pieces of software provided by, among other things: silicon wafer providers that manufacture integrated circuits; original equipment manufacturers (OEMs) that assemble the integrated circuits provided by the silicon wafer providers into electronic devices such as mobile phones; operating system vendors (OSVs) that provide the operating system 32 for such devices; and / or cloud platform providers that manage cloud servers that support services for multiple different clients via the cloud.

[0094] However, there is a growing desire to provide secure computing environments for parties providing user-level code (which may typically be expected to execute as application 30 within a non-secure domain 86), environments that can be trusted not to leak information to other parties operating the code on the same physical platform. It may be desirable for such secure computing environments to be dynamically allocated at runtime and to be certified and provable, allowing users to verify adequate security guarantees on the physical platform before trusting the device to process potentially sensitive code or data. Users of such software may not want to trust a party providing a rich operating system 32 or hypervisor 34 that may typically operate in a non-secure domain 86 (or even if these providers are trusted, users may want to protect themselves from attackers who could compromise the operating system 32 or hypervisor 34). Furthermore, while a secure domain 84 can be used for such user-provided applications requiring secure processing, this can actually create problems for both users providing code that requires a secure computing environment and providers of existing code operating within a secure domain 84. For providers of existing code operating within security domain 84, adding arbitrary user-supplied code within the security domain would increase the attack surface of their code, which is likely undesirable. Therefore, it is strongly recommended that users not be allowed to add code to security domain 84. On the other hand, users providing code that requires a secure computing environment may be reluctant to trust that all providers of different fragments of code operating within security domain 84 have access to their data or code. If authentication or certification of code operating within a specific domain is required as a prerequisite for user-supplied code to perform its processing, it may be difficult to audit and authenticate all different fragments of code operating within security domain 84 provided by different software providers. This could limit opportunities for third parties to provide more secure services.

[0095] Therefore, as Figure 2 As shown, an additional domain 88 (referred to as a domain domain) is provided, which can be used by code introduced by such users to provide a secure computing environment orthogonal to any secure computing environment associated with components operating in security domain 24. Within a domain domain, the software executed may include multiple domains, each of which can be isolated from other domains by a Domain Management Module (RMM) 46 operating at exception level EL2. The RMM 46 can control the isolation between corresponding domains 42, 44 executing domain domain 88, for example, by defining access permissions and address mappings in a page table structure, in a manner similar to how the hypervisor 34 manages the isolation between different components operating in non-security domain 86. In this example, the domains include an application-level domain 42 executing at EL0 and a packaged application / operating system domain 44 executing across exception levels EL0 and EL1. It should be understood that it is not necessary to support both EL0 and EL0 / EL1 type domains simultaneously, and multiple domains of the same type can be created by the RMM 46.

[0096] Similar to security domain 84, domain 88 has its own physical address space allocated to it. However, while domain 88 and security domain 84 can each access the non-secure PAS associated with non-secure domain 86, they cannot access each other's physical address spaces. In this sense, the domain is orthogonal to security domain 84. This means that the code executing in domain 88 and security domain 84 is independent of each other. The code in the domain only needs to trust the switching code between the hardware RMM 46 and the management domain operating in root domain 82, which makes proof and authentication more feasible. Proof enables a given piece of software to request verification that the code installed on the device matches certain expected characteristics. This can be achieved by checking whether the hash of the program code installed on the device matches an expected value signed by a trusted party using a cryptographic protocol. RMM 46 and monitoring code 29 can be verified, for example, by checking whether the hash of the software matches the expected value signed by a trusted party, such as a silicon supplier that manufactures the integrated circuits including processing system 2 or an architecture supplier that designs processor architectures that support domain-based memory access control. This allows the user-provided codes 42 and 44 to verify whether the integrity of the domain-based architecture can be trusted before performing any security or sensitive functions.

[0097] Thus, it can be seen that the code associated with domains 42 and 44 (which was previously executed in non-secure domain 86, as shown by the dashed lines indicating gaps in the non-secure domain where these processes were previously executed) can now be moved to the domain domain, where they can have stronger security guarantees because their data and code will not be accessed by other code operating in non-secure domain 86. However, the fact that domain domain 88 and secure domain 84 are orthogonal and therefore cannot see each other's physical address spaces means that the provider of the code in the domain domain does not need to trust the provider of the code in the secure domain, and vice versa. The code in the domain domain can simply trust the trusted firmware that provides the monitoring code 29 for root domain 82 and the RMM 46 provided by the silicon provider or the provider of the instruction set architecture supported by the processor (which may already be inherently needed to be trusted when the code executes on its device), so that a secure computing environment can be provided to users without further trust relationships with other operating system vendors, OEMs, or cloud hosts.

[0098] This can be used in a range of applications and use cases, including, for example, mobile wallets and payment applications, game anti-cheating and anti-piracy mechanisms, operating system platform security enhancements, secure virtual machine hosting, confidential computing, and gateway processing for networked or IoT devices. It should be understood that users can find many other applications with useful support for this domain.

[0099] To support security assurances provided to the domain, the processing system may support a proof reporting function, in which firmware images and configurations (e.g., monitoring code images and configurations or RMM code images and configurations) are measured at boot time or runtime, and domain content and configurations are measured at runtime, so that the domain owner can trace back the relevant proof reports to known implementations and certifications, thereby making a trust decision on whether to operate on the system.

[0100] like Figure 2 As shown, a separate root domain 82 is provided for managing domain switching, and this root domain has its own isolated root physical address space. Even for systems with only non-secure domains 86 and secure domains 84 but no domain 88, the creation of the root domain and the isolation of its resources from the secure domains allows for a more robust implementation, but it can also be used in implementations that do not support domain 88. Root domain 82 can be implemented using monitoring software 29 provided (or certified) by the silicon provider or architect, and can be used to provide secure boot functionality, trusted boot measurement, on-chip system configuration, debug control, and management of firmware updates for firmware components provided by other parties (such as OEMs). Root domain code can be developed, certified, and deployed by the silicon provider or architect without relying on the final device. In contrast, secure domain 84 can be managed by the OEM to implement certain platform and security services. The management of non-secure domain 86 can be controlled by operating system 32 to provide operating system services, while domain 88 allows for the development of new forms of trusted execution environments that can be dedicated to user or third-party applications while being isolated from the existing security software environment in secure domain 84.

[0101] Figure 3 Another example of a processing system 2 used to support these technologies is schematically shown. (Using the same reference numerals, it is shown alongside...) Figure 1 The same components. Figure 3The address translation circuit 16 is shown in more detail, comprising a Stage 1 memory management unit 50 and a Stage 2 memory management unit 52. The Stage 1 MMU 50 is responsible for translating virtual addresses into physical addresses (when triggered by EL2 or EL3 code) or into intermediate addresses (when triggered by EL0 or EL1 code in a certain operating state, requiring further Stage 2 translation by the Stage 2 MMU 52). The Stage 2 MMU translates intermediate addresses into physical addresses. The Stage 1 MMU can be based on a page table controlled by the operating system for translations initiated from EL0 or EL1, a page table controlled by the hypervisor for translations from EL2, or a page table controlled by the monitoring code 29 for translations from EL3. Conversely, the Stage 2 MMU 52 can be based on a page table structure defined by the hypervisor 34, RMM 46, or security partition manager 14, depending on which domain is used. Dividing these translations into two phases in this way allows the operating system to manage address translation for itself and applications (assuming they are the only operating system running on the system), while RMM 46, Hypervisor 34, or SPM40 can manage isolation between different operating systems running in the same domain.

[0102] like Figure 3 As shown, the address translation process using address translation circuit 16 can return a security attribute 54, which, combined with the current exception level 15 and the current domain 14 (or security state), allows access to a portion of a specific physical address space (identified by a PAS identifier or "PAS TAG") in response to a given memory access request. This provides an example of an authorization circuit (20) configured to be associated with a translated second memory address, the data (PAS TAG) indicating a region identifier associated with the translated second memory address.

[0103] The physical address and PAS identifier can be found in the Particle Protection Table 56, which provides the previously described particle protection information. In this example, the PAS filter 20 is shown as a Particle Memory Protection Unit (GMPU) that verifies whether the physical address of the selected PAS access request is allowed, and if so, allows the transaction to be passed to any cache 24 or interconnect 8 that is part of the system architecture of the memory system.

[0104] The GMPU 20 allows memory to be allocated to separate address spaces while providing strong hardware-based isolation guarantees and offering spatial and temporal flexibility as well as efficient sharing schemes in terms of how physical memory is allocated to these address spaces. As previously described, execution units in the system are logically partitioned into virtual execution states (domains or "worlds"), with one execution state (root world) located at the highest exception level (EL3), which is called the "root world" and manages the allocation of physical memory to these worlds.

[0105] A single system physical address space is virtualized into multiple "logical" or "architectural" physical address spaces (PASs), each of which is an orthogonal address space with independent consistency properties. The system physical address is mapped to a single "logical" physical address space by extending the system physical address with a PAS label.

[0106] Allowing a subset of the logical-physical address space to be accessed in a given world. This is implemented by hardware filter 20, which can be attached to the output of memory management unit 16.

[0107] The world uses fields in the translation table descriptor of the page table used for address translation to define the security attributes (PAS tags) for that access. Hardware filter 20 has access rights to a table (granular protection table 56 or GPT) that defines granular protection information (GPI) for each page in the system's physical address space, which indicates the associated PAS tag and (optionally) other granular protection attributes.

[0108] In some examples, so-called Level 0 (L0) GPT checks and Level 1 (L1) GPT checks are provided. L0 information indicates the type of memory associated with the PA and at least indicates whether a so-called side effect is likely to occur on a read access. For example, in the case of a PA assigned to an input / output device (which may provide data for reading from a first-in-first-out (FIFO) or other registers), the action of reading data from that PA by retrieving a data item from the FIFO register can change the data provided in response to the next read, such that the retrieved data is no longer available for access by the next read operation. On the other hand, reading data from DRAM typically does not suffer from such side effects and does not change the data to be read by the next operation at the same PA.

[0109] Therefore, an L0 GPT check can be used (as a useful additional benefit) to detect whether such side effects will potentially be experienced. If the answer is no, then by initiating a read operation at that PA, there is no direct risk to the data integrity at that particular PA.

[0110] The L0 GPT information upon which the check is based can be relatively coarse-grained, for example, with a 1 GB granularity. Therefore, the size of the L0 GPT data to be queried as part of the L0 GPT check can be relatively small (in the physical address space potentially one data item per GB). This, in turn, allows for relatively easy caching of the L0 GPT data, enabling the L0 GPT check to be performed with relatively little impact on latency to the processes discussed below.

[0111] Generally speaking, performing an L0 GPT check is an example of a permission circuit (such as a GMPU) performing another operation to detect the storage type applicable to a given second (e.g., physical) memory address, at least the first storage type or a different second storage type applicable to the given second memory address. For example, the first storage type could be a storage type where the data stored at the given address is not changed by a read operation from the given address (that is, a storage type that does not suffer the "side effects" as described above).

[0112] For example, L1 GPT checks can provide permission information and PAS tags.

[0113] Hardware filter 20 checks the world ID and security attributes against the particle's GPI and determines whether access can be granted, thereby forming a particle-shaped memory protection unit (GMPU).

[0114] For example, GPT 56 can reside in on-chip SRAM or off-chip DRAM. If stored off-chip, GPT 56 can be protected for integrity by an on-chip memory protection engine that can use encryption, integrity, and freshness mechanisms to maintain the security of GPT 56.

[0115] Positioning GMPU 20 on the requester side of the system (e.g., on the MMU output) instead of the completer side allows for page-level access permission allocation, while allowing interconnect 8 to continue hashing / stripping the page across multiple DRAM ports.

[0116] Transactions remain tagged with the PAS tag because they propagate throughout system architecture 24, 8 until they reach a location defined as a physical alias point 60. This allows filtering to be positioned on the master side (requester side) without compromising security guarantees compared to filtering on the slave side (requester side). Because the transaction propagates throughout the system, the PAS tag can be used as a security-in-depth mechanism for address isolation: for example, a cache can add a PAS tag to the address label in the cache, preventing accesses to the same PA with an incorrect PAS tag from hitting the cache and thus improving resistance to side-channel attacks. The PAS tag can also be used as a context selector for a protection engine attached to the memory controller, which encrypts data before writing it to external DRAM. Examples of such a protection engine will be discussed below.

[0117] A Physical Alias ​​Point (PoPA) is the location in the system where the PAS tag is stripped and the address is translated from a logical physical address back to a system physical address. The PoPA can be located on the system's completer side below the cache, where physical DRAM is accessed using the cryptographic context resolved via the PASTAG. Alternatively, it can be located above the cache to simplify system implementation at the cost of reduced security.

[0118] At any point in time, the world can request a page to be transitioned from one PAS to another. This request is made to monitoring code 29 at EL3, which checks the current state of the GPI. EL3 may allow only specific groups of transitions to occur (e.g., from a non-secure PAS to a secure PAS, but not from a domain PAS to a secure PAS). To provide a clean transition, the system supports a new instruction – “Data Cleanup and Invalidate to Physical Alias ​​Point” – which EL3 can submit before transitioning a page to a new PAS. This ensures that any residual state associated with the previous PAS is flushed from any cache upstream of PoPA 60 (closer to the requester side than PoPA 60).

[0119] Another feature achievable by attaching the GMPU 20 to the host side is efficient memory sharing between worlds. It might be desirable to grant shared access to a physical particle to a subset of N worlds, while preventing other worlds from accessing that physical particle. This can be achieved by adding "restricted sharing" semantics to the particle's protection information, while forcing it to use a specific PAS tag. As an example, the GPI could indicate that a physical particle can only be accessed by "Domain World" 88 and "Secure World" 84, while being tagged with the PAS tag of "Secure PAS" 84.

[0120] An example of the aforementioned characteristics is the ability to rapidly change the visibility properties of specific physical particles. Consider the case where each world is allocated a dedicated PAS that is only accessible to that world. For a specific particle, that world can request to make it visible to the insecure world at any point in time by changing its GPI from "exclusive" to "restricted sharing with the insecure world," without altering the PAS association. This increases the particle's visibility without requiring costly cache maintenance or data copying operations.

[0121] Figure 1 or Figure 3 The device can be implemented as a so-called System-on-Chip (SoC), a so-called Network-on-Chip (NoC), or as a discrete component in various corresponding examples.

[0122] Figure 4 The concept of aliases on physical memory provided to the hardware by the corresponding physical address space is illustrated. As previously described, each of the domains 82, 84, 86, and 88 has its own corresponding physical address space 61.

[0123] At the point when the address translation circuit 16 generates the physical address, the physical address has a value within a certain range 62 supported by the system, which is the same regardless of which physical address space is selected. However, in addition to generating the physical address, the address translation circuit 16 can also select a specific physical address space (PAS) based on the current domain 14 and / or information in the page table entries used to derive the physical address. Alternatively, instead of the address translation circuit 16 performing the PAS selection, an address translation circuit (e.g., MMU) can output the physical address and information derived from the page table entries (PTE) for selecting the PAS, which the PAS filter or GMPU 20 can then use to select the PAS.

[0124] The selection of PAS for a given memory access request can be restricted according to the rules defined in the table below, based on the current domain that the processing circuit 10 is operating on when issuing the memory access request:

[0125]

[0126] For those domains where multiple physical address spaces are available, information from page table entries used to provide access to physical addresses is used to select among the available PAS options.

[0127] Thus, at the point in time when the PAS filter 20 outputs the memory access request to system architectures 24, 8 (assuming it passes any filtering checks), the memory access request is associated with the physical address (PA) and the selected physical address space (PAS).

[0128] From the perspective of memory system components (such as caches, interconnects, snooping filters, etc.) operating prior to the Physical Alias ​​Point (PoPA) 60, the corresponding physical address space 61 is viewed as a completely separate address range corresponding to different system locations within memory. This means that, from the perspective of the pre-PoPA memory system components, the address range identified by a memory access request is actually four times the size of the range 62 that can be output in address translation, because the PAS identifier is actually treated as an additional address bit next to the physical address itself, so that the same physical address PAx can be mapped to multiple alias physical addresses 63 in different physical address spaces 61 depending on which PAS is selected. These alias physical addresses 63 all actually correspond to the same memory system location implemented in the physical hardware, but the pre-PoPA memory system components treat the alias address 63 as a separate address. Thus, if any pre-PoPA cache or snooping filter exists to allocate entries for such addresses, the alias address 63 will be mapped to a different entry with separate cache hit / miss decisions and separate consistency management. This reduces the likelihood or effectiveness of an attacker using a cache or consistency side channel as a mechanism to probe operations in other domains.

[0129] The system may include more than one PoPA 60 (e.g., as discussed below). Figure 14 (as shown in the image).

[0130] At each PoPA 60, the aliased physical address is shrunk to a single de-aliased address 65 in the system physical address space 64. The de-aliased address 65 is provided downstream of any subsequent PoPA component such that the system physical address space 64, which actually identifies the memory system location, is once again the same size as the range of physical addresses that can be output in the address translation performed on the requester side. For example, at PoPA 60, the PAS identifier can be stripped from these addresses, and for downstream components, these addresses can be simply identified using their physical address values ​​without specifying the PAS. Alternatively, for some cases where some completer-side filtering is expected for memory access requests, the PAS identifier may still be provided downstream of PoPA 60, but may not be interpreted as part of the address, so that the same physical address appearing in different physical address spaces 60 will be interpreted downstream of the PoPA as involving the same memory system location, but the supplied PAS identifier can still be used to perform any completer-side security checks.

[0131] Figure 5This illustrates how the granular protection table 56 can be used to divide the system physical address space 64 into blocks of access allocations within a specific architecture physical address space 61. The granular protection table (GPT) 56 defines which portions of the system physical address space 65 are allowed to be accessed from each architecture physical address space 61. For example, the GPT 56 may include multiple entries, each corresponding to a physical address granule of a certain size (e.g., 4K pages), and may be defined as the PAS allocated to that granule, which can be selected from non-secure domains, secure domains, domains, and root domains. By design, if a particular granule or group of granules is allocated to a PAS associated with one of these domains, it can only be accessed within the PAS associated with that domain and not within PASs of other domains. However, it should be noted that while granules allocated to secure PASs (for example) cannot be accessed from within the root PAS, the root domain 82 can access the physical address granule by specifying PAS selection information in its page table to ensure that the virtual address associated with a page in that region of physically addressed memory is translated into a physical address in the secure PAS (not the root PAS). Thus, the sharing of data across domains can be controlled at the point in time when a PAS is selected for a given memory access request (to the extent permitted by the accessibility / inaccessibility rules defined in the previously described table).

[0132] However, in some specific implementations, in addition to allowing access to physical address particles within the allocated PAS defined by the GPT, the GPT can also use other GPT attributes to mark certain regions of the address space as shared with another address space (e.g., an address space associated with a domain of lower or orthogonal privileges (which typically does not allow it to select an allocated PAS for access requests to that domain)). This facilitates temporary sharing of data without requiring changes to the allocated PAS of a given particle. For example, in Figure 5 In the GPT, region 70 of the domain PAS is defined as being allocated to the domain domain and therefore is normally inaccessible from the non-secure domain 86 because the non-secure domain 86 cannot select the domain PAS for its access requests. Since the non-secure domain 26 cannot access the domain PAS, non-secure code typically cannot see the data in region 70. However, if a domain temporarily wishes to share some data in its allocated region of memory with the non-secure domain, it can request the monitoring code 29 operating in the root domain 82 to update GPT 56 to indicate that region 70 will be shared with the non-secure domain 86, and this allows region 70 to also be accessed from the non-secure domain 86. Figure 5The non-secure PAS access shown on the left-hand side does not require changing which domain is the domain allocated to region 70. This improves performance because the operations used to allocate different domains to specific memory regions can be more performance-intensive, involving a greater degree of cache / TLB invalidation and / or data zeroing in memory or data copying between memory regions (which can be improper if the sharing is expected to be only temporary).

[0133] therefore, Figure 5 The arrangement provides an example of a memory with multiple memory partitions, each data memory partition being associated with a partition identifier and having a corresponding physical address range within the physical address space.

[0134] As an example of an access control circuit, the GMPU is configured to detect access control information:

[0135] To detect a region identifier (e.g., a PAS tag) associated with a second memory address, the region identifier being selected from a plurality of region identifiers, each region identifier indicating permission to access a corresponding set of memory partitions, wherein for at least one of the region identifiers, the corresponding set of memory partitions includes a subset of one or more, but not all, memory partitions; and

[0136] The detected region identifier is compared with the partition identifier associated with the second memory address (e.g., PAS identified by the translation circuit).

[0137] Protect the engine

[0138] Figure 6 and Figure 7 A schematic example of a so-called protection engine that can be associated with physical memory 600 is provided.

[0139] The protection engine provides encryption and decryption circuitry for encrypting data stored in memory 600 and decrypting data retrieved from memory 600. The encryption and decryption circuitry is configured to apply corresponding encryption and decryption from a set of encryption and decryption methods to different domains of PAS, such that data encrypted for that domain to a given domain or memory partition cannot be decrypted by applying decryption for another domain.

[0140] The protection engine can utilize PAS tags to apply encryption to data to be stored in memory by applying encryption and decryption selected based on the PAS tag (data indicating the region identifier) ​​associated with the physical memory address, and to apply decryption to decrypt data retrieved from memory at the translated second (physical) memory address.

[0141] refer to Figure 6 Each area is provided with corresponding encryption / decryption circuits 610, 612, 614, 616, and the control / selection circuit 620 selects an appropriate encryption / decryption circuit in response to a PAS tag associated with a specific memory access transaction.

[0142] exist Figure 7 In this embodiment, a single encryption / decryption circuit 700 is provided, wherein the control circuit 710 sets parameters such as encryption / decryption keys and / or encryption / decryption algorithms or algorithm characteristics according to the PAS tag.

[0143] The purpose of using the protection engine is to add another layer of security to the other measures provided here.

[0144] TLB Operation Overview

[0145] As described above, the memory management unit 16 can be associated with the translation back buffer (TLB) 18. The operational aspects of this arrangement are determined by… Figure 8 The schematic flowchart shows that, at step 800, MMU 16 receives a conversion request. At step 810, MMU 16 checks if the required conversion exists in TLB 18. If not, MMU 16 obtains the required conversion using the techniques discussed below and stores it in TLB at step 820.

[0146] Following step 820 or the "yes" result of step 810, at step 830, a data service conversion request is made from the data stored by the TLB.

[0147] MMU Operation Overview

[0148] Address translation occurs between a first memory address (such as a virtual address VA) and a second memory address (such as a physical address PA or an intermediate physical address IPA), and can be performed using a process called page table traversal (PTW). This process involves querying a so-called page table for memory translation information. Page tables are provided as a hierarchical structure, such that an entry accessed in the first page table provides a pointer to the relevant next translation information entry in the next page table.

[0149] Therefore, in the example, the first (input) memory address to the conversion process may include either a virtual memory address or an intermediate physical address; and the second (output) memory address from the process may include either an intermediate physical address or a physical memory address.

[0150] More specifically, the PTW process involves traversing a hierarchical set of so-called page tables to reach a specific VA. In the case of single-stage memory translation, the output can be a PA. In the case of multi-stage memory address translation, the process can be more complex. Accessing the page table itself requires a PA, so a translation stage may be needed to obtain the PA of the next required table at each access to the next table in the hierarchy.

[0151] Figure 9 The diagram schematically illustrates an example of so-called single-stage memory address translation, where the first memory address is a virtual address (VA) 900, and the second memory address is a so-called physical memory address (PA). Figure 9 The valid TLB entry 910 generated by the process at least represents a mapping between VA 900 and the translated PA. The mapping can be represented on a page or other memory region basis, such that a single mapping stored in the TLB as TLB entry 910 maps a contiguous set of virtual addresses to a corresponding contiguous set of physical addresses, for example, mapping a page (e.g., a page of 4k memory address) to a corresponding page of physical address.

[0152] The address of the first page table in the hierarchical structure is provided by the "Translation Table Based Register" (TTBR). The location of the first translation information entry 930 is provided by the memory address defined by the TTBR and at least a portion of the VA 900 to be translated. These two components form the address 920 of the first translation information entry L0[VA] 930. Looking up the first translation information entry 930 provides address information that can be combined with additional bits of the VA 900 to generate address 935 to access the next translation information entry 940. Similarly, the data stored at this translation information entry, concatenated with additional bits of the VA 900, provides the address 945 of entry 950. The translation information stored at entry 950, concatenated with additional bits of the VA 900, provides the address 955 of the final translation information entry 960, where the data stored at entry 960 is concatenated with the final bits of the VA 900 to form a valid TLB entry 910.

[0153] As a working example, the VA to be converted is formed as a 48-bit value. Different parts of the VA are used at different stages of the PTW process.

[0154] To obtain the first entry in the page table hierarchy, the base address stored in the TTBR is obtained. The first part of the VA (e.g., the 9 most significant bits) is added as an offset to the base address to provide the address of the entry in the L0 table. This lookup provides the base address of the L1 table.

[0155] In the second iteration, another part of the VA (e.g., the next 9 bits of the VA [38:30]) is formed as an offset from the base address of the L1 table to provide the address of the entry in the L1 table.

[0156] For example, the process is repeated for L2 and L3 table accesses using the following offset bits [29:21] and [20:12]. Finally, the page table entry in the L3 table provides the page address and potentially provides some access permissions associated with the physical memory page. The remainder of the VA (e.g., the least significant 12 bits [11:0]) provides the page offset within the memory page defined by the last page table entry, but in an exemplary system where information is stored in consecutive four-byte (e.g., 32-bit) portions, portion [11:2] can provide the required offset to the address of the appropriate 32-bit word.

[0157] Page table entries may also provide an indication of whether a page has been written (the so-called "change bit"), an indication of the last time a page was used (the "access bit") to allow cache eviction, etc., without optionally providing other parameters.

[0158] This method of using page tables provides an example of a hierarchical structure of translation information for a given first memory address, including translation information entries, where data indicating the translation information address of the next translation information entry is indicated by the preceding translation information entry. For example, data indicating the translation information address of the next translation information entry could indicate the first memory address applicable to the next translation information entry; and the translation circuitry could be configured to perform a translation operation to generate the corresponding translation information address.

[0159] Two-stage MMU Overview

[0160] In a so-called two-stage MMU, the VA is still translated to the PA, but this is done via a two-stage process, where the VA is translated to a so-called intermediate physical address (IPA), which is then translated to the required PA. The TTBR_EL1 lookup and the stage 1 MMU page table lookup provide the IPA instead of the PA, and each of those IPAs must undergo a stage 2 translation, which may even require looking up the next page table entry.

[0161] Two-stage MMUs are used for various reasons, such as to provide further isolation between the processing element and / or the processes executed on that processing element and the physical memory provided by the entire system. For example, a translation from VA to IPA can be based on page tables (translation information entries) established and controlled by the operating system, for example, at a first security level (such as so-called exception level 1 (EL1)). A translation from IPA to PA can be handled more securely, for example, at a higher security or exception level in the exception level hierarchy (such as EL2), under the control of a so-called hypervisor, such that operations at EL1 cannot access system resources associated with EL2.

[0162] One function of this arrangement is as follows: Figure 9 Each individual stage shown now requires a further translation from IPA to PA, represented by the translation table base address register TTBR_EL1, to access the next translation information entry in physical memory.

[0163] Therefore, refer to Figure 10 Upon receiving VA 1000 for translation, stage 1 TTBR entry 1010 is accessed at EL1. This produces the IPA of the first translation information entry 1020. However, this IPA 1010 must be translated to PA 1015 via stage 2 MMU in order to access entry 1020 in physical memory. The translation to PA involves querying stage 2 TTBR under EL2 and performing a multi-stage page table traversal to generate information 1015, which provides the complete physical address of the next translation table entry 1020 when combined with the bits of VA 1000. Similar processing is required for each level in the hierarchical structure of page table accesses to generate a valid TLB entry 1030.

[0164] Two-stage MMU involving GPT inspection

[0165] Turn now Figure 11 The diagram illustrates the sequence of operations, as involved in a two-stage MMU, in which the two-stage GMPU operates the GMPU in L0GPT and L1GPT to access and obtain permission information for each physical address.

[0166] Considering Figure 11 Each of the operations shown, namely the table lookup for VA conversion, the table lookup for IPA conversion, and the permission lookup performed by the GMPU, requires memory access. Figure 11 The number of visits involved in the process may be very large.

[0167] It should be noted that, Figure 11In some other examples discussed below, the GMP you are examining is referred to as "Stage 3GMPU". This term indicates that it follows the last stage of the MMU transition, and it is used even in the case of a single-stage MMU.

[0168] Single-stage MMU involved in GPT inspection

[0169] Figure 12 A similar arrangement is illustrated schematically, but for a single-stage MMU configuration, where each of the TTBR access and the four page table accesses requires potentially two further memory accesses to L0GPT and L1GPT.

[0170] Memory access cost

[0171] Assuming a “cold” (initially unfilled) TLB, estimates of the number of memory accesses required in various configurations can be derived. In the following example, a working assumption is made: the page table has four levels; however, it should be noted that embodiments of the invention are applicable to various depths or numbers of levels in the page table structure (and the cost may vary upwards for a larger number of levels or downwards for a smaller number of levels, but remains net cost relative to the exemplary embodiments of this disclosure). A related diagram of a four-level page table structure is shown below:

[0172]

[0173] GMPU checks for MMU access can be completely or partially omitted and / or postponed.

[0174] In an exemplary implementation, for certain operations performed by the MMU, at least a portion of the GMPU checks can be omitted or "elided." In other examples, at least a portion of the GMPU checks can be postponed. In either case, the results of the operations can be used before the corresponding GMPU checks have been completed, either because they were postponed from starting or because they were never started.

[0175] Omission and / or postponement can be performed for some, but not all, accesses (i.e., it can be performed selectively), as discussed in the embodiments described below. Omission and / or postponement can be requested or indicated by the MMU, for example, using control signal 21, or can be controlled by the GMPU based on which type of memory access is being initiated by the MMU (this can again optionally utilize control information via connection 21). In such examples, the access circuitry can selectively allow access even when the (full or partial) GMPU check has not yet been completed.

[0176] In the deferred example, the permission circuit is configured to postpone initiating the operation of detecting permission information for the next transition information entry until after access to that next transition information entry is initiated.

[0177] In at least some examples, these operations involve reading translation information performed by the MMU, or at least some of such reading operations. This provides examples where, when access to a translation information address involves a read access, the access circuitry is configured to access the translation information address before the authorization circuitry has completed its authorization information detection operation; and when access to a translation information address involves a write access, the access circuitry is configured to access the translation information address only if the authorization information indicates that memory access to the translation information address is permitted.

[0178] As background to the discussion of these exemplary implementations, it should be noted that the MMU does not actually require information provided by the GMPU from the GPT to form correct page table accesses. Related techniques will be discussed below.

[0179] The MMU hardware itself can be trusted such that the stored content read by the MMU is invisible to the host or other software; that is, individual instances of translation information are used only within the MMU and are not provided as output to external hardware or actual software. In the case of (at least partially) omission of the GPT check, this can provide an example of permission circuitry being configured not to perform permission information detection operations with respect to at least some of the translation information addresses. Furthermore, the translation circuitry is configured not to provide the translation information retrieved from the translation information addresses as output to circuitry outside the translation circuitry, with respect to the translation information addresses where permission information detection has not yet been completed.

[0180] It should be noted that, Figure 11 and Figure 12 In the arrangement shown, the primary performance impact caused by the number of memory accesses pertains to read operations performed by the MMU. Write operations performed by the MMU are slightly less common (e.g., writing to access or change bits of the page descriptor in the page table), and although the omission of the types described herein can apply to both write and read operations, at least in principle, in the example implementation described below, the omission is limited to read operations. This has the potential benefit that the security risk of allowing certain MMU read operations to proceed without a completed GMPU check is perceived to be slightly lower than the security risk associated with allowing the MMU to write data without a proper GMPU check. Therefore, in at least the exemplary implementation, all writes to memory (whether performed by the MMU or initiated by any other aspect of the system) are restricted to being permitted only if verified by the arrangement through a full GMPU check.

[0181] Safety features

[0182] To avoid or at least mitigate security risks by allowing MMU read access to translation information to omit and / or postpone GMPU checks, the following security features can be provided by the hardware design:

[0183] (a) External hardware and software do not have direct access to the data read by the MMU.

[0184] In other words, any data values ​​read from the MMU are guaranteed to remain private within the MMU (in the exemplary implementation). Other exemplary measures that may optionally be applied (individually or collectively) are as follows:

[0185] (b) Translation faults are adequately handled, and any granular protection faults occurring relative to page table (translation information) access are reported to the process operating under EL3. This provides an example where the translation circuit is configured to detect a translation fault with respect to a given translation operation when the use of the translation information by the translation circuit does not provide a valid address translation; and in response to detecting a translation fault, the translation circuit is configured to control the access control circuit to perform an access control information detection operation with respect to any translation information address accessed as part of a given translation operation.

[0186] (c) Memory encryption and decryption can be achieved, for example, through... Figure 6 or Figure 7 The technique utilizes, for example, separate keys and / or algorithms for each world or domain to remain in the proper place. This can mitigate so-called side-channel analysis of the conversion behavior.

[0187] (d) Page table traversal processes suppress read access to the input / output address space. This can be achieved by omitting only a portion of the GMPU checks but retaining a portion of the GMPU checks associated with the memory type of each address (e.g., by providing L0GPT checks but omitting or deferring L1GPT checks), thus allowing MMU access only to memory regions that do not have the “side effects” discussed above. Examples of this technique will be discussed below.

[0188] (e) Omit and / or defer restrictions to MMU read access (that is, provide a full GMPU check for MMU write access).

[0189] Cache and memory access

[0190] Regarding the cache storage device in cache 24 shown in the example above, attempting to access a secure cache line using, for example, a non-secure PAS tag will not even reveal the PA in the cache.

[0191] If data is written to the cache using a "mistaken" PAS, it is harmless because it cannot be subsequently accessed or written back to main memory. Instead, it will simply remain in the cache until it is overridden by the regular cache management and eviction policies operated by the cache itself.

[0192] Another level of security is provided through the previously discussed encryption arrangement, and mentioned at point (c) above. This uses memory encryption associated with each PAS, such that if an “erroneous” PAS tag is associated with a PA, an attempt can be made to decrypt the contents of the memory at a specific address, but that attempt will fail.

[0193] These arrangements provide examples where a cache memory associates a corresponding region identifier with each data item held by the cache memory; and the cache memory is configured to suppress access to a data item associated with a given region identifier in response to a memory access associated with data indicating a different region identifier.

[0194] Example - Single-stage, L1GPT check omitted

[0195] exist Figure 13 In this process, for each proposed access to PA, a Level 0 GPT check (LOGPT) is performed to detect the type of memory region as discussed above, and in particular whether the memory region involves an input / output device or memory that can be read without “side effects” as long as the read operation itself changes the data stored at that address.

[0196] As mentioned above, the GPT data required for this particular check can be relatively compact, for example, one data item per GB, and therefore in the exemplary arrangement, it is cached in a custom cache maintained by the MMU or in the system cache, making the performance penalty of obtaining the LOGPT data for a particular memory access relatively low.

[0197] However, in Figure 13 In the example, except for the final stage that causes the TLB entries to be filled, the L1GPT check is omitted or omitted for all memory accesses involved in a single-stage MMU operation.

[0198] Therefore, this arrangement allows speculative loading of data that has not yet undergone L1GPT checks. To achieve this, for example, a default PAS tag value can be assumed for the data access by associating the default PAS tag with the GMPU associated with the access. In other examples, the PAS tag for page table traversal can be directly derived from the "safe state" associated with the page table, optionally combined with an optional bit in the Phase 1 or Phase 2 page table (this bit is called "NS" to indicate whether the state is "unsafe"). Therefore, in such examples, GPT is not required to submit the correct page table access (e.g., initiated by the present invention). The GMPU in such examples only needs to verify that the PAS tag is a PAS tag "allowed" for the safe state according to the table provided above.

[0199] Provide a final check to verify the final address to be populated into the TLB entry.

[0200] exist Figure 12 Of the approximately 14 memory accesses required in the arrangement, Figure 13 and Figure 12 The potential savings in the comparison involve approximately four memory accesses during address translation. In an exemplary arrangement, the five memory accesses in the remaining memory accesses could be L0GPT data relative to the cache.

[0201] This use of L0GPT checks provides an example of the following: an access circuit is configured to access a translation information address without requiring the permission circuit to have completed the operation of detecting permission information to indicate whether memory access to the translation information address is permitted only if the memory type applicable to the translation information address is the first memory type discussed above (rather than, for example, a second memory type, such as a second memory type that may potentially suffer from "side effects").

[0202] Example - Two-stage, L1GPT check omitted

[0203] Figure 14 It shows the relationship with Figure 11 The layout is similar, but the L1GPT check is omitted again, except that it leads to the final filling of TLB entries. This reduces the aforementioned 74 accesses by 24, with the remaining 25 accesses being relative to the cached L0GPT data.

[0204] Disable omission

[0205] Optionally, the code under EL3 can be disabled for omissions in the event of page table access failures or certain types of failures, such as general protection failures.

[0206] Optionally, if a conversion or other failure occurs, the code under EL3 may require a full GPT check to be re-executed regarding the MMU conversion.

[0207] Another example—single-stage, delayed L1GPT check.

[0208] Figure 15 and Figure 16 Deferred checking is provided, with two arrangements allowing the MMU to load unchecked data with assumed PAS labels. Note that... Figure 15 and Figure 16 Both (for illustrative purposes) involve single-stage MMU operation, but the corresponding techniques can be used for two-stage MMU operation. Note that [the text abruptly ends here, likely due to an incomplete sentence or missing information]. Figure 15 and Figure 16 The relevant L0GPT checks are performed, but these checks are assumed to occur before the relevant address is used in a memory access.

[0209] exist Figure 15 In this example, L0 page table access 1500 is performed before the L1GPT check 1505, which uses the page table base address provided by TTBR_EL1, is completed. In the example shown, the two accesses are initiated roughly simultaneously, but this technique only requires that the check 1505 is not completed when the relevant information is actually used to access the first level of the page table at 1500.

[0210] Figure 15 The example does indeed impose a requirement in some instances where the PAS associated with a read access is verified before the data obtained from the read access can be used or cached. This is indicated by the vertical dashed line 1510 representing the GPT checkpoint. In other words, data read from L0 page table access 1500 is not used further until the address of that data is obtained (i.e., the page table base address in TTBR_EL1 itself is verified).

[0211] Similarly, L1 page table entry access 1520 can be initiated before its address (the output of access 1500) has been verified by step 1525, but in Figure 15 In the example, the next-level address information read by access 1520 cannot be used in subsequent access 1530 until the L1GPT check 1525 associated with the address that performed access 1520 is completed.

[0212] Figure 15The arrangement can be implemented in some examples as a dual or parallel problem of GPT L1 checking and the next page table entry read operation (e.g., a parallel or dual problem of accessing 1505 and 1500). This means that the page table entry read operation is not fully checked at its address and is performed with the assumed PAS. However, by providing a GPT checkpoint (such as checkpoint 1510) upon completion of the read, any memory access failures or translation failures arising from reads of so-called “bad” data are handled synchronously, as these failures occur at or in response to the MMU operation in which they occur. The auxiliary speculative page table entry read is not generated by unchecked data, which can help avoid cache side-channel attacks or other problems.

[0213] exist Figure 15 At the GPT checkpoint in the system, access / change bit updates can be performed (note that, as mentioned above, in the exemplary implementation, any write operation performed by the MMU requires a full L1GPT check prior to the specific implementation), and the TLB can be populated and the cache can be traversed.

[0214] exist Figure 16 In another arrangement, illustrated schematically, loading or reading unchecked data with the assumed PAS value is again permitted, and the loaded data value itself can be used as input for subsequent load operations. However, the assumed PAS value is verified by performing a corresponding L1GPT check before the final data can be committed or cached.

[0215] refer to Figure 16 In the first read operation 1600, the page table base address provided by TTBR_EL1 is used, and the result can be used in the second read operation 1610, and so on up to the fourth read operation 1620 associated with the fourth page table traversal operation. Separately, a chain of L1GPT checks is initiated such that the base address provided by TTBR_EL1 undergoes L1GPT check 1605, and assuming the check passes at checkpoint 1607, it is determined whether the second L1GPT check 1615 of the output of read operation 1600 has passed 1617, etc. The L1GPT check chain continues to L1GPT check 1625 of the address or address portion read from operation 1620, where passing check 1625 1627 is a condition for (a) the filling of TLB entry 1630 and the prefetching of data 1640 at the translated address.

[0216] As mentioned above, in Figure 16 At the GPT checkpoint in the system, access / change bit updates can be performed (note that, as mentioned above, in the exemplary implementation, any write operation performed by the MMU requires a full L1GPT check prior to the specific implementation), and the TLB can be populated and the cache can be traversed.

[0217] Two-stage MMU example

[0218] In a two-stage MMU, any of the techniques described herein can be applied to one stage but not to the other (in either case), or can be applied to both stages.

[0219] Other examples

[0220] Another example of selectively allowing access is as follows.

[0221] The permission circuitry can be selected, or controlled by the conversion circuitry, to select a separate arrangement for each PTE access (or a subgroup of PTE accesses), in other words, to postpone, omit, or retain (without postponing or omitting) the corresponding full or partial permission checks.

[0222] As an example, if both Phase 1 and Phase 2 are enabled, the permission circuit can be configured (either by itself or under the control of the switching circuit):

[0223] • For each Stage 1 MMU read operation, perform the corresponding permission check before the Stage 1 read, or postpone its completion by not exceeding the time point at which the use result drives the subsequent Stage 2 read.

[0224] • For final stage 2 read operations (which obtain data that defines the output address of the requested memory address translation), perform a permission check before the stage 2 read, or postpone its completion by not exceeding the time point before the commit or otherwise use of the output address.

[0225] • For all other stage 2 read operations, permission circuit checks are omitted.

[0226] This specific implementation can suppress ghostly revelation attacks by attackers using a phase 2 table controlled by the attacker as the content of phase 1 checks that can be publicly omitted.

[0227] More generally, different omission and / or deferral patterns can be used, such as random or pseudo-random patterns.

[0228] Overview of Exemplary Technologies

[0229] As discussed above, the various exemplary arrangements contemplate at least the following options and variations, all of which are within the scope of this disclosure as defined in the appended claims:

[0230] a) Check at least some permission information (e.g., GPT) (e.g.) Figure 13 , Figure 14 (The omission of all or part of the text)

[0231] b) Continue accessing transformation information (e.g., PTE) while simultaneously initiating a GPT check, or at least ensure that the GPT check is not completed when the PTE access is initiated.

[0232] c) As in (b), the result of the GPT check of the PTE access is required before utilizing the PTE access of the next MMU operation (e.g., before the next PTE access, information retrieved from the PTE access is initiated). Figure 15 )

[0233] d) As in (b), the results of GPT checks that require PTE access before committing the results (e.g., before initiating any writes to results such as TLB entries or fetches of translated addresses). Figure 16 )

[0234] e) As in any of (b) to (d), where a portion of the GPT check (such as L0GPT) is performed before the relevant PTE access, but in the example, the remaining portion of the GPT check and L1GPT is a portion with deferred completion.

[0235] Example (c) is a so-called “lockstep” variant, in which the GPT check is initiated in parallel with the memory access of the page table traversal, but the GMPU check itself is deferred to a point in time before the result of the memory access is used (e.g., to drive the next traversal).

[0236] Invention Content and Method

[0237] Figure 17 This is a schematic flowchart illustrating a method that includes:

[0238] Performing a conversion operation (at step 1700) to generate a converted second memory address in the second memory address space as a conversion of the first memory address in the first memory address space includes generating the converted second memory address based on conversion information stored at one or more conversion information addresses;

[0239] Perform (at step 1710) the operation of detecting permission information to indicate whether memory access to the given second memory address is permitted for the given second memory address;

[0240] Access (at step 1720) the data stored at the given second memory address when the permission information indicates that memory access to the given second memory address is permitted; and

[0241] Access (at step 1730) the translation information address without requiring the permission circuitry to have completed the operation of detecting permission information to indicate whether memory access to the translation information address is permitted.

[0242] General device features

[0243] Based on the technical operations discussed above Figure 1 and Figure 3 An example of an arrangement is provided, which includes:

[0244] Conversion circuit 16 (50, 52) is configured to perform a conversion operation to generate a converted second memory address in a second memory address space as a conversion of a first memory address in a first memory address space, wherein the conversion circuit is configured to generate the converted second memory address based on conversion information stored at one or more conversion information addresses;

[0245] Permission circuits 20 and 22 are used to perform operations to detect permission information to indicate whether memory access to a given second memory address is permitted; and

[0246] Access circuit 20, configured to access data stored at a given second memory address when permission information indicates that memory access to a given second memory address is permitted; and

[0247] The access circuit is configured to access the translation information address without requiring the permission circuit to have already performed the operation of detecting the permission information to indicate whether memory access to the translation information address is permitted.

[0248] In an exemplary arrangement, the conversion circuit 16 is capable of operating with respect to memory access transactions, each memory access transaction being associated with a first memory address for conversion, and the conversion circuit associating a converted second memory address with each memory access transaction; and

[0249] The permission circuit 20 is configured to perform a detection permission information operation relative to the translated second memory address for each memory access transaction (e.g., L1GPT checks 1300, 1400, 1532, 1625), and the access circuit is configured to provide the memory access transaction with the result of access to the translated second memory address only if the permission data permits access to the translated second memory address.

[0250] Specific implementation of the simulator

[0251] Figure 18A specific implementation of a usable simulator is illustrated. While the previously described embodiments implement the invention in terms of means and methods for operating specific processing hardware supporting the technologies involved, an instruction execution environment according to the embodiments described herein can also be provided, which is implemented using a computer program. Such computer programs are generally referred to as simulators, in part because they provide a software-based implementation of a hardware architecture. Types of simulator computer programs include emulators, virtual machines, models, and binary converters, including dynamic binary converters. Typically, a simulator implementation can run on a host processor 1430 that supports the simulator program 1410, which optionally runs a host operating system 1420. In some arrangements, multiple emulation layers may exist between the hardware and the provided instruction execution environment and / or multiple different instruction execution environments provided on the same host processor. Historically, powerful processors were required to provide simulator implementations that execute at a reasonable speed, but this approach may be reasonable in certain situations, such as when it is desirable to run code native to another processor for compatibility or reuse reasons. For example, a simulator implementation may provide additional functionality to the instruction execution environment that is not supported by the host processor hardware, or provide an instruction execution environment that is typically associated with a different hardware architecture. An overview of the simulation is given in the following literature: “Some Efficient Architecture Simulation Techniques”, Robert Bedichek, Winter 1990 USENIX Conference, pp. 53-63.

[0252] With respect to embodiments previously described with reference to specific hardware constructions or features, in simulated embodiments, equivalent functionality may be provided by suitable software constructions or features. For example, specific circuitry may be implemented as computer program logic in simulated embodiments. Similarly, memory hardware such as registers or cache memory may be implemented as software data structures in simulated embodiments. One or more of the hardware elements referenced in the previously described embodiments are present in an arrangement on host hardware (e.g., host processor 1430), and where appropriate, some simulated embodiments may utilize the host hardware.

[0253] The simulator program 1410 may be stored on a computer-readable storage medium (which may be a non-transitory medium) and provides a program interface (instruction execution environment) to the target code 1400 (which may include applications, operating systems, and management programs), the same as the interface of the hardware architecture modeled by the simulator program 1410. Therefore, the simulator program 1410 can be used to execute the program instructions of the target code 1400 from within the instruction execution environment, allowing the host computer 1430, which does not actually possess the hardware features of the device 2 discussed above, to emulate these features. This is useful, for example, for allowing testing of the target code 1400 being developed for a new version of the processor architecture before a hardware device that actually supports the architecture is available, since the target code can be tested by running it within a simulator executed on a host device that does not support the architecture.

[0254] The simulator code includes processing program logic 1412 that simulates the behavior of the processing circuit 10. For example, it includes instruction decoding program logic that decodes the instructions of the target code 1400 and maps the instructions to corresponding instruction sequences in the native instruction set supported by the host hardware 1430 to perform functions equivalent to the decoded instructions. Processing program logic 1412 also simulates the processing of code in different exception levels and domains as described above. Register emulation program logic 1413 maintains data structures in the host processor's host address space, emulating the architecture register states defined according to the target instruction set architecture associated with the target code 1400. Therefore, it is not as... Figure 1 Instead of storing this architectural state in hardware register 12 as in the example, it is stored in the memory of the host processor 1430, where register emulation logic 1413 maps register references of instructions in target code 1400 to corresponding addresses to obtain simulated architectural state data from host memory. This architectural state may include the previously described current domain indicator 14 and current exception level indicator 15.

[0255] The simulation code includes address translation logic 1414 and filtering logic 1416, which respectively reference the same page table structure as described above and the functionality of the GPT 56 emulated address translation circuit 16 and PAS filter 20. Therefore, address translation logic 1414 translates the virtual address specified by target code 1400 into a simulated physical address within the PAS (from the target code's perspective, this refers to a physical location in memory), but in reality, these simulated physical addresses are mapped to the host processor's (virtual) address space by address space mapping logic 1415. Filtering logic 1416 performs a lookup of particle protection information in the same manner as the aforementioned PAS filter to determine whether memory access triggered by the target code is allowed to proceed.

[0256] therefore, Figure 18 The arrangement provides an example of a computer program for controlling a host data processing device to provide an instruction execution environment for executing object code. The computer program includes:

[0257] The conversion logic is used to perform a conversion operation to generate a converted second memory address in a second memory address space as a conversion of a first memory address in a first memory address space, wherein the conversion logic is configured to generate the converted second memory address based on conversion information stored at one or more conversion information addresses.

[0258] The permission logic is used to perform operations that detect permission information to indicate whether memory access to a given second memory address is permitted; and

[0259] Access logic, which allows access to data stored at a given second memory address when permission information indicates that memory access to a given second memory address is permitted;

[0260] The access logic is configured to selectively allow the translation logic to access the translation information address without requiring the permission logic to have completed the operation of detecting permission information to indicate whether memory access to the translation information address is permitted.

[0261] In this application, the phrase "configured as..." is used to mean that the elements of the device have a configuration capable of performing the defined operations. In this context, "configuration" means the arrangement or manner of interconnection of hardware or software. For example, the device may have dedicated hardware that provides the defined operations, or a processor or other processing device may be programmed to perform the function. "Configured as" does not mean that the elements of the device need to be changed in any way to provide the defined operations.

[0262] While exemplary embodiments of the invention have been described in detail herein with reference to the accompanying drawings, it should be understood that the invention is not limited to those precise embodiments, and various changes and modifications can be made therein by those skilled in the art without departing from the scope of the invention as defined by the appended claims.

Claims

1. A data processing apparatus, comprising: A conversion circuit is configured to perform a conversion operation to generate a converted second memory address in a second memory address space as a conversion of a first memory address in a first memory address space, wherein the conversion circuit is configured to generate the converted second memory address based on conversion information stored at one or more conversion information addresses. A permission circuit, the permission circuit being used to perform an operation of detecting permission information to indicate whether memory access to a given second memory address is permitted; and An access circuit, the access circuit being configured to allow access to data stored at the given second memory address when the permission information indicates that memory access to the given second memory address is permitted; The access circuit is configured to selectively allow access to the translation information address by the translation circuit without the permission circuit having already performed the operation of detecting permission information to indicate whether memory access to the translation information address is permitted.

2. The apparatus according to claim 1, wherein: When the access to the translation information address involves a read access, the access circuit is configured to selectively allow access to the translation information address by the translation circuit without the permission circuit having already completed the operation of detecting permission information; and When the access to the translation information address involves a write access, the access circuitry is configured to allow access to the translation information address only if the permission information indicates that memory access to the translation information address is permitted.

3. The apparatus of claim 1 or claim 2, wherein the authorization circuit is configured to perform another operation of detecting a storage type applicable to a given second memory address, at least a first storage type or a different second storage type applicable to the given second memory address.

4. The apparatus of claim 3, wherein the access circuitry is configured to selectively allow access to the translation information address by the translation circuitry without the permission circuitry having already performed the operation of detecting permission information to indicate whether memory access to the translation information address is permitted only if the storage type applicable to the translation information address is the first storage type.

5. The apparatus of claim 4, wherein the first storage type is a storage type in which the data stored at a given address is not altered by a read operation from the given address.

6. The apparatus according to claim 1, wherein: The permission circuit is configured to perform the operation of detecting permission information without regard to at least some of the conversion information addresses; and The conversion circuit is configured not to provide conversion information retrieved from the conversion information address as an output to a circuit outside the conversion circuit, and the operation of the detection permission information has not yet been completed regarding the conversion information address.

7. The apparatus of claim 1, wherein the conversion information applicable to a conversion of a given first memory address comprises a hierarchical structure of conversion information entries, wherein data indicating the conversion information address of the next conversion information entry is indicated by the preceding conversion information entry.

8. The apparatus according to claim 7, wherein: The data indicating the translation information address of the next translation information entry is a first memory address applicable to the next translation information entry; and The conversion circuit is configured to perform the conversion operation to generate a corresponding conversion information address.

9. The apparatus of claim 7 or claim 8, wherein the permission circuit is configured to postpone initiating the operation of detecting permission information for the next conversion information entry until after access to the next conversion information entry has been initiated.

10. The apparatus according to claim 1, wherein: The conversion circuit is capable of operating on memory access transactions, with each memory access transaction associated with a first memory address for conversion, and the conversion circuit associating a converted second memory address with each memory access transaction; and The permission circuit is configured to perform the operation of detecting permission information relative to the translated second memory address for each memory access transaction, and the access circuit is configured to provide the memory access transaction with the result of access to the translated second memory address only if the permission data permits access to the translated second memory address.

11. The apparatus according to claim 1, wherein: The first memory address includes either a virtual memory address or an intermediate physical address; and The second memory address includes an intermediate physical address or a physical memory address.

12. The apparatus of claim 11, comprising a memory having a plurality of memory partitions, each data memory partition being associated with a partition identifier and having a corresponding physical address range within a physical address space.

13. The apparatus of claim 12, wherein the permission circuit, if configured, is configured to perform the operation of detecting permission information: To detect a region identifier associated with a second memory address, the region identifier being selected from a plurality of region identifiers, each region identifier indicating permission to access a corresponding set of the memory partitions, wherein for at least one of the region identifiers, the corresponding set of the memory partitions includes a subset of one or more, but not all, of the memory partitions; and The detected region identifier is compared with the partition identifier associated with the second memory address.

14. The apparatus of claim 13, comprising: An encryption and decryption circuit, wherein the encryption and decryption circuit is used to encrypt data stored in the memory and to decrypt data retrieved from the memory; The encryption and decryption circuitry is configured to apply a corresponding encryption and corresponding decryption from a set of encryption and corresponding decryption to each memory partition, the set of encryption and corresponding decryption such that data encrypted by the corresponding encryption for a given memory partition cannot be decrypted by applying the decryption to another memory partition.

15. The apparatus of claim 14, wherein the authorization circuitry is configured to associate data indicating the region identifier associated with the translated second memory address with the translated second memory address.

16. The apparatus of claim 15, wherein the encryption and decryption circuitry is configured to apply decryption by applying decryption selected according to the data indicating the region identifier associated with the translated second memory address to decrypt data retrieved from the memory at the translated second memory address.

17. The apparatus of claim 15 or claim 16, comprising one or more cache memories to hold data retrieved from the memory and / or stored in the memory; The cache memory associates a corresponding region identifier with each data item held by the cache memory; The cache memory is configured to suppress access to data items associated with a given region identifier in response to memory access associated with data indicating a different region identifier.

18. The apparatus according to claim 1, wherein: The conversion circuit is configured to detect a conversion failure with respect to a given conversion operation when the use of the conversion information by the conversion circuit does not provide a valid address translation. and In response to the detection of a conversion failure, the conversion circuit is configured to control the permission circuit to perform the operation of the detection permission information with respect to any conversion information address accessed as part of the given conversion operation.

19. The apparatus according to claim 1, comprising: A processor, the processor being configured to execute program instructions at the current exception level in a hierarchical structure of exception levels, each exception level being associated with a security privilege such that instructions executed at a higher exception level can access resources that cannot be accessed by instructions executed at a lower exception level. The processor needs to execute instructions at the highest exception level among the exception levels in order to set the data for the permission circuit to detect permission information.

20. A data processing method, comprising: Performing a conversion operation to generate a converted second memory address in a second memory address space as a conversion of a first memory address in a first memory address space includes generating the converted second memory address based on conversion information stored at one or more conversion information addresses; Perform the operation of detecting permission information to indicate whether memory access to the given second memory address is permitted for the given second memory address; When the permission information indicates that memory access to the given second memory address is permitted, the data stored at the given second memory address is accessed. as well as Optionally accessing a translation information address without having to complete the operation of detecting the permission information to indicate whether memory access to the translation information address is permitted.