Apparatus, method and computer program for performing translation table entry load / store operations
The conversion table entry loading/storage circuit directly identifies the target conversion table entries using software-defined address information, which solves the problem of low performance in the conversion table structure update in the prior art, and improves the efficiency of conversion table updates, especially during virtual machine migration.
Patent Information
- Application Number
- CN202380085837.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-20
- Filing Date
- 2023-11-17
- Publication Date
- 2025-07-22
AI Technical Summary
When updating the conversion table structure, software is required to roam the multi-level conversion table structure to identify the conversion table entry address to be updated, resulting in poor performance, especially when a large number of conversion table entries need to be updated during virtual machine migration, which affects performance.
Through the conversion table entry loading/storage circuit, the target conversion table entries are identified using software-defined address information, avoiding the roaming of explicit instruction sets in the hardware, and directly performing the conversion table entry loading/storage operation, and improving efficiency using the hardware cache and address conversion circuit.
Improves the performance of conversion table structure updates, reduces the overhead of software roaming, and improves the efficiency of processors when updating conversion table entries, especially in virtual machine migrations and other software use cases.
Smart Images

Figure CN120359502A_ABST
Abstract
Description
[0001] This technology relates to the field of data processing.
[0002] A processing system can perform address translation to convert between an input address space and an output address space based on address mapping information obtained from a translation table structure including translation table entries. By supporting address translation, different software written using conflicting input address space definitions can be mapped onto a common output address space to resolve any address conflicts between such software. The translation table structure can also specify access permissions or memory region attributes for controlling access to the memory of corresponding regions of the address space.
[0003] At least some examples provide an apparatus including: processing circuitry for processing instructions; address translation circuitry for converting between an input address space and an output address space based on address mapping information obtained from a translation table structure including translation table entries; and translation table entry load / store circuitry for performing a translation table entry load / store operation for at least one target translation table entry address in response to a translation table entry load / store trigger instruction processed by the processing circuitry, the at least one target translation table entry address being selected according to software-defined address information identifying a selected address in the input address space, each target translation table entry address including the address of a leaf translation table entry providing the address mapping information for converting the selected address from the input address space to the output address space or the address of a branch translation table entry traversed in a translation table walk operation for obtaining the leaf translation table entry providing the address mapping information for converting the selected address; wherein for at least one variant of the translation table entry load / store operation, for a given target translation table entry among the at least one target translation table entries, the translation table entry load / store operation supports clearing access tracking metadata of the given target translation table entry from a first state indicating that at least one load / store access to a corresponding region of the input address space has occurred to a second state indicating that no load / store access to the corresponding region of the input address space has occurred.
[0004] At least some examples provide a method that includes: using a processing circuit to process instructions; and using an address translation circuit to perform a translation between an input address space and an output address space based on address mapping information obtained from a translation table structure including translation table entries; and in response to the processing circuit processing a translation table entry load / store trigger instruction, performing a translation table entry load / store operation for at least one target translation table entry address, the at least one target translation table entry address being selected according to software-defined address information identifying a selected address in the input address space, each target translation table entry address including the address of a leaf translation table entry providing the address mapping information for translating the selected address from the input address space to the output address space or the address of a branch translation table entry traversed in a translation table walk operation for obtaining the leaf translation table entry providing the address mapping information for translating the selected address; wherein for at least one variant of the translation table entry load / store operation, for a given target translation table entry among the at least one target translation table entries, the translation table entry load / store operation supports clearing access tracking metadata of the given target translation table entry from a first state indicating that at least one load / store access to a corresponding region of the input address space has occurred to a second state indicating that no load / store access to the corresponding region of the input address space has occurred.
[0005] At least some examples provide a computer program for controlling a host data processing device to provide an instruction execution environment for executing target code, the computer program comprising: address translation program logic for translating between an input address space and an output address space based on address mapping information obtained from a translation table structure including translation table entries; and translation table entry load / store program logic for performing a translation table entry load / store operation for at least one target translation table entry address in response to a translation table entry load / store trigger instruction of the target code, the at least one target translation table entry address being selected according to software-defined address information identifying a selected address in the input address space, each target translation table entry address including the address of a leaf translation table entry providing the address mapping information for translating the selected address from the input address space to the output address space or the address of a branch translation table entry traversed in a translation table walk operation for obtaining the leaf translation table entry providing the address mapping information for translating the selected address; wherein for at least one variant of the translation table entry load / store operation, for a given target translation table entry in the at least one target translation table entry, the translation table entry load / store operation supports clearing access trace metadata of the given target translation table entry from a first state indicating that at least one load / store access has occurred to a corresponding region of the input address space to a second state indicating that no load / store access has occurred to the corresponding region of the input address space.
[0006] The computer program may be stored on a storage medium. The storage medium may be a transient storage medium or a non-transient storage medium.
[0007] Additional aspects, features, and advantages of the present technology will become apparent from the following description of examples read in conjunction with the accompanying drawings, in which:
[0008] Figure 1 An example of a data processing device is illustrated;
[0009] Figure 2 An example of an execution state of a processing circuit is illustrated;
[0010] Figure 3 An example of two-stage address translation is illustrated;
[0011] Figure 4 and Figure 5 Examples of translation table walks for stage 1 and stage 2 address translations are illustrated respectively;
[0012] Figure 6 Examples of branch translation table entries and leaf translation table entries are illustrated;
[0013] Figure 7 Illustrates examples of load and store instructions;
[0014] Figure 8 Illustrates examples of a translation table entry load instruction and a translation table entry store instruction, both of which are examples of translation table entry load / store trigger instructions;
[0015] Figure 9 Illustrates a method of processing instructions;
[0016] Figure 10 Illustrates a method of performing a translation table entry load / store operation in response to a translation table entry load / store trigger instruction;
[0017] Figure 11 Shows an example of a translation table entry load / store circuit;
[0018] Figure 12 Illustrates an example of a programming interface for a translation table entry load / store circuit;
[0019] Figure 13 Illustrates a method of controlling asynchronous processing of a translation table entry load / store operation based on parameters defined in a programming interface; and
[0020] Figure 14 Illustrates an emulator implementation.
[0021] An apparatus has: a processing circuit for processing instructions; and an address translation circuit for translating between an input address space and an output address space based on address mapping information obtained from a translation table structure including translation table entries. In a typical instruction set architecture, for software to modify information within a translation table entry of such a translation table structure, the software will use general load / store and arithmetic instructions to read the current value of each entry to be modified, determine what the modified value of each entry should be, and write each modified translation table entry back to memory. The load / store instructions used in such an update of the translation table structure will typically be general load / store instructions that specify the address of the memory location to be read / written in the load / store operation as the target address of the load / store. Thus, when such load / store instructions are used to implement a translation table entry update, software will need to identify the address of each translation table entry to be updated before performing the load / store on the address of each translation table entry to be updated. Identifying the address in software of the location storing a given translation table entry is typically not straightforward because it may require software traversal of a multi-level translation table structure, which may require the software to perform a relatively long sequence of memory accesses that are only relevant for identifying the address of the translation table entry to be updated for each translation table entry to be updated. This can be slow and may degrade performance, especially when many translation table entries need to be updated. An example of a scenario where this problem occurs is during the live migration of a virtual machine from one host processor to another, when most of the translation table entries associated with the migrated virtual machine may need to be updated during the migration.
[0022] In the example discussed below, a translation table entry load / store circuit performs a translation table entry load / store operation for at least one target translation table entry address in response to a translation table entry load / store trigger instruction processed by the processing circuit, the at least one target translation table entry address being selected according to software-defined address information identifying a selected address in the input address space. Each target translation table entry address includes the address of a leaf translation table entry providing the address mapping information for translating the selected address from the input address space to the output address space or the address of a branch translation table entry traversed in a translation table walk operation for obtaining the leaf translation table entry providing the address mapping information for translating the selected address.
[0023] Accordingly, an instruction that triggers a load / store operation to the address of a target translation table entry does not itself need to specify the address of the target translation table entry. Instead, the translation table entry load / store circuitry can use software-defined address information that identifies a selected address to identify the target translation table entry to be updated. Instead of triggering a load / store to the selected address itself, the translation table entry load / store circuitry performs a load / store operation on the address of at least one target translation table entry corresponding to the selected address. The at least one target translation table entry can include a leaf target translation table entry that provides address mapping information for translating the selected address and / or at least one branch translation table entry that provides a pointer for use in a translation table walk to locate the leaf translation table entry. The operation of identifying the target translation table entry address based on the selected address can be performed in hardware without an explicit instruction set in software for walking the translation table entry structure to locate the address of the entry to be updated. This can improve the performance of updating the translation table structure.
[0024] For at least one variant of the translation table entry load / store operation, for a given target translation table entry in the at least one target translation table entry, the translation table entry load / store operation supports clearing access trace metadata of the given target translation table entry from a first state that indicates that at least one load / store access to a corresponding region of an input address space has occurred to a second state that indicates that no load / store access to the corresponding region of the input address space has occurred. One or more items of access trace metadata can be specified in a given translation table entry, and each item of access trace metadata can track load / store accesses in a combined manner (e.g., only track whether any access to the region has occurred, regardless of whether it is a load or a store), or can be specific to tracking load accesses or store accesses (e.g., tracking store accesses can specifically be used to identify which regions of the address space may be "clean" (not written), and thus may not need to be written back to another storage device if evicted from the memory device currently storing the region of the address space). Accordingly, there can be one or more access trace metadata that can be recorded in the translation table structure to allow software to track the access patterns of how the memory address space is accessed on a region-by-region basis.
[0025] Typically, it is the responsibility of the software to determine when to clear any access trace metadata back to an initial state that indicates that no load / store access has occurred to the corresponding region of the input address space (e.g., this can be done at the start of a period for collecting access frequency or dirty state tracking information for a group of memory regions). In a typical instruction set architecture, this would require general-purpose load / store instructions of the type discussed above. However, when clearing access trace metadata, many translation table entries may need to have their access trace metadata cleared. Thus, clearing only the access trace metadata for the monitored regions of the address space to a first state can incur a significant performance cost, as each entry to be updated may require a software-managed walk of only the translation table structure to identify the address where the corresponding translation table entry is stored.
[0026] Accordingly, by supporting translation table entry load / store operations that, in at least one variant, support clearing access trace metadata using an instruction that does not require specifying the address of the translation table entry itself, but can use software-defined address information specifying the selected address that is used to translate the selected address itself or to provide a pointer along the path to a leaf translation table entry that translates the selected address, performance can be improved for a variety of software use cases.
[0027] In response to a translation table entry load / store trigger instruction, the translation table entry load / store circuit may control the address translation circuit to identify at least one target translation table entry address based on the selected address. For example, the address translation circuit may include a translation table walk circuit that, in response to a given address, may trigger a walk of the translation table structure to identify a series of memory accesses required to identify the address of the translation table entry for translating the given address. The address translation circuit may also include at least one address translation cache that caches information derived from a previous translation table walk performed by the translation table walk circuit. The cached information in the address translation cache may include information identifying a mapping between the input memory addresses used for the translation table walk and the translation table entry addresses of the corresponding translation table entries for translating those input memory addresses. Thus, the address translation circuit may generally already have hardware for identifying the address of the target translation table entry related to the translation of a given address in the input address space. This hardware can be efficiently reused for translation table entry load / store operations to avoid the software needing to trigger a software-specified related series of translation table walks via memory accesses. The hardware will generally perform the translation table walk faster than software, and in any case, if the relevant address information is already stored in the address translation cache, some portions of such translation table walks may be able to be eliminated by the hardware. In contrast, a software-managed translation table walk is less likely to benefit from such caching, and in any case, even in the absence of a cache, the software-managed translation table walk will generally be slower than the hardware-managed translation table walk. Therefore, it may be useful to provide a translation table entry load / store operation as an architecturally defined operation available to software that takes a selected address and triggers a load / store operation on the address of the corresponding translation table entry corresponding to that selected address, where the responsibility for identifying the mapping between the selected address and the translation table entry address is left to the hardware of the address translation circuit.
[0028] In response to a translation table entry load / store trigger instruction, there is no need to trigger any load / store operation on the selected address specified by software-defined address information. Thus, the translation table entry load / store trigger instruction may cause the translation table entry load / store circuit to trigger a load / store on at least one target translation table entry address without triggering any load / store on the selected address itself.
[0029] Different specific implementations may support different variants of the translation table entry load / store operation. Some specific implementations may support only a single variant, but the variant may vary depending on the specific implementation. Other specific implementations may support two or more different variants, where at least one software-defined information is used to distinguish which specific variant will be executed in response to a given instance of the translation table entry load / store trigger instruction. For example, the specific variant of the operation to be executed in response to a given instance of the translation table entry load / store trigger instruction may be set by any of the following architectural features:
[0030] ● Parameters of the translation table entry load / store trigger instruction itself (e.g., the opcode of the instruction or another instruction field within the encoding of the instruction).
[0031] ● Presence / absence of a "prefix" instruction executed before the translation table entry load / store trigger instruction, where if present, the "prefix" instruction modifies the behavior of the operation to be executed in response to a subsequent instance of the translation table entry load / store trigger instruction. The prefix instruction may identify which variant of the operation will be executed by the subsequent translation table entry load / store trigger instruction. If no prefix instruction is included before the translation table entry load / store trigger instruction,
[0032] then the translation table entry load / store operation may be executed according to the default variant.
[0033] ● Control information stored in a register that identifies which variant of the operation will be executed. The register providing the control information may be, for example, a general-purpose register referenced by the instruction, or a predetermined system register not explicitly referenced by the instruction.
[0034] ● Information stored at a given address in memory (e.g., in a memory-mapped register or in an entry of a buffer structure referenced based on a software-programmed base address), which can be read by the translation table entry load / store circuit to identify the variant of the translation table entry load / store operation to be executed.
[0035] ● Information specified within a selected address defined by software-defined address information. For example, since the selected address is used to represent a translation table entry, and translation table entries are typically defined on a per-region basis, some address bits of the selected address may be sub-region bits that only distinguish different addresses within a region all corresponding to the same translation table entry. Therefore, those sub-region bits are not meaningful for identifying which translation table entry corresponds to the selected address and can thus be reused to encode operation variant information that identifies the variant of the translation table entry load / store operation to be executed for a given instance of the translation table entry load / store trigger instruction.
[0036] One way in which variants of the translation table entry load / store operation can vary can be with respect to how, based on the selected address, it is determined which translation table entry is identified as the target translation table entry.
[0037] In some particular implementations, for at least one variant of the translation table entry load / store operation, at least one translation table entry includes a leaf translation table entry. In such a case, the translation table entry loaded or stored by the translation table entry load / store operation can be the entry that provides the address translation mapping for translating the selected address. In such variants, in addition to allowing clearing of access trace metadata, the translation table entry load / store operation can also support updating the leaf translation table entry to change at least one of the following: the address mapping information for translating the selected address; the access permission information indicating which types of memory access operations are permitted; and the memory attribute information for controlling the disposition of memory access to the selected address.
[0038] In some particular implementations, for at least one variant of the translation table entry load / store operation, at least one translation table entry includes a translation table entry at a specified level of the translation table structure, whether the translation table entry is a leaf translation table entry or a branch translation table entry. For example, in a translation table structure having a certain maximum number of levels (e.g., 4 levels), leaf translation table entries (which provide address translation mappings rather than pointers to additional translation tables for providing additional translation table entries) may be defined at different levels of the table structure. For example, an entry at level 2 can be encoded as a leaf translation table entry to indicate address mapping information for a larger address space region than would be the case if the entry at level 2 were a branch translation table entry pointing to another translation table and the leaf translation table entry were defined at level 3. Thus, variants of the translation table entry load / store operation can be provided where the target translation table entry to be loaded or stored is at the specified level of the translation table structure reached in the translation table walk for the selected address, whether the entry at that level is a branch or a leaf. For example, this option can be useful when software wishes to set new translation table information for a given block of memory corresponding to a specified table level, whether that block of memory was previously defined with general characteristics for the entire block (using a leaf translation table entry at the specified level) or was subdivided into smaller blocks (with different characteristics, or the same compatible characteristics, where a branch translation table entry at the specified level points to an additional table that can define separate entries for each subdivision).
[0039] In some specific implementations, for at least one variant of the translation table entry load / store operation, at least one translation table entry includes a leaf translation table entry and each branch translation table entry traversed in the translation table walk operation used to obtain the leaf translation table entry. Thus, the method can apply the translation table entry load / store operation in a hierarchical manner to each translation table entry on the path of the translation table structure traversed in the translation table walk at the selected address, such that loading or storing can be performed on two or more translation table entries in response to the same translation table entry load / store trigger instruction. For example, in some specific implementations of the translation table structure, the access trace metadata mentioned above can exist elsewhere at various levels of the table, and not just in the leaf entries, and thus it can be used to provide a variant of the translation table entry load / store operation that can, in a single instruction, clear the access trace metadata in the leaf translation table entry corresponding to the selected address and in all branch translation table entries traversed in the path to the leaf translation table entry. Thus, for some variants of the translation table entry load / store operation, based on a single selected address, two or more addresses of the corresponding translation table entries can each undergo the translation table entry load / store operation.
[0040] In some specific implementations, for at least one variant of the translation table entry load / store operation, at least one translation table entry includes: the leaf translation table entry when the leaf translation table entry is effectively defined for the selected address; and, when a valid leaf translation table entry is not defined for the selected address, the final valid branch translation table entry reached in the traversal of the translation table structure at the selected address. Thus, with this method, the loaded / stored entry can be the final valid translation table entry that can be reached in the translation table walk at the selected address, whether the final valid translation table entry is a leaf entry or a branch entry, and regardless of which level of the translation table structure the final valid translation table entry appears in. For example, this can be helpful for software when constructing a new translation table structure.
[0041] Another way in which variants of the translation table entry load / store operation can differ can be in what type of load / store operation is performed on a given target translation table entry whose address is identified based on the selected address.
[0042] For example, at least one variant of the translation table entry load / store operation that supports clearing of access trace metadata can include at least one of the following:
[0043] ● The store variant of the translation table entry load / store operation, where this store variant of the translation table entry load / store operation is used to update a given target translation table entry with an updated value specified by a store data operand. For example, the store variant can write a new value (specified as an operand of the operation) to a memory location corresponding to the address of the given target translation table entry.
[0044] ● The exchange variant of the translation table entry load / store operation, where this exchange variant of the translation table entry load / store operation is used to update a given target translation table entry with an updated value specified by an exchange data operand and to load a software-accessible location with the pre-update value or the post-update value of the given target translation table entry. This variant can be useful if, in addition to updating the given target translation table entry, subsequent operations that depend on the information within the entry need to be performed.
[0045] ● The atomic compare-and-swap variant of the translation table entry load / store operation, where this atomic compare-and-swap variant of the translation table entry load / store operation is used to determine whether the result of a comparison between a given translation table entry and a compare operand satisfies a comparison condition, and in response to determining that the result of the comparison satisfies the comparison condition, to update the given target translation table entry based on the exchange operand of the compare-and-swap variant. This operation is performed atomically such that the result of the translation table entry load / store operation is consistent with the result that would occur if a read operation to obtain the value to be compared with the compare operand and a write operation to write the updated value to the given target translation table entry occurred, with no intervening write to the address of the given target translation table entry between the read and write of the compare-and-swap.
[0046] ● The atomic bit update variant of the translation table entry load /
[0047] store operation, where this atomic bit update variant of the translation table entry load / store operation is used to set or clear one or more specified bits of a given translation table entry, where the specified bits are identified by a bit selection operand (e.g., an index identifying a single bit to be updated or a mask identifying for each bit whether that bit should be updated). Depending on the variant of the operation, the bit update can set one or more specified bits to 1 or 0. In some embodiments, the bit update operation may require a read of the address of the given translation table entry followed by a write to that address, because the write can be performed at a word granularity greater than one bit, and thus a read operation may be required to read the other bits of the same word as the bit to be updated so that a new value including the updated bit can be written back (where the other bits remain the same as previously read). Since there can be separate read and write operations, atomicity can again be enforced on the bit update operation to ensure that the result is consistent with the result that would occur if there were no intervening write to the address of the given target translation table entry between the read and write of the atomic bit update operation.
[0048] For atomic compare-and-swap variants and atomic bit-update variants, atomicity can be implemented in different ways, such as by locking access to the relevant memory location to prevent intermediate write operations during the period between the read and the write, or by speculatively allowing the read and write to proceed by assuming that there are no intermediate write operations (without locking access to the location), but providing techniques for detecting intermediate writes to the location such that the translation table entry load / store operation can be cancelled and repeated to ensure atomicity in the event that an intermediate write is detected.
[0049] For each of the store variants, swap variants, atomic compare-and-swap variants, and atomic bit-update variants described above, these variants may be able to clear the access trace metadata as mentioned above. These variants may also be able to perform other kinds of updates to the translation table entries, such as updating the address mapping information, access permission information, or memory attribute information as previously discussed. The operands of these variants of the translation table entry load / store operation may define the specific values to be written to the target translation table in order to specify which type of information within the entry is to be updated. These variants may also support setting the access trace metadata to a first state.
[0050] In addition to the variants that support clearing of the access trace metadata as mentioned above, the translation table entry load / store circuitry may also support a load variant of the translation table entry load / store operation to load at least one target translation table entry into at least one software-accessible register. This can be useful such that software can perform more complex manipulation of the loaded translation table entry using conventional arithmetic instructions that operate on the data stored in the software-accessible register. The previously mentioned store variant or compare-and-swap variant of the translation table entry load / store operation can later be used to write back the updated value of the translation table entry once those operations are completed.
[0051] In response to a translation table entry load / store trigger instruction, the translation table entry load / store circuitry is configured to perform an error reporting action in response to identifying that an error condition has occurred, the error condition including one of the following: no valid leaf translation table entry is defined for the selected address; and a valid leaf translation table entry for the selected address is defined at a level outside of the expected level of the translation table structure. The expected level may be defined as an operand of the translation table entry load / store operation, or may be part of the definition of which variant of the operation is being performed.
[0052] Support for error responses may be useful because sometimes software may request the load / store of a translation table entry using one of the variations of the operations discussed above, but in fact the current configuration of the translation table structure may be different from what the software expects, such that applying the load / store to the entry returned by the operation may risk an error occurring, which may lead to inappropriate setting of translation table information.
[0053] For example, software may expect a leaf translation table entry to be defined for a given address, but if no such leaf translation table entry exists, applying an update to some other location (e.g., an address that is not a valid translation table entry at all) may risk leaking information about the address space layout to a software process that may not be trusted to view the mapped address. This may also risk other errors in memory access control (e.g., if an update address mapping expected to be applied to a leaf entry is incorrectly applied to a branch entry, this may cause the mapped address for the expected leaf to be alternatively treated as a pointer to another translation table, which may cause subsequent levels of page table walks to give incorrect results. Similarly, software may expect it is setting update information for a given block of memory of a size corresponding to a leaf entry expected at one level of the table, but if the leaf is found at a different level of the table than the expected level, applying the update at the wrong level may cause the update information to be applied to a memory address region that is smaller or larger than the expected size. Thus, by providing a mechanism for reporting the following errors: no valid leaf is identified for the selected address or the valid leaf is at a wrong level of the translation table structure than the expected level, this can reduce the risk of inappropriate setting of translation table information. Additionally, supporting error responses allows the use of instructions without the need for a coarse-grained locking structure in software, thus providing a performance benefit.
[0054] The error reporting action can be implemented in a variety of ways. For example, the error reporting action may include signaling a fault (raising an exception), which may cause the processing of the current software to be interrupted, or setting an error status indication in a software-accessible register or other storage location to identify that an error has occurred.
[0055] In some examples, in response to a translation table entry load / store triggering instruction, the translation table entry load / store circuit is configured to update at least one software-accessible register with comprehensive information that specifies one of the following: the level of the translation table structure that defines a valid leaf translation table entry for the selected address; and the information specified by at least one target translation table entry. For example, the comprehensive information may be written to the software-accessible register regardless of whether any error has occurred. Even in a no-error scenario, it may still be useful for software to know the information from the translation table entry that the translation table entry load / store operation is targeted at.
[0056] The processing circuit can support a given instruction set architecture. The translation table entry load / store trigger instruction can be implemented in different ways within the instruction set architecture.
[0057] In some examples, the translation table entry load / store trigger instruction designates a selected address as an operand of the translation table entry load / store trigger instruction, and the translation table entry load / store trigger instruction has an instruction opcode different from that of the load / store instruction, for triggering a load / store operation to be performed for the address designated as the operand of the load / store instruction. The operand defining the selected address can be an immediate operand directly specified in the encoding of the instruction, or can be an operand stored in a register referenced by a register field specified in the encoding of the instruction.
[0058] Accordingly, dedicated instructions (or multiple dedicated instructions, supporting different variants of the translation table entry load / store operation) can be defined in the instruction set architecture used by the processing circuit, separate from the general-purpose load / store instructions that apply their load / store to the address designated as the operand of the instruction itself rather than to the address of the corresponding translation table entry.
[0059] By way of this example, the translation table entry load / store circuit can be the general-purpose load / store unit provided in the processing circuit for handling general-purpose load / store instructions, but the processing circuit can have a mechanism for (based on whether a translation table entry load / store instruction or a general-purpose load / store instruction is being executed) differentiating whether the address for which the general-purpose load / store unit initiates a load / store operation is designated as the operand of the instruction itself or the address of the corresponding translation table entry identified by the address translation circuit (such as for translating the address designated as the operand of the instruction). Alternatively, for some embodiments, the translation table entry load / store circuit can be implemented within the address translation circuit (rather than the load / store unit), because the address translation circuit may already have a mechanism for controlling page table walks or obtaining translation table entries from a cache structure, and thus it may be more efficient to implement the translation table entry load / store operation by extending the scope of operations available to the address translation circuit (e.g., adding compare-and-swap or bit clear / set functionality) once the address of the corresponding translation table entry has been identified based on the selected address.
[0060] In other examples, the translation table entry load / store trigger instruction includes a store instruction that specifies a predetermined translation table entry load / store trigger address as the store address operand, for which the store trigger translation table entry load / store circuit to that address performs a translation table entry load / store operation. Thus, with this method, there is no need to provide encoding space within the instruction set architecture for a dedicated type of instruction for triggering translation table entry load / store operations. Instead, existing general-purpose store instructions can be used to trigger translation table entry load / store operations, where the store address operand of the store instruction differentiates whether the store should be treated as a normal store operation (triggering a store to the address specified by the store address operand) or a translation table entry load / store trigger instruction (where the store address operand is an address that is mapped to a special address for triggering translation table entry load / store operations). With this method, for example, a predetermined translation table entry load / store trigger address can be the address of a memory-mapped register that, when written to, causes the translation table entry load / store circuit to perform a translation table entry load / store operation.
[0061] With this method, the translation table entry load / store circuit can obtain software-defined address information from a memory-based data structure accessed from a software-programmable base address. When a general-purpose store instruction is reused to trigger a translation table entry load / store operation, its address operand has been used to specify a predetermined translation table entry load / store trigger address, so the selected address used to identify which translation table entry to load / store is alternatively obtained from the memory-based data structure.
[0062] When the software-defined address information specified by the memory-based data structure specifies multiple selected addresses, the translation table entry load / store circuit can perform a translation table entry load / store operation for each of the selected addresses in response to executing a single instance of a store instruction that acts as a translation table entry load / store trigger instruction. Thus, one advantage of using a store to a memory-mapped location to trigger a translation table entry load / store operation is that a single instruction executed by software can trigger multiple instances of translation table entry load / store operations for each selected address specified by the memory-based data structure.
[0063] With this method, the translation table entry load / store circuit can operate asynchronously and does not need to use separate instances of instructions to explicitly mark each selected address. The selected addresses to be processed may have been previously stored in the memory-based data structure before triggering the translation table entry load / store operation. With the asynchronous method, the set of loads / stores required for each translation table entry corresponding to a defined set of selected addresses may be performed by the processing circuit in the context of the ongoing execution of other instructions, which may be helpful for performance.
[0064] The translation table entry load / store circuit can update a software-accessible location to specify a progress indicator that indicates the progress made when performing translation table entry load / store operations for multiple selected addresses. This can be helpful because if any error occurs while processing one of the selected addresses, the processing may stop, and this can help the software to identify how many of the selected addresses have been successfully processed, such that the address that caused the error can be identified.
[0065] In the case where the translation table entry load / store circuit is triggered to perform translation table entry load / store operations for multiple selected addresses in response to a store instruction processed by the processing circuit in a higher privilege execution state, and the processing circuit then switches to a lower privilege execution state, the translation table entry load / store circuit is able to continue processing the remaining addresses among the multiple selected addresses after switching to the lower privilege execution state. Again, this reflects an asynchronous approach where applying translation table entry load / store operations to a set of selected addresses can continue in the background while processing continues at the processing circuit based on the execution of other instructions. It may be particularly useful to allow this processing to continue despite the reduction in privilege at the processing circuit because the responsibility for updating translation table entries often lies with software executing in a higher privilege execution state, but when those translation table entries are being updated, the higher privilege software may not require other processing, and thus it may be useful to allow a lower privilege process to make some forward progress while executing instructions on the processing circuit simultaneously.
[0066] As mentioned above, a memory-based data structure can be used to define one or more selected addresses for which translation table entry load / store operations are to be applied. The selected addresses defined in such a data structure can be non-consecutive addresses that need not be adjacent within a given range.
[0067] Another option that can be used in a method of providing special instructions in an instruction set architecture, or in an asynchronous accelerator method of applying an operation to each address specified in a defined memory structure, can be to specify range information that identifies a contiguous range of addresses for which translation table entry load / store operations are to be applied as software-defined address information. Thus, for at least one variant of the translation table entry load / store operation, the software-defined address information specifies a range of addresses, and the translation table entry load / store circuit is configured to perform translation table entry load / store operations for each address within that range that is a selected address.
[0068] The address translation circuit can support two-stage address translation between the virtual address space and the physical address space based on a first translation table structure that provides address mapping information for translation between the virtual address space and an intermediate address space, and a second translation table structure that provides address mapping information for translation between the intermediate address space and the physical address space. The techniques discussed above can be applied to the loading / storing of entries in the first translation table structure or the second translation table structure. Thus, some variations of the translation table entry loading / storing operation can be defined for specific stages of address translation.
[0069] For example, the translation table entry loading / storing circuit is configured to support at least one of the following:
[0070] ● A first-stage variation of the translation table entry loading / storing operation, where the selected address includes a virtual address specified in the virtual address space, and at least one target translation table entry includes at least one translation table entry of the first translation table structure; and
[0071] ● A second-stage variation of the translation table entry loading / storing operation, where the selected address includes an intermediate address specified in the intermediate address space, and at least one target translation table entry includes at least one translation table entry of the second translation table structure.
[0072] Thus, the first-stage variation can be applicable to operating system software to manage updates to the stage 1 translation table, while the second-stage variation can be applicable to hypervisor software to manage the stage 2 translation table.
[0073] Multiple variations of the translation table entry loading / storing operation have been discussed above. It should be understood that a given implementation can support any one or more of these variations, provided that there is at least one variation that supports the clearing of access tracking metadata as discussed above.
[0074] The techniques discussed above can be implemented within a data processing apparatus that has hardware circuitry provided for implementing the processing circuitry, address translation circuitry, and translation table entry loading / storing circuitry discussed above.
[0075] However, the same technique can also be implemented within a computer program that executes on a host data processing device to provide an instruction execution environment for the execution of object code. Such a computer program can control the host data processing device to simulate the architectural environment that would be provided on a hardware device that actually supports object code according to a certain instruction set architecture, even if the host data processing device itself does not support that architecture. The computer program can have address translation program logic and translation table entry load / store program logic that mimics the functionality of the address translation circuit and the translation table entry load / store circuit discussed above, including support for translation table entry load / store operations. For example, such a simulation program may be useful when traditional code written for one instruction set architecture is executed on a host processor that supports a different instruction set architecture. Additionally, since the execution of software on the simulated execution environment can enable software testing to proceed in parallel with the ongoing development of hardware devices that support a new architecture, the simulation can allow software development for a newer version of an instruction set architecture to begin before the processing hardware that supports that new architecture version is ready. The simulation program can be stored on a storage medium, which can be a non-transitory storage medium.
[0076] Specific example of a data processing device
[0077] Figure 1Schematically illustrates an example of a data processing apparatus 2. The data processing apparatus has a processing pipeline 4 (an example of a processing circuit) including a plurality of pipeline stages. In this example, the pipeline stages include: a fetch stage 6 for fetching instructions from an instruction cache 8; a decode stage 10 for decoding the fetched program instructions to generate micro-operations (decoded instructions) to be processed by the remaining stages of the pipeline; a dispatch stage 12 for checking whether the operands required for the micro-operations are available in a register file 14 and dispatching the micro-operations for execution once the required operands for a given micro-operation are available; an execution stage 16 for performing a data processing operation corresponding to the micro-operation by processing the operands read from the register file 14 to generate a result value; and a write-back stage 18 for writing the processed result back to the register file 14. It should be understood that this is merely an example of a possible pipeline architecture, and other systems may have additional stages or different stage configurations. For example, in an out-of-order processor, a register renaming stage may be included for mapping architectural registers specified by program instructions or micro-operations to physical register specifiers that identify physical registers in the register file 14. In some examples, there may be a one-to-one relationship between the program instructions decoded by the decode stage 10 and the corresponding micro-operations processed by the execution stage. There may also be a one-to-many or many-to-one relationship between the program instructions and the micro-operations, such that, for example, a single program instruction may be split into two or more micro-operations, or two or more program instructions may be fused to be processed as a single micro-operation.
[0078] The execution stage 16 includes a plurality of processing units for performing different categories of processing operations. For example, the execution units may include: a scalar arithmetic / logic unit (ALU) 20 for performing arithmetic or logical operations on scalar operands read from the register 14; a floating-point unit 22 for performing operations on floating-point values; a branch unit 24 for evaluating the result of a branch operation and accordingly adjusting a program counter representing the current execution point; and a load / store unit 26 for performing load / store operations to access data in a memory system 8, 30, 32, 34. A memory management unit (MMU) 28 (which is an example of an address translation circuit) is provided for performing an address translation between a virtual address specified by an operand of a data access instruction by the load / store unit 26 and a physical address identifying a storage location of data in the memory system. The MMU has a translation lookaside buffer (TLB) 29 for caching translation data from a page table stored in the memory system, where the page table entries of the page table define an address translation mapping and access permissions, and the access permissions, for example, control whether a given process executing on the pipeline is allowed to read, write, or execute instructions from a given memory region.
[0079] In this example, the memory system includes a level-1 data cache 30, a level-1 instruction cache 8, a shared level-2 cache 32, and a main system memory 34. It should be understood that this is merely an example of a possible memory hierarchy, and other arrangements of caches may be provided. The particular types of processing units 20 to 26 shown in the execution stage 16 are merely an example, and other embodiments may have different sets of processing units or may include multiple instances of the same type of processing unit, such that multiple micro-operations of the same type can be processed in parallel. It should be understood that Figure 1 is a simplified representation of some components of a possible processor pipeline implementation, and the processor may include many other elements not illustrated for the sake of brevity. Although Figure 1 a single processor core with access rights to the memory 34 is shown, the apparatus 2 may also have one or more additional processor cores that share access rights to the memory 34, where each core has a corresponding cache 8, 30, 32.
[0080] Figure 2 is a diagram illustrating different execution states (also referred to as exception levels) in which the processing circuit 4 can operate when executing instructions. In this example, there are four exception levels EL0, EL1, EL2, EL3, where the exception level EL0 is the lowest privilege exception level and the exception level EL3 is the highest privilege exception level. Generally, when executing at a higher privilege exception level, the processing circuit may have access rights to some memory locations or registers 14 that are inaccessible to lower, less privileged exception levels.
[0081] In this example, the exception level EL0 is used to execute applications managed by the corresponding operating system or virtual machine executing at the exception level EL1. In the case where multiple virtual machines coexist on the same physical platform, a hypervisor operating at EL2 may be provided to manage the corresponding virtual machines. Although Figure 2 an example of a hypervisor managing virtual machines and virtual machines managing applications is shown, it is also possible for the hypervisor to directly manage applications at EL0.
[0082] Although not necessary, some embodiments may implement secure and non-secure domains as separate hardware partitions for the processing circuit. The data processing system 2 may have hardware features implemented within the processor and memory system to ensure isolation of access to data and code associated with software processes operating in the secure domain from processes operating in the non-secure domain. For example, something such as that provided by Cambridge, UK Limited Hardware architectures such as architectures. Alternatively, other hardware-enforced security partition architectures may be used. Secure applications (trusted services) may operate at exception level EL0 in the secure domain, and a secure (trusted) operating system or virtual machine may operate at exception level EL1 in the secure domain. In some specific implementations, EL2 in the secure state is not supported and the hypervisor can only execute in non-secure EL2. In other specific implementations, a secure hypervisor that executes at secure EL2 may be supported, as indicated by the asterisk in Figure 2 . In some examples, a security monitor that executes at exception level EL3 may be provided to manage the transition between the non-secure domain and the secure domain. Other specific implementations may monitor the transition between secure domains in hardware, such that a security monitor may not be required.
[0083] Address translation
[0084] One task performed by the MMU 28 is the address translation between virtual addresses (VA) and physical addresses (PA). Software executed on the processing circuitry 4 specifies memory locations using virtual addresses, but these virtual addresses may be translated by the MMU 28 into physical addresses that identify the memory system locations to be accessed. The benefit of using virtual addresses is that it allows the management software (such as an operating system (OS)) to control the view of the memory presented to the software. The OS can control which memory is visible, the virtual addresses at which the memory is visible, and what accesses are permitted to the memory. This allows the OS to sandbox applications (hide the resources of one application from another application) and provide an abstraction from the underlying hardware. Another benefit of using virtual addresses is that the OS can present multiple segmented physical regions of memory as a single contiguous virtual address space to an application. Virtual addresses are also beneficial to software developers who do not know the exact memory addresses of the system when writing their applications. With virtual addresses, software developers do not need to worry about physical memory. The application knows that the OS and the hardware should work together to perform the address translation.
[0085] In fact, each application may use its own set of virtual addresses, which will be mapped to different locations in the physical system. When the operating system switches between different applications, it reprograms the mapping. This means that the virtual addresses for the current application will be mapped to the correct physical locations in memory.
[0086] Virtual addresses are translated into physical addresses through mapping. The mapping between virtual addresses and physical addresses is stored in a translation table (sometimes called a page table). The translation table is stored in memory and is managed by software (usually the OS or the hypervisor). The translation table is not static, and the table can be updated as the needs of the software change. This changes the mapping between virtual addresses and physical addresses.
[0087] For memory accesses performed when the processing circuitry 4 is in a subset of the execution states, specifically when the processing circuitry 4 is in non-secure EL0 or non-secure EL1, two-stage address translation is used, as Figure 3 shown (for other execution states, one stage of address translation using the stage 1 page table is sufficient). Thus, virtual addresses from non-secure EL0 and non-secure EL1 are translated using two sets of tables. These tables support virtualization and allow the hypervisor to virtualize the view of the physical memory seen by a given virtual machine (VM) (the virtual machine corresponding to the guest operating system and the applications controlled by that guest operating system). This set of translations controlled by the OS is referred to as stage 1. The stage 1 table translates the virtual address to an intermediate physical address (IPA - an example of the intermediate address mentioned previously). In stage 1, the OS considers the IPA to be the physical address space. However, the hypervisor controls a second set of translations, which is referred to as stage 2. This second set of translations translates the IPA to a physical address.
[0088] The stage 1 translation table and the stage 2 translation table are implemented as a hierarchical table structure that includes multiple levels of translation tables as shown respectively for stage 1 and stage 2 Figure 4 and Figure 5 shown. In this example, both the stage 1 table and the stage 2 table can have up to 4 levels of page tables, namely level 0 (L0), level 1 (L1), level 2 (L2), and level 3 (L3).
[0089] To locate the physical address mapping for a given address, a translation table walk including one or more translation table lookups is performed. A translation table walk is the set of lookups required to translate a virtual address to a physical address. For the non-secure EL1 and EL0 translation schemes, this set includes lookups for both stage 1 translation and stage 2 translation. The information returned by a successful translation table walk using the stage 1 lookup and the stage 2 lookup includes:
[0090] ● The physical address required (translated based on the stage 1 mapping to the intermediate address and the stage 2 mapping to the physical address).
[0091] ● The access rights and / or memory attributes of the target memory region, which provide information on how to control access to that memory region. This information can include the stage 1 access rights and / or attributes defined in the stage 1 table structure and the stage 2 access rights and / or attributes defined in the stage 2 table structure.
[0092] To traverse either the stage 1 structure or the given one of the stage 2 structures, based on the address specified in the translation table base address register (TTBR for stage 1, VTTBR_EL2 for stage 2), the walk starts with a read at the highest level (L0) for the initial lookup. Each translation table lookup returns a descriptor that indicates one of the following:
[0093] ● The entry is the last entry of the traversal of the stage 1 structure or the stage 2 structure, and this last entry provides the address mapping being sought. If the entry is in the final L3, then the entry is called a page descriptor (D_Page), while if the entry that provides the last entry of the walk is at one of the higher levels, then the entry is called a block descriptor (D_Block). The page descriptor and the block descriptor may be collectively referred to as "leaf" translation table entries, where "leaf" refers to the last entry of the walk that provides the address mapping. The last entry of the traversal contains the output address (OA — i.e., IPA for stage 1 or PA for stage 2)
[0094] and the permissions and attributes for access. If a block descriptor is found at a higher level in the translation table structure, this means that the block descriptor represents a memory region larger than the 4 kB memory page represented by a single entry at L3 (the specific sizes represented by the block descriptors at L1 and L2 depend on the number of bits used for indexing into the L1 table or the L2 table; in this example, the L1 and L2 block descriptors represent 1 GB regions and 2 MB regions, respectively).
[0095] ● An additional level to be looked up. In this case, the entry is called a table descriptor (D_Table) or a "branch" translation table entry because the entry provides the translation table base address for the lookup at the additional level in the table. The table descriptor may also optionally provide other hierarchical attributes that can be applied to the final translation. The encoding of the translation table entries at level 1 and level 2 differentiates between the block descriptor and the table descriptor.
[0096] ● The descriptor is invalid. In this case, a translation fault occurs on the memory access.
[0097] Figure 4Illustrates using the corresponding bits of the virtual address provided as the input address for table lookup to index the stage 1 translation table. The base address of the highest level L0 table is read from TTBR, and the base addresses of the L1, L2, and L3 tables are indicated by the addresses stored in the index table descriptors in the L0, L1, and L2 tables respectively (if the block descriptor is not identified in the L1 or L2 table, when the block descriptor is found in the index entry of L1 or L2, traversal stops at this level because the output address mapping has been found). The specific entry to be selected within a given level of the stage 1 translation table is determined based on the index values a, b, c, d, which correspond to a certain subset of the bits of the virtual address provided as the input address for lookup. Figure 4 Illustrates which bits of the input address are used for each index value a, b, c, d in a specific example. The address of the relevant entry in a given table is obtained by adding a multiple of the index bits a, b, c, or d to the base address of that given table, which is determined based on the address specified in TTBR or the table descriptor of the previous level (the multiplier is applied to the index value corresponding to the size of a translation table entry).
[0098] Similarly, Figure 5 Illustrates using the corresponding bits of the intermediate address provided as the input address for stage 2 table lookup to index the stage 2 translation table. This index building is similar to that for stage 1 Figure 4 as shown, but uses a different base address register VTTBR_EL2 to provide the base address of the L0 table. As shown in the example of Figure 5 , for stage 2 lookup, based on the value stored in the control register VTCR_EL2.SL0 which can specify whether the lookup should start at L0 or L1, the starting level at the beginning of the stage 2 translation table walk may be changed. If the stage 2 lookup starts at L0, the index building for levels 0, 1, 2, 3 uses the index values a, b1, c, d which are respectively similar to the index values in Figure 4 for stage 1. If the stage 2 lookup starts at L1, the index building is performed in a similar manner, but now a larger number of index bits b2 are used at the highest level (L1) of the lookup, as shown in Figure 5 . Providing a variable starting level is not an essential feature and can be omitted if desired. Although not shown in Figure 4 , it will also be possible to provide a variable starting level for the lookup at stage 1.
[0099] In fact, when performing a full translation table walk that includes both stage 1 and stage 2 translations, each stage 1 table base address obtained from the TTBR and the table descriptors accessed in the stage 1 L0, L1, L2 translation tables will be the intermediate address required for its own translation using the stage 2 translation table. Therefore, in the case where the translation table walk does not encounter any block descriptors but proceeds all the way to L3 where a page descriptor is found, the full page table walk process can include accessing page tables at multiple levels in the following sequence:
[0100] ● Stage 2 translation of the base address of the stage 1 L0 page table to a physical address (the stage 1 L0 base address is typically an intermediate physical address since the stage 1 translation is configured by the operating system). The stage 2 translation includes 4 lookups (stage 2 L0; stage 2 L1; stage 2 L2; stage 2
[0101] L3).
[0102] ● Stage 1 L0 lookup of the entry at the address obtained based on the L0 index part "a" of the target virtual address and the translated stage 1 L0 base address to obtain the stage 1 L1 base address (intermediate physical address)
[0103] ● Stage 2 translation of the stage 1 L1 base address to a physical address (also including 4 lookups).
[0104] ● Stage 1 L1 lookup of the entry at the address obtained based on the L1 index part "b" of the target virtual address and the translated stage 1 L1 base address to obtain the stage 1 L2 base address (intermediate physical address)
[0105] ● Stage 2 translation of the stage 1 L2 base address to a physical address (also including 4 lookups)
[0106] ● Stage 1 L2 lookup of the entry at the address obtained based on the L2 index part "c" of the target virtual address and the translated stage 1 L2 base address to obtain the stage 1 L3 base address (intermediate physical address)
[0107] ● Stage 2 translation of the stage 1 L3 base address to a physical address (also including 4 lookups)
[0108] ● Stage 1 L3 lookup of the entry at the address obtained based on the L3 index part "d" of the target virtual address and the translated stage 1 L3 base address to identify the target intermediate physical address corresponding to the target virtual address.
[0109] ● Stage 2 translation of the stage 1 L3 base address to a physical address (also including 4 lookups).
[0110] ● Stage 1 L3 lookup of the entry at the address obtained based on the L3 index part "d" of the target virtual address and the translated stage 1 L3 base address to identify the target intermediate physical address corresponding to the target virtual address.
[0111] ● Stage 1 L3 lookup of the entry at the address obtained based on the L3 index part "d" of the target virtual address and the translated stage 1 L3 base address to identify the target intermediate physical address corresponding to the target virtual address.
[0112] ● Stage 1 L3 lookup of the entry at the address obtained based on the L3 index part "d" of the target virtual address and the translated stage 1 L3 base address to identify the target intermediate physical address corresponding to the target virtual address.
[0113] ● Stage 2 translation from the target intermediate physical address to the target physical address, where the latter represents the location in the memory to be accessed corresponding to the original target virtual address (also including 4 lookups).
[0114] Thus, in the absence of any cache and assuming that the starting level of Stage 2 is L0, this translation will include a total of 24 lookups. If the starting level of Stage 2 is L1, this can reduce the number of lookups to 19 (one less lookup for each of the 5 Stage 2 translations performed). However, as can be seen from the above sequence, performing the entire page table walk process can be very slow because it may require a large number of accesses to the memory to traverse each level of the page tables in each stage of the address translation. This is why it is generally desirable to cache the information derived from the translation table walk in the TLB 29 of the MMU 28. The cached information can include not only the final Stage 1 address mapping from VA to IPA, the final Stage 2 mapping from IPA to PA, or the combined Stage 1 and Stage 2 mapping from VA to PA (derived from the previous lookups in the Stage 1 structure and Stage 2 structure), but also the entries of the higher-level page tables from the Stage 1 table and Stage 2 table can be cached within the TLB 29 of the MMU 28. This can allow bypassing at least some steps of the full page table walk, even if the final-level address mapping for a given target address is not currently in the address translation cache.
[0115] Therefore, a TLB 29 that supports a "walk cache" (cache of pointers from branch translation table entries) can provide a faster route to identify the address in the memory that stores the location of the branch translation table entry or leaf translation table entry corresponding to a particular address. Even if it does not support a walk cache, the MMU 28 can support a translation table walk circuit in hardware that can generate a sequence of memory accesses required to traverse the page table structure based on the input address to be translated (which is faster than when software has to explicitly execute a series of load / store / arithmetic instructions based on the input address to be translated to calculate the index value as possible for the translation table), use this index to generate the address for reading the next-level translation table entry, load the address of the next-level translation table entry, and then repeat any additional levels based on the pointer loaded for the previous level until a leaf entry providing the address mapping is found.
[0116] Figure 6 An example of a branch translation table entry 50 and a leaf translation table entry 52 is illustrated. Both types of translation table entries have an encoding that identifies that this is a valid translation table entry. For example, in Figure 6In the example, a valid translation table entry needs to have the least significant bit set to 1. This encoding helps prevent an instruction address pointer (which is typically aligned to a multiple of the instruction size and thus expected to have some lower bits equal to 0) from being accidentally treated as a valid translation table entry. It should be understood that there may also be other ways to identify valid translation table entries. During translation table traversal, if the loaded value at a given level of the traversal is not a validly encoded translation table entry, a fault may be signaled.
[0117] In this example, the second least significant bit is used to distinguish between a branch translation table entry 50 and a leaf translation table entry 52 for a given level detection in translation table traversal beyond the maximum level supported in the traversal (in the above example, the maximum level is level 3). Thus, for the above example, the second least significant bit is used for levels 0, 1, and 2 to distinguish between a table descriptor and a block descriptor. If the second least significant bit is 1, the entry is a table descriptor, and if the second least significant bit is zero, the entry is a block descriptor. A page descriptor (leaf entry at the maximum supported level 3) is encoded as 1 using the second least significant bit. Again, it should be understood that this is just an example encoding, and other encodings can be used to distinguish different types of descriptors.
[0118] The branch translation table entry (table descriptor) 50 specifies the next level table address 54, which provides a pointer to the table at the next level of the translation table structure. The next level table address 54 serves as a base address relative to which the addresses of the individual entries in the next level table can be calculated based on an offset that is derived from an index value selected from the bits of the address being translated.
[0119] In contrast, the leaf translation table entry 52 specifies address mapping information 56, which provides the address mapping for mapping the address to be translated from the input address space to the output address space. For a stage 1 table, the input address space is the virtual address space and the output address space is the intermediate address space, while for a stage 2 table, the input address space is the intermediate address space and the output address space is the physical address space.
[0120] The leaf translation table entry 52 may also specify multiple fields 58 for encoding access rights and / or memory attributes. The access rights may specify what types of memory accesses are permitted to the corresponding region of the address space. For example, the access rights may specify whether reading, writing, and / or using the region for instruction fetch of executable instructions is permitted. The memory attributes may specify other characteristics of the memory region, which may control how memory accesses are performed when access is permitted based on these rights. For example, these attributes may specify characteristics such as whether caching of data from the corresponding memory region is permitted, whether the region is defined as device memory such that reordering or merging of different memory accesses to the device memory is not permitted, and so on. The access rights and memory attributes may be explicitly encoded within the fields 58 of the translation table entry, but may also be defined using an indirect reference to a register. For example, the fields 58 of the leaf translation table entry 52 may specify a register number that identifies a given register and / or a register field identifier that identifies a specific field within the register, and the access rights and / or memory attributes may be encoded by values stored in the register or register field referenced by the translation table entry. Using indirect permission / attribute specification with registers can be used to allow software to quickly update the permissions of many translation table entries that all reference the same permission / attribute field by a single update to a register, rather than requiring updates to many different translation table entries in memory. Additionally, in embodiments where each field of the permission indirect register has more bits than the corresponding permission field of the translation table entry, this indirectness can help support more types of permissions / attributes than the limited encoding space for permissions within the entry would likely support. It should be understood that some embodiments may define the access rights and memory attributes through a combination of explicit encoding information within the translation table entry format itself and indirect reference information within a register. Although Figure 6 the branch translation table entries in Figure 6 are not shown as specifying any access rights or memory attributes, in other examples, the branch translation table entries may also specify some rights or attributes that apply to the corresponding block of memory covered by the branch translation table entry.
[0121] Leaf translation table entries (and in some examples, branch translation table entries as well) may also specify one or more access trace metadata, which can be used to provide information about which regions have been accessed by load / store operations. In this example, there are two pieces of access trace metadata: an access flag (AF) 70 and a dirty bit modifier (DBM) 72. It should be understood that it is not necessary to provide these two types of access trace metadata, and other types of access trace metadata (e.g., a counter that counts the access frequency to the corresponding memory region) can be used. These two forms of access trace metadata can have a first state and a second state. The first state indicates that at least one load / store access has occurred to the corresponding region of the input address space, and the second state indicates that no load / store access has occurred to the corresponding region of the input address space. The access flag 70 is used to track read accesses to the region, and the DBM 72 can be used to track write accesses.
[0122] Periodically, the operating system software can set the access flag 70 in the entries corresponding to a set of memory regions to be monitored to the second state (indicating zero access). When a read access is made to one of these memory regions, the access flag 70 in the corresponding stage 1 block or page descriptor can be set to the first state (if it has not been set after a previous access). In some examples, the store memory access for setting the access flag 70 in the leaf translation table entry 52 corresponding to the target address specified by another load access can be automatically triggered in hardware by the MMU 28 when processing the load to the target address, rather than requiring an explicit software instruction to write to the leaf translation entry 52.
[0123] After one cycle of monitoring, the operating system can check the access flag 70 of the entries it is monitoring to assist in performing operations that may benefit from information about the access frequency of specific pages. For example, the operating system can maintain an additional tracking data structure in memory, which has an entry per memory region that tracks the number of times the memory region has been accessed. And thus, at the end of each cycle of monitoring, the entry in this additional tracking structure corresponding to the memory region with the access flag 70 set can be incremented. After multiple cycles of monitoring, this additional tracking structure will provide an indication of the relative frequency of access to the corresponding memory regions. This can provide useful information for controlling operations such as paging, where the information can be used to identify the pages with the lowest access frequency in memory, and compared with other pages with higher access frequencies, the corresponding data of these pages with the lowest access frequencies can be preferentially paged to an external storage device. The specific use formed by the access flag 70 can vary based on software, but generally, providing at least one bit for tracking whether any read access has occurred to the corresponding region of memory may be useful for software.
[0124] Similarly, DBM 72 assists in tracking which pages have been written to. If the operating system wishes to track whether a given page has been written to, when the page is mapped or at the start of a monitoring period, the operating system can set the access rights for the page to "read-only" (even if the page is expected to be allowed to be written to) and set DBM bit 72 to a second state (indicating no previous write access to the page). For an access rights fault caused by writing to a read-only page when DBM bit 72 is set, the operating system can determine, based on the set DBM bit 72, that this is not a "true" violation of the read-only rights, and instead cause the operating system to update a data structure stored in memory that tracks the pages affected by write requests, update DBM bit 72 to a first state (indicating that at least one previous write access has occurred), and update the write access rights for the page to indicate that the page can now be written to without triggering a fault. After a cycle of monitoring, the tracking data structure in memory can be used by software to determine whether, when paging a particular region, modified data from that region must be written back to an external storage device, or (in the case where no writes have occurred) whether the data stored in memory can be discarded when paging that region, because if the corresponding data in external memory is clean, it can be assumed that the data is still the same.
[0125] In other specific implementations, the permission information 58 of a translation table entry can serve as access tracking metadata for tracking whether a given page has been written to. For example, all pages can initially be set to "read-only" as mentioned above, and a DBM indicator 72 can be set for pages that should actually be readable / writable but are only temporarily "read-only" because they have not been written to yet. The DBM indicator 72 (when set) can indicate to the hardware of the MMU 28 whether it is allowed to update the permission information 58 of a translation table entry to add write permissions when a write to a read-only region is detected. Thus, in a similar manner to the example above, a first write to a previously unwritten page can be detected from the fact that the page is read-only when DBM bit 72 is set to a state indicating that the hardware is allowed to update to add write permissions; and subsequently, when write permissions are added, the permission field can be used to determine that the page has been written to. Thus, in this example, the first state of the access tracking metadata (indicating at least one previous write access) can be when the permission information 58 indicates read / write permissions, and the second state (indicating no previous write access) can be when the permission information 58 indicates read-only when DBM bit 72 is set.
[0126] Thus, for both the access flag 70 and the DBM bit 72 (and / or the permission field 58, provided that the permission field is used to indicate whether a page has been written to previously), the responsibility for clearing the access trace metadata back to a second state (indicating that no previous load / store access to the corresponding region of the address space has occurred) lies with software rather than hardware. If the type of translation table entry load / store operation discussed below is not supported, this will require a general-purpose store instruction that specifies the address of the memory location where the translation table entry to be updated will be stored as its destination address.
[0127] General load / store instruction
[0128] Figure 7 Examples of general-purpose load / store instructions are illustrated for comparison with Figure 8 comparison.
[0129] A load instruction is an instruction that triggers a read of a location in the memory systems 30, 32, 34 and returns the read data to a register. A general-purpose load instruction specifies the destination register Xn into which the read data will be loaded and a target address operand (in this example, specified using the value stored in register Xm), which identifies the memory location to be read, for example using a virtual address in the virtual address space.
[0130] A store instruction is an instruction that causes stored data obtained from a register to be stored to a location in the memory system. A general-purpose store instruction specifies the source register Xn from which the stored data is obtained and a target address operand (in this example, specified using the value stored in register Xm), which identifies the memory location to which the stored data will be written.
[0131] Thus, for a general-purpose load / store instruction, the load / store operation is performed on the memory location identified by the address specified by the address operand of the instruction.
[0132] Typically, such general-purpose load / store instructions are needed when the update is to a translation table. A general-purpose load instruction can read the current contents of a translation table entry, and a general-purpose store instruction can update the contents of a translation table entry. However, since the load / store operation will be performed at the location corresponding to the address specified in the address operand of the load / store instruction, this means that software will first need to identify the address of the location where the relevant translation table entry is stored. This can be complex, as described above Figure 4 and Figure 5As shown, this may require the software to traverse through the translation table structure based on a specific address of interest, and the corresponding translation table entry (which provides the address mapping for the specific address or provides a pointer on the path of traversing through the translation table to find the entry that provides the mapping for the specific address) will be updated for the specific address of interest. This traversal will require the software to execute multiple instructions (e.g., several related load instructions, and arithmetic instructions for calculating the addresses of the entries to be read at various levels of traversal). This can be slow. Since it is often necessary to update a large number of translation table entries (e.g., when clearing access trace metadata at the end of an accounting cycle for monitoring the access frequency to memory, or when migrating a virtual machine from one host processor to another host processor), this may incur significant performance costs.
[0133] Translation table load / store trigger instruction
[0134] In contrast, Figure 8 illustrates an example of a translation table entry load / store trigger instruction that controls a translation table entry load / store circuit to perform a translation table entry load / store operation for at least one target translation table entry address, where the at least one target translation table entry address is selected according to software-defined address information that identifies a selected address in an input address space that is translated to an output address space by a given translation table structure. Different from Figure 7 the general load / store instructions shown, in the translation table entry load / store operation triggered by the translation table entry load / store trigger instruction processed by the processing circuit 4, the load / store operation is performed for at least one target translation table entry address that is selected based on the selected address defined in the software-defined address information, rather than performing a load / store on the selected address itself. The at least one target translation table entry address is selected by the address translation circuit 28 as one or more addresses of one or more target translation table entries corresponding to the selected address. Each target translation table entry is either a leaf translation table entry that provides the actual address mapping information for translating the selected address from the input address space to the output address space or a branch translation table entry traversed in the translation table traversal operation for obtaining the leaf translation table entry corresponding to the selected address.
[0135] Thus, the instruction that triggers the translation table entry load / store operation itself does not need to specify the address of the location where the entry is stored, because the address of the translation table entry to be loaded or stored can alternatively be derived in hardware by the address translation circuit 28 based on an existing mechanism for caching translation table entries in the TLB 29 (e.g., a page table walk cache that caches pointers to entries of the translation table structure) or based on a page table walk circuit implemented using hardware circuit logic that can trigger a series of memory accesses based on the input address to walk through the translation table. Thus, by supporting this operation, software can manipulate translation table entries without performing page table walks in software. The page table walk is performed by hardware and allows utilization of the walk cache in the MMU 28. The benefits are reduced software complexity and a reduction in the total runtime spent in the operation, thus improving performance. It also allows omission of locking structures, further saving time, because software agents that manipulate translation table entries do not need to acquire coarse-grained locks while performing the manipulation.
[0136] This also means that software that manipulates the page table does not need to generate a separate set of page table entries that point to the page table entries of interest, because software that manipulates the page table does not need to issue explicit load / store to the addresses of the page table entries. This is useful because by not explicitly mapping page table entries in the address space visible to the software that manipulates the page table entries, it avoids the situation where the software may corrupt the page table entry when another load / store operation that is not expected to update the page table incorrectly sets its address operand (either due to an accident caused by an error or maliciously based on an attacker exploiting a memory error that causes the address operand to be incorrect). If the address operand of another load / store is accidentally set to the address of a translation table entry, the memory access will trigger a fault because there is no valid translation table entry defined for that address in the input address space. However, the hardware may still be able to follow the trail of pointers defined in other translation table entries to walk through the translation table structure, even if there are no page table entries that provide an input address to output address mapping for some of the addresses indicated by the translation table pointers.
[0137] In some examples, the instruction set architecture supported by the processing circuit 4 can support a new set of memory access instructions (different from general-purpose load / store operations) that trigger translation table entry load / store operations. Figure 8 Two examples of such instructions are illustrated: a translation table entry load trigger instruction and a translation table entry store trigger instruction.
[0138] In this example, the translation table entry load trigger instruction specifies the destination register Xa and an address operand, which in this example is provided using the value stored in register Xb. Similarly, in this example, the translation table entry store trigger instruction specifies the source register Xa and an address operand, which in this example is provided using the value stored in register Xb. The address operand may specify the address itself or an offset relative to a reference address, such as the current value of the program counter indicating the execution point reached by the program being executed. For the load / store variants of this instruction, the address operand is an example of software-defined address information and specifies a selected address.
[0139] For both variants, in response to the instruction being decoded through decode stage 10 and processed at execution stage 16, the MMU 28 obtains the address of the leaf translation table entry corresponding to the selected address from its address translation cache 29 (if cached); or if there is no existing cached information to provide the address of the leaf translation table entry, the MMU 28 triggers a page table walk to obtain the address of the leaf translation table entry. For the load variant of this instruction, a load is performed to read the leaf translation table entry from the memory system 30, 32, 34 and write the loaded translation table entry to the destination register Xa. For the store variant of this instruction, the stored data obtained from the source register Xa is written to the location in the memory system 30, 32, 34 corresponding to the address of the leaf translation table entry.
[0140] Figure 8 Two examples of executable translation table entry load / store trigger operations are shown, but multiple other variants may also be provided, as follows.
[0141] First, the variants of this operation may differ in terms of which translation table entries the load or store operation is directed at. In Figure 8 the example, the target translation table entry is a leaf translation table entry that translates the selected address specified using the address operand Xb. However, in other examples, one or more target translation table entries may be selected in different ways, such as:
[0142] ● Selecting the translation table entry obtained for the selected address at a specified level of the page table walk as the target translation table entry (e.g., the level 1 translation table entry obtained in the page table walk for the selected address, whether the level 1 translation table entry is a leaf translation table entry or a branch translation table entry);
[0143] ● Select the leaf translation table entry and each branch translation table entry on the path to the leaf translation table entry for translating the selected address as the target translation table entry. For example, this variant of the translation table entry store instruction can be used to clear the access flags at each level of the translation table structure on the path to the leaf.
[0144] ● Select the final valid translation table entry reached in the translation table walk for the selected address as the target translation table entry, whether the final valid entry is a branch translation table entry or a leaf translation table entry (this may be useful when the table structure is still only partially constructed).
[0145] In addition, there may be variants of the operations for the target stage 1 or stage 2 translation table entries respectively. For the stage 1 variant of the operation, the selected address specified by the address operand Xb is a virtual address, and the load / store operation is performed on at least one stage 1 translation table entry corresponding to that address. For the stage 2 variant of the operation, the selected address specified by the address operand Xb is an intermediate address, and the load / store operation is performed on at least one stage 2 translation table entry corresponding to that address.
[0146] In addition, variants of the operation can be defined that differ in the specific operations applied to each target translation table entry:
[0147] ● For Figure 8 the load variant of the instruction in, the operation is a load that loads the translation table entry into the destination register;
[0148] ● For Figure 8 the store variant of the instruction in, the operation is a store that writes the store data obtained from the source register Xa to the memory location associated with the target translation table entry address.
[0149] Additional variants can be defined as follows:
[0150] CASS1 Xa,Xb,[Xc] — Atomic compare and swap stage 1 translation table entry.
[0151] Xa provides the compare operand;
[0152] Xb provides the swap operand;
[0153] Xc provides the address operand for identifying the selected address.
[0154] In response to the instruction, control the MMU 28 to obtain at least one target translation table entry address of at least one target stage 1 translation table entry corresponding to the selected address (where the target entry is determined according to any of the above examples). For each such target translation table entry address, initiate an atomic compare and swap operation, where the atomic compare and swap includes:
[0155] ● Read data from the memory location corresponding to the target translation table entry address;
[0156] ● Compare the read data with the compare operand;
[0157] ● If the comparison of the read data with the compare operand satisfies the comparison condition, write the swap operand to the memory location corresponding to the target translation table entry address,
[0158] If the comparison condition is not satisfied, the write does not occur;
[0159] ● Return a status indication (e.g., in a control register or by setting condition status flags) that indicates whether the comparison condition is satisfied.
[0160] The read and write operations in the compare and swap are performed atomically, such that the compare and swap operation is seen by other observers of the memory location corresponding to the target translation table entry address as a single indivisible operation. This means that the result of the compare and swap is equivalent to the result that would be obtained if no other writes to the target translation table entry address occurred between the read and the write. Atomicity can be implemented in different ways, such as by locking access to the address to prevent intermediate write operations, or by continuing to allow access to the location corresponding to the address, but providing a mechanism for detecting when an intermediate write has occurred between the read and write of the atomic compare and swap operation, and canceling and re-executing the atomic compare and swap operation in the event of a detected intermediate write. In some specific implementations, the memory systems 30, 32, 34 may have a mechanism to allow the atomic compare and swap operation to be performed locally at a location close to where the data is stored, thereby reducing the latency between the read and write performed. Other examples may not support this and may require data transfer between the processing pipeline 4 and the memory to support the atomic compare and swap, such that the compare part of the operation can be completed at the processing pipeline 4.
[0161] SWAPS1 Xa,Xb,[Xc] — Swap stage 1 translation table entry
[0162] Xa is the destination register, Xb provides the swap data operand, and Xc provides the address operand for identifying the selected address. This instruction causes the MMU to obtain the target translation table entry address of the stage 1 translation table entry based on the selected address specified by the operand Xc. A swap operation is performed to update the translation table entry at the target translation table entry address with the updated value specified by the swap data operand in Xb, and the destination register Xa is loaded with the value before the update of the target translation table entry before the update or the value after the update of the target translation table entry after the update.
[0163] BITSETS1 Xa,[Xb]
[0164] In response to this instruction, the control MMU 28 is caused to obtain at least one target translation table entry address of at least one target stage 1 translation table entry corresponding to the selected address (where the target entry is determined according to any of the examples above). Xa provides a bit position operand or a bit mask that specifies the position within the translation table entry of at least one bit to be set to 1. Xb provides the address operand for identifying the selected address.
[0165] For each such target translation table entry address, an atomic operation is initiated to:
[0166] ● Read data from the memory location corresponding to the target translation table entry address;
[0167] ● Update one or more bits of the target translation table entry at any bit position specified by the bit position operand Xa to set the one or more bits to 1 while keeping all other bits of the translation table entry unchanged;
[0168] ● Write the updated value of the translation table entry back to the memory location corresponding to the target translation table entry address.
[0169] Similarly, this is done atomically with the atomicity as discussed above for the compare and swap variant. A similar bit clear variant of this instruction may be provided, which sets the specified bit to 0 instead of 1.
[0170] For each of the variants described above for stage 1, a similar stage 2 variant may also be provided, where the address operand is alternatively interpreted as an intermediate address (instead of the virtual address for the stage 1 variant), and the target translation table entry is an entry in the stage 2 table (thus, any walk for obtaining the address of the target translation table entry is based on the stage 2 base address in VTTBR_EL2, rather than the stage 1 base address in TTBR for the stage 1 variant of the operation).
[0171] In addition to all of the above, instructions can be provided to support different variants for the final level of describing roaming: In a first variant, the instruction encoding indicates "perform an operation regardless of the result of the final level of roaming". In a second variant, the instruction encoding indicates "the final level of roaming is expected to be level X", and if this does not match the current configuration of the translation table (i.e., if the leaf entry of the selected address is not at a level other than level X), the hardware raises a fault condition or signals another error response (e.g., sets an error code in a register).
[0172] In addition to all of the above, instructions can handle error conditions based on control. Software may use the instructions inappropriately such that leaf-level page table entries are not found. In these cases, the hardware may generate an exception or fill a register to indicate "operation failed". Regardless of the reporting mechanism, an error code indicating the nature of the error is reported.
[0173] Different variants of an operation can be distinguished by any of the following: the instruction opcode, another field in the instruction encoding, a previous instruction that modifies the behavior of the instruction, or control information stored in a control register or other storage location. Additionally, some of the lower bits of the selected address specified by the address operand can be used to encode the variant of the operation, since those lower bits (sub-page or sub-region address bits) are not needed to identify the address of the page corresponding to the translation table entry to be updated.
[0174] At least the store variant, the compare-and-swap variant, and the bit-set / clear variant can be examples of variants that support clearing the access tracking metadata 70, 72, 58 in the translation table entries. However, these instructions can also be used to update other translation table information, such as the address mapping 56, the table pointer 54, and the access rights or memory attributes 58.
[0175] Variants can also be provided that identify multiple selected addresses, and for each of these selected addresses, a translation table entry load / store operation is applied to that address. For example, the instruction can specify information that defines a range of addresses, and each address in that range can be considered a selected address for a corresponding instance of the translation table entry load / store operation to be performed on it. Thus, range-based operations can be performed on more than one translation table entry, and appropriate translations are performed on multiple leaf-level translation table entries and / or higher-level entries that lead to these leaf-level translation table entries.
[0176] In an example where dedicated instructions are supported in an instruction set architecture for triggering translation table entry load / store operations, the translation table entry load / store circuitry can be considered to include a load / store unit 26 (also used for general load / store operations) and an MMU 28 (address translation circuitry).
[0177] Method
[0178] Figure 9 A method of data processing is shown. At step 200, instructions are processed (e.g., decoded and executed) by a processing circuit. At step 202, for any load / store instruction, the address translation circuit 28 performs a translation between an input address space and an output address space based on mapping information 56 obtained from a translation table structure.
[0179] Figure 10 A method of processing a translation table entry load / store operation (which may be performed at Figure 9 step 200) is shown. In response to a translation table entry load / store trigger instruction processed at step 210, at step 212, the address translation circuit 28 selects at least one target translation table entry address according to software-defined address information that identifies a selected address in an input address space to be translated by a given translation table structure. Each target translation table entry address includes the address of a leaf translation table entry that provides address mapping information for translating the selected address or the address of a branch translation table entry traversed in a translation table walk operation for obtaining the leaf translation table entry. For example, the at least one target translation table entry address may be obtained from an address translation cache 29 or by performing a translation table walk for the selected address.
[0180] At step 214, the address translation circuit 28 determines whether an error condition has been identified, such as the non-existence of a valid translation table entry defined for the selected address, or the leaf translation table entry is identified at a different level in the translation table structure than expected. If an error condition is identified, at step 216, an error reporting action is performed, such as signaling a fault or setting an error status indication in a register. This may notify the software that the current configuration of the translation table structure is not as expected.
[0181] If no error condition is identified, at step 218, the translation table entry load / store circuits 26, 300 perform a load / store operation for the at least one target translation table entry address. The load / store may be any translation table entry load / store operation variant among the translation table entry load / store operation variants discussed above. For at least one of these variants, the load / store supports clearing of access trace metadata.
[0182] At step 220, summary information is updated in a software-accessible register. The summary information may provide information about the updated translation table entry, such as specifying the level of the translation table structure in which the translation table entry has been found and / or specifying information about access permissions or memory attributes in the entry (e.g., whether the entry specifies a read-only region of memory).
[0183] Accelerator example
[0184] Figure 11 Illustrates another example of apparatus 2, in which, in this example, the translation table entry load / store circuit 300 is provided as an accelerator separate from the load / store unit 26 for performing conventional load / store operations. The translation table entry load / store accelerator 300 has access to the address translation circuit 28 such that it can trigger a lookup in the TLB or any walk cache 29 and trigger the walk circuit 302 of the address translation circuit 28 to perform a hardware-managed translation table walk for a given address specified by the translation table entry load / store circuit 300. Alternatively, the accelerator 300 can have the address translation circuit 28 provided by a local MMU-like structure that is separate from the main MMU used by the load / store unit 26 for general load / store. The translation table entry load / store circuit 300 can also initiate load / store operations on the memory system 304 (including caches 30, 32 and memory 34), which can be performed asynchronously without specifically executing individual load / store instructions for each address to be loaded / stored.
[0185] Figure 12 Illustrates a programmer interface for accelerator 300, which includes one or more memory mapped registers. Memory mapped registers are registers accessed by the processing circuit 4 by performing a general load / store operation that designates a predetermined address allocated to represent the memory mapped register as the target address of the load / store. The memory mapped registers provide:
[0186] ● A base address 310 for indicating a physical address that indicates the start of a memory-based circular buffer structure 320 that serves as software-defined address information that defines one or more addresses, each of which will be used as the "selected address" for a translation table entry load / store operation. The circular buffer 320 is located in a contiguous region of the physical address space. Software can allocate the addresses of the corresponding translation table entries to be loaded / updated to this circular buffer. The hardware of the accelerator 300 can read this buffer to identify the addresses to be processed using translation table entry load / store operations.
[0187] ● A size parameter 312 such that the accelerator 300 can detect the location of the end of the buffer 320.
[0188] ● The "Run" parameter 314, which indicates to the accelerator 300 whether it should perform a translation table entry load / store operation at the next address in the circular buffer 320. When the software sets this "Run" parameter to a first state (e.g., 1), the accelerator 300 begins working through the buffer to perform a translation table entry load / store operation for each address indicated in the buffer. When the accelerator 300 reaches the end of the buffer or encounters an error, the hardware of the accelerator clears this "Run" parameter to a second state (e.g., 0). When the run parameter is in the second state, the accelerator 300 does not perform a translation table entry load / store operation.
[0189] ● The "Fault" field 316, which is used to record fault information that indicates any errors detected by the accelerator 300 during the performance of a translation table entry load / store operation. The software can clear this "Fault" field to a first state (e.g., 0) before starting the operation, and the hardware 300 can set this fault field to a second state (e.g., 1) when a fault occurs (and also change "Run" to the second state to stop further execution of the translation table entry load / store operation). When in the second state, the fault field can also specify further information about the cause of the fault.
[0190] ● The "Progress" field 318, which provides a "current index" that identifies the most recent entry in the memory-based buffer 320 that has been processed using a translation table entry load / store operation. The index 318 can be used to calculate the memory address to read for the next entry in the buffer relative to the base address 310. The index is incremented each time an address from the buffer 320 is successfully processed using a translation table entry load / store operation. If an error occurs, the index provided by the progress indicator 318 can be used by the software to identify which address caused the error.
[0191] The various pieces of information 310 to 318 shown in the memory-mapped register can be encoded in the register in different ways. For example, in some examples, there can be a separate register for each piece of information, or some information can be combined into the same register. For example, the base address 310 and the size 312 can be encoded in one memory-mapped register (at a given address A), and the run 314, fault 316, and progress 318 indicators can be encoded in a second memory-mapped register (at a given address B).
[0192] Thus, using this method, the instruction that triggers the translation table entry load / store instruction can be a general-purpose store instruction processed by the processing circuit 4, and the general-purpose store instruction is mapped to the address of the memory mapped register including the run indicator 314 as its target address. This avoids the need to consume encoding space in the instruction set architecture for a dedicated instruction for triggering the translation table entry load / store operation. This method may also be helpful because the accelerator 300 enables the translation table entry load / store operation to be applied to multiple addresses in the context of other instructions being processed on the processing circuit 4 in an asynchronous or semi-asynchronous manner, and the processing circuit can continue with other operations. Thus, this improves the throughput of the processing operation. Software can program the relevant addresses in the buffer 320, program the base address / size indicators 310, 312, set the "current index" in the progress indicator 318 to 0, and set RUN = 1 and FAULT = 0. Then, the software can turn to perform other work and later check the RUN bit to see if the process is completed and no faults are identified.
[0193] When RUN = 1, the hardware of the accelerator 300 can then iterate as follows:
[0194] 1. Read the entry from the circular buffer = [base address + (8 * current index)]. This entry is a VA or IPA.
[0195] 2. Obtain the translation table entry address corresponding to the VA or IPA read from this entry. For example, look up the VA or IPA in the snoop cache 29 and / or use the current stage 1 translation table base address or stage 2 translation table base address obtained from the TTBR or VTTBR_EL2 configuration to trigger translation table snoop to find the leaf entry and / or branch entry for this VA or IPA, and perform a load, store, compare and swap, bit set or bit clear operation on this entry. For example, this can be used to clear the access flag 70, dirty bit 72 or other access trace metadata.
[0196] 3. Increment the "current index". If this is the end of the circular buffer (e.g., determined based on whether the current index >= 2^size, where the size parameter 312 is encoded as the log(2) of the number of entries in the buffer), then set RUN = 0 and stop. Otherwise, repeat from step 1 for the next entry in the buffer 320.
[0197] If the snoop encounters a fault, or the leaf descriptor is invalid (which will result in a translation fault), then the hardware of the accelerator 300 configures RUN = 0 and FAULT = 1 and stops. This does not generate a data abort or other exception. The software checks the RUN / FAULT bits and the "current index" to check the progress / error.
[0198] Once set to run, the translation table entry load / store accelerator 300 can continue to operate even if there is a transition to a lower privilege state at the processing circuitry 4. For example, if the accelerator 300 has been configured by software executing in EL2 to apply updates to the respective stage 2 translation table entries, then the accelerator 300 can continue to operate through the buffer 320 to apply translation table entry load / store operations to each address in the buffer even if the processing circuitry 4 performs an exception return to EL1 or EL0. Later, when the processing circuitry 4 transitions back to EL2, the software at EL2 can check the run / fault / progress indicators 314, 316, 318 to check the progress and whether the operation was successful. Thus, the method means that the translation table entry load / store operations can be performed asynchronously in the context of other instructions being executed on the processing circuitry 4.
[0199] Figure 13 is a flowchart illustrating a method of performing translation table entry load / store operations using the accelerator 300. At step 250, the accelerator 300 determines whether a run indicator 314 (set in a memory mapped register) is set to a first state (e.g., RUN = 1). If the run indicator is in a second state (RUN = 0), then the accelerator takes no action.
[0200] When the run indicator 314 is determined to be set to the first state, then at step 252, the accelerator determines the next buffer address to be read. The next buffer address is determined based on a base address 310 and a progress indicator 318 (e.g., adding a multiple of an index provided by the progress indicator 318 to the base address 310). At step 254, the accelerator 300 controls the address translation circuitry 28 to obtain at least one translation table entry address of at least one translation table entry for translating a selected address read from an entry in the buffer 320 corresponding to the next buffer address. Thus, at step 254, a load operation is performed to load the address from the next buffer entry identified by the next buffer address, and then the selected address read from that entry is provided to the address translation circuitry 28, which can use the selected address to trigger a lookup in the snoop cache 29 and / or use the translation table walk of the walk circuitry 302.
[0201] At step 256, accelerator 300 and / or address translation circuit 28 determine whether an error condition is detected. For example, an error may be detected if a walk of the translation table structure does not find a valid translation table entry for the selected address, or if a leaf translation table entry is identified at an error level of the table structure. An error may also occur if the loading of a buffer entry based on the next buffer address violates the access rights for that address. If an error is detected, then at step 258, the fault reporting register 316 is updated to indicate that a fault has occurred and the cause of the fault. Additionally, accelerator 300 clears the run indicator 314 to a second state (e.g., RUN = 0).
[0202] If no error is detected at step 256, then at step 260, accelerator 300 performs a load / store operation on each of one or more addresses returned as the address of at least one translation table entry. The load / store can be any of the variants discussed above, such as a load variant, a store variant, a compare and swap variant, or a bit set / clear variant. For the load variant, when this is performed by accelerator 300, a second buffer (identified by a second base address) may be provided to be written with a translation table entry loaded based on the address identified at step 254. When using one of the store variant / compare and swap variant / bit clear variant / bit swap variant, additional operands for these operations (e.g., the set store data, the compare and swap value, or the bit positions for the bit clear / swap operation) may be obtained from an additional structure in memory or may be encoded in the same buffer as buffer 320 used to provide the selected address. However, for some specific implementations, such additional operands may not be required (e.g., if accelerator 300 is dedicated to accessing flag 70 and / or clearing DBM 72, the operations to be performed may be implied to be atomic operations to read a translation table entry, update flag 70 and DBM 72 to indicate a state of zero previous access, and write back the updated translation table entry, and thus explicit operands may not be required).
[0203] At step 262, the accelerator 300 increments the progress indicator 318 (in the memory mapped register) to reflect that the load / store operation for the current buffer entry has been successful, such that any additional iterations of the load / store of the translation table entries will be performed for the next entry. At step 264, the accelerator 300 detects whether all valid addresses in the buffer 320 have been processed based on the progress indicator 318 and the size indicator 312, and if not, the method returns to step 250 to continue running. The check to see if the run indicator 314 is set to the first state is repeated in each iteration of the loop because software may choose to explicitly stop the progress of the accelerator 300 by clearing the run indicator to the second state, even if not all addresses have been processed. If at step 264, the accelerator 300 detects that all addresses in the buffer have been processed, then at step 266, the run indicator 340 is cleared to the second state (e.g., RUN = 0) by the hardware, and then the accelerator stops its operation to wait for reprogramming for a later instance of the execution of the load / store operation of the translation table entries.
[0204] Emulator
[0205] Figure 14Illustrative simulator implementations that may be used are described. Although the previously described embodiments implement the present invention in terms of apparatus and methods for operating specific processing hardware that supports the technology of interest, it is also possible to provide an instruction execution environment implemented by using a computer program in accordance with the embodiments described herein. Such computer programs are commonly referred to as simulators because they provide software-based implementations of hardware architectures. Types of simulator computer programs include emulators, virtual machines, models, and binary translators (including dynamic binary translators). In general, simulator implementations may run on a host processor 1330 that optionally runs a host operating system 1320 and supports a simulator program 1310. In some arrangements, there may be multiple layers of simulation between the hardware and the provided instruction execution environment and / or multiple different instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations that execute at a reasonable speed, but such approaches may be justified in certain cases, such as when it is necessary to run native code of another processor for compatibility or reuse reasons. For example, a simulator implementation may provide an instruction execution environment with additional functionality not supported by the host processor hardware, or provide an instruction execution environment typically associated with a different hardware architecture. An overview of simulation is given in "Some Efficient Architecture Simulation Techniques" (Robert Bedichek, Winter 1990 USENIX Conference, pages 53 - 63).
[0206] In cases where embodiments have been previously described with reference to specific hardware architectures or features, in a simulation embodiment, equivalent functionality may be provided by appropriate software architectures or features. For example, a specific circuit may be implemented as computer program logic in a simulation embodiment. Similarly, memory hardware such as registers or caches may be implemented as software data structures in a simulation embodiment. In arrangements where one or more of the hardware elements referred to in the previously described embodiments are present on the host hardware (e.g., host processor 1330), some simulation embodiments may utilize the host hardware as appropriate.
[0207] The simulator program 1310 can be stored on a computer-readable storage medium (which can be a non-transitory medium) and provides a program interface (instruction execution environment) for the target code 1300 (which can include an application, an operating system, and a hypervisor). This program interface is the same as the interface of the hardware architecture modeled by the simulator program 1310. Therefore, the program instructions of the target code 1300 (including the translation table entry load / store trigger instructions described above) can be executed from within the instruction execution environment using the simulator program 1310, such that the host computer 1330, which does not actually have the hardware features of the device 2 discussed above, can emulate these features.
[0208] Therefore, the simulator program 1310 can have a handler logic 1312 that emulates the state of the processing circuit 4 described above. For example, the handler logic 1312 can emulate a transition of the execution state in response to an event that occurs during the simulated execution of the target code 1300 and perform processing operations. An instruction decoding program logic 1314 (which can be regarded as part of the handler logic) decodes the instructions of the target code 1300 and maps these instructions to the corresponding instruction sets in the native instruction set of the host device 1330. A register emulation program logic 1316 maps the register access requested by the target code to an access to the corresponding data structures maintained on the host hardware of the host device 1330, such as by accessing data in the registers or the memory 1332 of the host device 1330. The memory access program logic 1318 has an address translation program logic 1319 for implementing address translation, page table walking, and access control checking in a manner corresponding to the MMU 28 described in the hardware implementation embodiment above, and also has an additional function of mapping the simulated entity address obtained by address translation based on the translation table defined for the target code 1300 to the host virtual address for accessing the host memory 1332. These host virtual addresses themselves can be translated into host physical addresses using the standard address translation mechanism supported by the host (translating the host virtual address into the host physical address is outside the scope controlled by the simulator program 1310).
[0209] Therefore, by supporting the translation table entry load / store operation with the same functionality as discussed above within the simulator program 1310, the target code 1300 written for the device 2 that supports this operation in hardware is presented with the same architecture interface (e.g., CPU instructions in the instruction set architecture and / or the accelerator programming interface as Figure 12 shown), which would be available in the hardware device 2, such that it can also be executed on the host device 1330 that does not have this hardware.
[0210] In the present application, the phrase "configured to" is used to mean that an element of a device has a configuration capable of performing the defined operation. In this context, "configuration" means an arrangement or manner of interconnection of hardware or software. For example, the device may have dedicated hardware that provides the defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not mean that the device element needs to be changed in any way to provide the defined operation.
[0211] In the present application, a list of features prefaced by the phrase "at least one of" means that any one or more of those features may be provided individually or in combination. For example, "at least one of [A], [B], and [C]" encompasses any of the following options: A alone (without B or C), B alone (without A or C), C alone (without A or B), a combination of A and B (without C), a combination of A and C (without B), a combination of B and C (without A), or a combination of A, B, and C.
[0212] Although the illustrative embodiments of the present invention have been described in detail herein with reference to the accompanying drawings, it should be understood that the present invention is not limited to those exact embodiments, and that various changes and modifications may be effected therein by those skilled in the art without departing from the scope of the invention as defined by the appended claims.
Claims
1. An apparatus, the apparatus comprising: processing circuitry for processing instructions; address translation circuitry for translating between an input address space and an output address space based on address mapping information obtained from a translation table structure including translation table entries; and translation table entry load / store circuitry for performing a translation table entry load / store operation for at least one target translation table entry address in response to a translation table entry load / store trigger instruction processed by the processing circuitry, the at least one target translation table entry address being selected according to software-defined address information identifying a selected address in the input address space, each target translation table entry address including the address of a leaf translation table entry providing the address mapping information for translating the selected address from the input address space to the output address space or the address of a branch translation table entry traversed in a translation table walk operation for obtaining the leaf translation table entry providing the address mapping information for translating the selected address; wherein for at least one variant of the translation table entry load / store operation, for a given target translation table entry among the at least one target translation table entries, the translation table entry load / store operation supports clearing access trace metadata of the given target translation table entry from a first state indicating that at least one load / store access to a corresponding region of the input address space has occurred to a second state indicating that no load / store access to the corresponding region of the input address space has occurred.
2. The apparatus according to claim 1, wherein in response to the translation table entry load / store trigger instruction, the translation table entry load / store circuitry is configured to control the address translation circuitry to identify the at least one target translation table entry address based on the selected address.
3. The device according to any one of claims 1 and 2, wherein, For at least one variant of the translation table entry load / store operation, the at least one translation table entry includes the leaf translation table entry.
4. The apparatus according to claim 3, wherein for at least one variant of the translation table entry load / store operation, the translation table entry load / store operation supports updating the leaf translation table entry to change at least one of: the address mapping information for translating the selected address; access permission information indicating which types of memory access operations are permitted; memory attribute information for controlling the disposition of memory access to the selected address.
5. The apparatus according to any one of the preceding claims, wherein for at least one variant of the translation table entry load / store operation, the at least one translation table entry includes a translation table entry at a specified level of the translation table structure, whether the translation table entry is the leaf translation table entry or the branch translation table entry.
6. The apparatus according to any one of the preceding claims, wherein for at least one variant of the translation table entry load / store operation, the at least one translation table entry includes the leaf translation table entry and each branch translation table entry traversed in the translation table walk operation for obtaining the leaf translation table entry.
7. The apparatus according to any one of the preceding claims, wherein for at least one variant of the translation table entry load / store operation, the at least one translation table entry includes: the leaf translation table entry when the leaf translation table entry is effectively defined for the selected address; and in the case where a valid leaf translation table entry is not defined for the selected address, the final valid branch translation table entry reached in the traversal of the translation table structure at the selected address.
8. The apparatus according to any one of the preceding claims, wherein the at least one variant of the translation table entry load / store operation that supports clearing of the access trace metadata includes at least one of the following: a store variant of the translation table entry load / store operation, the store variant of the translation table entry load / store operation being for updating the given target translation table entry to an update value specified by a store data operand; a swap variant of the translation table entry load / store operation, the swap variant of the translation table entry load / store operation being for updating the given target translation table entry to an update value specified by a swap data operand and loading a software-accessible location with a pre-update value or a post-update value of the given target translation table entry; an atomic compare-and-swap variant of the translation table entry load / store operation, the atomic compare-and-swap variant of the translation table entry load / store operation being for determining whether a result of a comparison between the given translation table entry and a compare operand satisfies a comparison condition, and in response to determining that the result of the comparison satisfies the comparison condition, updating the given target translation table entry based on a swap operand of the compare-and-swap variant; and an atomic bit-update variant of the translation table entry load / store operation, the atomic bit-update variant of the translation table entry load / store operation being for setting or clearing at least one designated bit of the given translation table entry, the at least one designated bit being identified by a bit selection operand.
9. The apparatus according to any one of the preceding claims, wherein the translation table entry load / store circuit is configured to support a load variant of the translation table entry load / store operation to load the at least one target translation table entry into at least one software-accessible register.
10. The apparatus according to any one of the preceding claims, wherein in response to the translation table entry load / store trigger instruction, the translation table entry load / store circuit is configured to perform an error reporting action in response to identifying that an error condition has occurred, the error condition including one of the following: no valid leaf translation table entry is defined for the selected address; and Define a valid leaf translation table entry for the selected address at a level outside the expected level of the translation table structure.
11. The apparatus according to any one of the preceding claims, wherein in response to the translation table entry load / store trigger instruction, the translation table entry load / store circuit is configured to update at least one software-accessible register using comprehensive information that specifies one of the following: The level of the translation table structure that defines a valid leaf translation table entry for the selected address; and The information specified by the at least one target translation table entry.
12. The apparatus according to any one of the preceding claims, wherein the translation table entry load / store trigger instruction designates the selected address as an operand of the translation table entry load / store trigger instruction, and the translation table entry load / store trigger instruction has an instruction opcode different from that of the load / store instruction for triggering a load / store operation to be performed on the address designated as the operand of the load / store instruction.
13. The apparatus according to any one of claims 1 to 11, wherein the translation table entry load / store trigger instruction includes a store instruction that designates a predetermined translation table entry load / store trigger address as a store address operand, and for the store address operand, storing to the address triggers the translation table entry load / store circuit to perform the translation table entry load / store operation.
14. The apparatus according to claim 13, wherein the translation table entry load / store circuit is configured to obtain the software-defined address information from a memory-based data structure accessed based on a software-programmable base address; and When the software-defined address information specifies a plurality of selected addresses, the translation table entry load / store circuit is configured to perform the translation table entry load / store operation for each of the selected addresses in response to executing a single instance of the store instruction that serves as the translation table entry load / store trigger instruction.
15. The apparatus according to claim 14, wherein the translation table entry load / store circuit is configured to update a software-accessible location to specify a progress indicator that indicates the progress completed when performing the translation table entry load / store operation for the plurality of selected addresses.
16. The apparatus according to any one of claims 14 and 15, wherein In the case where the translation table entry load / store circuit is triggered to perform the translation table entry load / store operation for the plurality of selected addresses in response to the store instruction processed by the processing circuit in a higher privilege execution state and the processing circuit then switches to a lower privilege execution state: The translation table entry load / store circuit is capable of continuing to process the remaining addresses among the plurality of selected addresses after switching to the lower privilege execution state.
17. The apparatus according to any one of the preceding claims, wherein for at least one variant of the translation table entry load / store operation, the software-defined address information specifies an address range, and the translation table entry load / store circuit is configured to perform the translation table entry load / store operation for each address in the range that is the selected address.
18. The apparatus according to any one of the preceding claims, wherein: the address translation circuit is configured to support two-stage address translation between the virtual address space and the physical address space based on a first translation table structure providing address mapping information for translation between the virtual address space and an intermediate address space and a second translation table structure providing address mapping information for translation between the intermediate address space and the physical address space; and the translation table entry load / store circuit is configured to support at least one of: a first-stage variant of the translation table entry load / store operation, wherein the selected address includes a virtual address specified in the virtual address space, and the at least one target translation table entry includes at least one translation table entry of the first translation table structure; and a second-stage variant of the translation table entry load / store operation, wherein the selected address includes an intermediate address specified in the intermediate address space, and the at least one target translation table entry includes at least one translation table entry of the second translation table structure.
19. A method, the method comprising: using a processing circuit to process instructions; and using an address translation circuit to perform translation between an input address space and an output address space based on address mapping information obtained from a translation table structure including translation table entries; and in response to the processing circuit processing a translation table entry load / store trigger instruction, performing a translation table entry load / store operation for at least one target translation table entry address, the at least one target translation table entry address being selected according to software-defined address information identifying a selected address in the input address space, each target translation table entry address including the address of a leaf translation table entry providing the address mapping information for translating the selected address from the input address space to the output address space or the address of a branch translation table entry traversed in a translation table walk operation for obtaining the leaf translation table entry providing the address mapping information for translating the selected address; wherein for at least one variant of the translation table entry load / store operation, for a given target translation table entry among the at least one target translation table entries, the translation table entry load / store operation supports clearing access tracking metadata of the given target translation table entry from a first state indicating that at least one load / store access to a corresponding region of the input address space has occurred to a second state indicating that no load / store access to the corresponding region of the input address space has occurred.
20. A computer program comprising instructions for controlling a host data processing device to provide an instruction execution environment for executing object code, the computer program comprising: Address translation program logic for translating between an input address space and an output address space based on address mapping information obtained from a translation table structure including translation table entries; And Translation table entry load / store program logic for performing a translation table entry load / store operation for at least one target translation table entry address in response to a translation table entry load / store trigger instruction of the object code, the at least one target translation table entry address being selected according to software-defined address information identifying a selected address in the input address space, each target translation table entry address including the address of a leaf translation table entry providing the address mapping information for translating the selected address from the input address space to the output address space or the address of a branch translation table entry traversed in a translation table walk operation for obtaining the leaf translation table entry providing the address mapping information for translating the selected address; Wherein for at least one variant of the translation table entry load / store operation, for a given target translation table entry in the at least one target translation table entry, the translation table entry load / store operation supports clearing access tracking metadata of the given target translation table entry from a first state indicating that at least one load / store access to a corresponding region of the input address space has occurred to a second state indicating that no load / store access to the corresponding region of the input address space has occurred.
21. A storage medium storing the computer program according to claim 20.
Citation Information
Cited By
Methods, products, and media for use in ai accelerator chip
CN121918888A