TLB invalidation instruction

The TLB invalidation instruction with domain-specific control enhances processor performance in multi-core and multi-chiplet systems by efficiently managing TLB invalidation across non-overlapping domains, addressing the performance challenges of modern distributed processing systems.

WO2025210328A1PCT designated stage Publication Date: 2025-10-09ARM LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2025/050322
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-02
Filing Date
2025-02-20
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

In modern multi-core and multi-chiplet processing systems, the performance impact of translation lookaside buffer (TLB) invalidation operations is significant due to increased communication latencies and delays in completing TLB invalidation across numerous processor cores and chiplets, especially in systems where components are distributed across multiple integrated circuits.

Method used

A TLB invalidation instruction (TLBI) that specifies TLBI domains, allowing software to indicate which domains need to observe the invalidation event, thereby limiting the scope of TLB invalidation requests and reducing delays by using non-overlapping domains and hint information to avoid unnecessary broadcasts.

Benefits of technology

This approach improves processor performance by reducing the latency associated with waiting for acknowledgments and completing TLB invalidation operations, supporting scalable processor architectures in many-core and multi-chiplet systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025050322_09102025_PF_FP_ABST
    Figure GB2025050322_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A translation lookaside buffer invalidation (TLBI) instruction is described for instructing at least one translation lookaside buffer (TLB) to observe a TLB invalidation event. The TLBI instruction specifies TLBI domain information indicating, for each of two or more TLBI domains (including at least two non-overlapping TLBI domains), whether that TLBI domain is a target TLBI domain for which any TLBs are required to observe the TLB invalidation event or a non-target TLBI domain not required to observe the TLB invalidation event. TLBI broadcast control circuitry determines, based on the TLBI domain information specified by the TLBI instruction decoded by the instruction decoding circuitry, one or more TLBI request signals to be issued for instructing TLBs in one or more TLBI domains to observe the TLB invalidation event, and limits which TLBI domains are instructed to observe the TLB invalidation event based on the TLBI domain information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TLB INVALIDATION INSTRUCTION

[0002] The present technique relates to the field of data processing.

[0003] A data processing system may have at least one translation lookaside buffer (TLB) for caching address translation data used for translating addresses for memory system accesses. The address translation data in the TLB may depend on page table entries of one or more page tables which are stored in the memory system. By caching the address translation data in the TLB, addresses can be translated faster than if the page tables had to be looked up from memory every time an address translation is required. A TLB invalidation mechanism may be supported so that, if a change is made to the page tables (for example an operating system may change the memory mapping being used for a given software process or context), cached address translation information which is at risk of being out of date following the change to the page tables can be requested to be invalidated from a TLB.

[0004] At least some examples of the present technique provide an apparatus comprising: instruction decoding circuitry to decode a translation lookaside buffer invalidation, TLBI, instruction for instructing at least one translation lookaside buffer, TLB, to observe a TLB invalidation event, the TLBI instruction specifying TLBI domain information indicating, for each TLBI domain of a plurality of TLBI domains, whether that TLBI domain is a target TLBI domain for which any TLBs in that TLBI domain are required to observe the TLB invalidation event or a nontarget TLBI domain for which any TLBs in that TLBI domain are not required to observe the TLB invalidation event, the plurality of TLBI domains including at least two non-overlapping TLBI domains; and TLBI broadcast control circuitry to determine, based on the TLBI domain information specified by the TLBI instruction decoded by the instruction decoding circuitry, one or more TLBI request signals to be issued for instructing TLBs in one or more TLBI domains to observe the TLB invalidation event; wherein the TLBI broadcast control circuitry is configured to limit which TLBI domains are instructed to observe the TLB invalidation event based on the TLBI domain information.

[0005] At least some examples of the present technique provide computer-readable code for fabrication of an apparatus comprising: instruction decoding circuitry to decode a translation lookaside buffer invalidation, TLBI, instruction for instructing at least one translation lookaside buffer, TLB, to observe a TLB invalidation event, the TLBI instruction specifying TLBI domain information indicating, for each TLBI domain of a plurality of TLBI domains, whether that TLBI domain is a target TLBI domain for which any TLBs in that TLBI domain are required to observe the TLB invalidation event or a non-target TLBI domain for which any TLBs in that TLBI domain are not required to observe the TLB invalidation event, the plurality of TLBI domains including at least two non-overlapping TLBI domains; and TLBI broadcast control circuitry to determine, based on the TLBI domain information specified by the TLBI instruction decoded by the instruction decoding circuitry, one or more TLBI request signals to be issued for instructing TLBs in one or more TLBI domains to observe the TLB invalidation event; wherein the TLBI broadcast control circuitry is configured to limit which TLBI domains are instructed to observe the TLB invalidation event based on the TLBI domain information.

[0006] At least some examples of the present technique provide a method comprising: decoding a translation lookaside buffer invalidation, TLBI, instruction for instructing at least one translation lookaside buffer, TLB, to observe a TLB invalidation event, the TLBI instruction specifying TLBI domain information indicating, for each TLBI domain of a plurality of TLBI domains, whether that TLBI domain is a target TLBI domain for which any TLBs in that TLBI domain are required to observe the TLB invalidation event or a non-target TLBI domain for which any TLBs in that TLBI domain are not required to observe the TLB invalidation event, the plurality of TLBI domains including at least two non-overlapping TLBI domains; and determining, based on the TLBI domain information specified by the TLBI instruction decoded by the instruction decoding circuitry, one or more TLBI request signals to be issued for instructing TLBs in one or more TLBI domains to observe the TLB invalidation event, to limit which TLBI domains are instructed to observe the TLB invalidation event based on the TLBI domain information.

[0007] At least some examples provide a computer program for controlling a host data processing apparatus to provide an instruction execution environment for execution of target program code, the computer program comprising: instruction decoding program logic to decode a translation lookaside buffer invalidation, TLBI, instruction for instructing at least one simulated translation lookaside buffer, TLB, to observe a TLB invalidation event, the TLBI instruction specifying TLBI domain information indicating, for each TLBI domain of a plurality of TLBI domains, whether that TLBI domain is a target TLBI domain for which any simulated TLBs in that TLBI domain are required to observe the TLB invalidation event or a non-target TLBI domain for which any simulated TLBs in that TLBI domain are not required to observe the TLB invalidation event, the plurality of TLBI domains including at least two non-overlapping TLBI domains; and TLBI control program logic to determine, based on the TLBI domain information specified by the TLBI instruction decoded by the instruction decoding program logic, one or more simulated TLBs in one or more TLBI domains which should observe the TLB invalidation event; wherein the TLBI broadcast control program logic is configured to limit which TLBI domains should observe the TLB invalidation event based on the TLBI domain information.

[0008] At least some examples provide a storage medium storing the computer-readable code or computer program described above. The storage medium may be a non-transitory storage medium.

[0009] Further aspects, features and advantages of the present technique will be apparent from the following description of examples, which is to be read in conjunction with the accompanying drawings, in which:

[0010] Figure 1 illustrates an example of a processing system comprising TLBI domains;

[0011] Figure 2 illustrates an example of a processing element comprising instruction decoding circuitry and TLB invalidation (TLBI) broadcast control circuitry; Figure 3 illustrates an example encoding for a TLBI instruction and a register operand of the TLBI instruction;

[0012] Figure 4 illustrates an example of shareability domains and TLBI domains;

[0013] Figures 5 and 6 illustrate examples of assigning TLBI domains in a processing system comprising multiple chiplets;

[0014] Figure 7 illustrates steps for processing a TLBI instruction;

[0015] Figure 8 illustrates steps for target domain remapping of virtual TLBI target domains to physical TLBI target domains based on target domain remapping information;

[0016] Figure 9 illustrates an example of architectural state stored in system registers for controlling target domain remapping;

[0017] Figure 10 illustrates an example of the target domain remapping information;

[0018] Figure 1 1 illustrates an example of target domain remapping based on the target domain remapping information;

[0019] Figure 12 illustrates a more detailed example of steps for processing of a TLBI instruction;

[0020] Figure 13 illustrates an example of hardware supporting over-invalidation where at least one non-target TLBI domain nevertheless observes the TLB invalidation event;

[0021] Figure 14 illustrates steps for processing a synchronisation barrier instruction based on tracking of outstanding TLB invalidation events; and

[0022] Figure 15 illustrates a simulation example.

[0023] A processor architecture may support a translation lookaside buffer invalidation instruction (TLBI instruction) for instructing at least one translation lookaside buffer (TLB) to observe a TLB invalidation event. In some processor architectures, TLBI instructions may only have an effect on local TLBs associated with the particular processor core which execute the TLBI instruction, so if address translation information from a given set of page tables has been cached in TLBs associated with multiple processor cores, the mechanism for causing each of those cores’ TLBs to invalidate corresponding page table information may be for one processor core to raise interprocessor interrupts which interrupt the software being processed on other cores and cause those cores to execute TLBI instructions with local effect. However, such interrupts have a significant performance impact, and so some processor architectures support TLBI instructions which trigger broadcasting of TLBI request signals to many TLBs across a processing system, including broadcasting to TLBs not immediately associated with the processor core executing the TLBI instruction.

[0024] Hence, an apparatus may comprises instruction decoding circuitry to decode a TLBI instruction for instructing at least one TLB to observe a TLB invalidation event, and TLBI broadcast control circuitry to determine, based on the decoded TLBI instruction, one or more TLBI request signals to be issued (e.g. broadcast) for instructing TLBs to observe the TLB invalidation event. Such TLBI instructions with support for broadcasting of TLBI request signals can improve processor performance by eliminating a need to interrupt each processor’s workload to allow for invalidation of cached address translation information in corresponding TLBs.

[0025] However, a problem with such TLBI broadcast mechanisms is that in recent years the number of processor cores in a processing system is greatly increasing (e.g. many-core systems may have 10s, 100s or even 1000s of processor cores). If TLBI request signals are broadcast to TLBs associated with many processor cores, there may be a very long delay in completing the TLB invalidation operation for that TLBI instruction. The processor core having the instruction decoding circuitry which decoded the corresponding TLBI instruction may need to wait for some guarantee that each of the targeted TLBs have guaranteed that the TLBI request will be actioned, before being able to proceed with subsequent processing operations, and so the performance impact of TLBI operations is becoming significant as processing systems scale to increasing numbers of TLBs and processor cores in a processing system. This is a particular problem in multi-ch iplet systems, in which the components of a multi-core processing system are distributed across two or more distinct integrated circuits. A growing trend in processing system design is that, rather than implementing all components of a computing system on a single integrated circuit, the components are distributed across two or more smaller integrated circuits, known as “chiplets”, to enable use of different process nodes for different parts of the system, enable modular system design and / or improve manufacturing yields because each smaller chiplet has a lower probability of being subject to manufacturing defects because the probability of a given integrated circuit suffering from such defects increases with chip size. However, communication latencies between one chiplet and another may be higher than communication within a single chiplet, and so this may have a significant impact on TLBI performance. Therefore, it can be desirable to support architectural features of the TLBI instruction that can enable better TLBI performance in multi-core and multi-chiplet systems.

[0026] In the examples discussed below, the TLBI instruction specifies TLBI domain information indicating, for each TLBI domain of a plurality of TLBI domains, whether that TLBI domain is a target TLBI domain for which any TLBs in that TLBI domain are required to observe the TLB invalidation event or a non-target TLBI domain for which any TLBs in that TLBI domain are not required to observe the TLB invalidation event. The TLBI domains include at least two nonoverlapping TLBI domains (domains for which any given TLB is a member of only one of those non-overlapping domains). The TLBI broadcast control circuitry can determine the one or more TLBI request signals based on the TLBI domain information specified by the TLBI instruction, so that the TLBI broadcast circuitry can limit which TLBI domains are instructed to observe the TLB invalidation event based on the TLBI domain information.

[0027] With this approach, the TLBI instruction takes as an input the TLBI domain information allowing software to express, to the hardware, hint information indicating which domains of TLBs need, or do not need, to observe the TLB invalidation event. For example, the software may have knowledge that certain cores in a multi-core system are carrying out functions that are independent of the set of page tables whose update required the TLBI instruction to be executed, and so can hint to the hardware that TLBs in certain TLBI domains do not need to observe the TLB invalidation event. Hence, the scope of the broadcasting of TLBI request signals can be limited, as the hardware now has hint information which it can use to avoid sending the TLBI request signals to domains which the software has indicated, using the TLBI domain information, as being non-target TLBI domains. This can reduce the delay associated with waiting for acknowledgements indicating that a TLBI invalidation operation is complete, and therefore support improved processing performance in modern many-core or multi-chiplet processing systems. Compared to cases where TLBI domains are specified as two or more concentric domains where each outer domain overlaps with the inner domain inside it, the use of nonoverlapping TLBI domains can better support a scalable processor architecture accommodating expansion of processing systems into many-core and multi-chiplet systems, as it enables better pinpointing of whether TLBs in two or more regions of a many-core system need to observe the TLB invalidation event, as those TLBs can be separated into distinct TLBI domains even when those regions are far from the processor core executing the TLBI instruction.

[0028] The TLBI domain information could be specified in different ways. In some examples, the TLBI instruction could specify an immediate field (directly within its instruction encoding) which specifies the TLBI domain information.

[0029] However, the encoding space available within the instruction encoding itself may be extremely constrained, which might limit the maximum number of TLBI domains supported if an immediate field is used to specify the TLBI domain information. Hence, to help support scaling to larger systems with greater number of TLBI domains, it can be useful for the TLBI instruction to specify a register identifier indicating a register storing a register operand, with the TLBI domain information being encoded within the register operand.

[0030] In some examples, the register operand which provides the TLBI domain information could be a dedicated TLBI domain identifying register operand which does not provide any other information, other than the TLBI domain information.

[0031] However, the register width may be large enough that it is possible to combine the TLBI domain information with other information in a single register. Hence, a more efficient instruction encoding can be achieved if the register identifier of the TLBI instruction identifies a register storing both the TLBI domain information and other information.

[0032] For example, the TLB invalidation event may comprise invalidation of TLB entries that satisfy at least one invalidation condition; and the at least one invalidation condition may depend on invalidation condition specifying information specified in the register operand. For example, the invalidation condition specifying information could specify criteria for determining which TLB entries would need to be invalidated by the TLBs in the target TLBI domains in response to the TLBI instruction. For example, the invalidation condition specifying information could specify an address used to filter TLB entries for invalidation based on the address they translate, or could specify an address translation context (associated with the set of page tables whose update required the TLBI instruction to be executed) so that each TLB in the target domains can filter whether a given TLB entry needs invalidating based on whether an address translation context of that TLB entry matches the address translation context specified for the TLBI instruction. It will be appreciated that the invalidation condition(s) to be applied could also depend on other information encoded directly in the TLBI instruction (e.g. part of the instruction opcode), as well as the invalidation condition specifying information provided by the register operand. Nevertheless, implementing the TLBI domain information in the same register operand already used to provide invalidation condition specifying information allows the TLBI domain information to be supported without any increase in the number of register operands specified by the TLBI instruction, which can help conserve instruction encoding space and make it easier for the TLBI instruction supporting the TLBI domain information to be backwards compatible with legacy software written for systems not supporting the TLBI domain information. For example, the instruction encoding of the TLBI instruction may be the same as in a legacy TLBI instruction not supporting TLBI domain information, but when supporting the use of TLBI domain information, the instruction decoding circuitry and TLBI broadcast control circuitry may respond differently to part of the register operand compared to systems designed based on a legacy version of the TLBI instruction.

[0033] The TLBI domain information is a hint expressed by software indicating which TLBI domains are, or are not, required to observe the TLB invalidation event. The TLBI broadcast control circuitry can use this information to determine which TLBI domains should be instructed to observe the TLB invalidation event. However, for a number of reasons, the TLBI domains which are actually instructed to observe the TLB invalidation event may not necessarily be the same as the target TLBI domains indicated in the TLBI domain information.

[0034] For example, some implementations may support the ability to remap virtual target TLBI domains specified in the TLBI domain information of the TLBI instruction to physical target TLBI domains which are actually required to observe the TLB invalidation event. Hence, in at least one operating state where virtualized remapping of target TLBI domains is enabled, the TLBI broadcast control circuitry may determine the one or more TLBI request signals based on the TLBI domain information specified by the TLBI instruction and target domain remapping information specifying a mapping between a set of virtual target TLBI domains specified using the TLBI domain information and a set of physical target TLBI domains that are actually required to observe the TLB invalidation event. This can be helpful in systems supporting virtualisation, where a hypervisor may manage multiple operating systems executing on the same physical platform. From the point of view of an operating system, the operating system may assume that a certain processing workload has been allocated to one or more processor cores associated with TLBs in a given TLBI domain, and so when changing page table information associated with that workload, may include the given TLBI domain in the set of virtual target TLBI domain specified using the I LBI domain information. However, the hypervisor may be aware that actually, the operations associated with the operating system have been assigned to physical processor cores in a different manner, so that the virtual view of the TLBI domains assumed by the operating system may not necessarily match the physical TLBI domains which actually depend on page table information for which the TLB invalidation event instructed by a TLBI instruction included in program code of the operating system is relevant. For example, the hypervisor might have migrated the workload associated with the operating system to a processor core in a different TLBI domain from the one expected by the operating system. Hence, providing architectural support for the target TLBI domains to be remapped based on target domain remapping information can allow the TLBI domains which actually require the TLB invalidation event to be observed to be identified without needing to the operating system to be exposed to information indicating how the operating system’s activities have been scheduled on particular cores of the physical platform.

[0035] In some examples, the target domain remapping information may comprise a plurality of remapping fields. Each remapping field is associated with a corresponding virtual TLBI domain and specifies which of a plurality of physical TLBI domains are to be regarded as a physical target TLBI domain when the corresponding virtual TLBI domain is indicated as being a virtual target TLBI domain by the TLBI domain information. The set of physical target TLBI domains which are required to observe the TLB invalidation event can comprise any of the plurality of physical TLBI domains that are indicated as being a physical target TLBI domain by any one or more of a subset of remapping fields that are associated with the set of virtual target TLBI domains indicated by the TLBI domain information. Remapping fields associated with virtual TLBI domains indicated as non-target TLBI domains by the TLBI domain information are ignored for the purpose of selecting the physical target TLBI domains (so even if one of the remapping fields associated with a given virtual non-target TLBI domain specifies a given physical TLBI domain as being a target TLBI domain, that given physical TLBI domain would not be selected as a physical target TLBI domain unless one of the other remapping fields associated with a virtual target TLBI domain specifies that given physical TLBI domain as a physical target TLBI domain). With this approach, the hypervisor has flexibility to map an indication that a given virtual TLBI domain is a target TLBI domain to any combination of target / non-target domain settings for each of the plurality of physical TLBI domains.

[0036] In some examples, virtualized remapping of TLBI domains may be considered permanently enabled, without any ability for software to control whether virtualized remapping of TLBLI domains is enabled or disabled.

[0037] However, in some examples, it can be helpful to support the ability for the TLBI broadcast control circuitry to determine whether virtualized remapping of target TLBI domains is enabled based on virtualized remapping enable control information configurable by software. Hence, if the system is operating in a non-virtualized mode in which a single operating system is resident on the hardware platform and the operating system has direct control of the scheduling of tasks on physical processor cores, then virtualized remapping of target TLBI domains can be disabled to reduce complexity and avoid overheads associated with maintenance of the target domain remapping information. When virtualized remapping of target TLBI domains is disabled, there may be a one-to-one mapping of virtual target TLBI domains to physical target TLBI domains, so that the physical target TLBI domains are simply the set of (virtual) target TLBI domains specified by the TLBI domain information. When virtualized remapping is disabled, the selection of the one or more TLBI domains for which TLBI request signals are issued may be controlled independent of the target domain remapping information. The virtualized remapping enable control information can be information specified in at least one control register. The at least one control register used to specify the virtualized remapping enable control information may comprise one or more registers to which write access is denied for instructions executed with less privilege than a hypervisor-level privilege. Hence, a hypervisor may control whether or not virtualized remapping of target TLBI domains is enabled or disabled at a given time.

[0038] In some examples, the target domain remapping information comprises information specified in at least one control register. The at least one control register used to specify the target domain remapping information may comprise one or more registers to which write access is denied for instructions executed with less privilege than a hypervisor-level privilege. Hence, the hypervisor can control which particular mappings are used between virtual and physical target TLBI domains, based on the way in which it has chosen to schedule the activities of certain guest operating systems on the physical processor cores of a processing system.

[0039] In some examples, when virtualized remapping of target TLBI domains is enabled, the virtualized remapping may occur in at least one operating state, but there may be at least one other operating state of the apparatus for which virtualized remapping of target TLBI domains is not performed even when indicated as enabled by the virtualized remapping enable control information. For example, virtualized remapping of target TLBI domains may not be relevant when executing in an operating state with hypervisor-level privilege (or a higher level of privilege), as the hypervisor or other software more privileged than the hypervisor may already have a view of the actual activity on physical processor cores of the system, so does not need to have its indications of TLBI domains remapped. In some examples, the at least one operating state (for which the virtual TLBI domains specified by the TLBI domain information are remapped to physical TLBI domains when virtualized remapping is enabled) may comprise an operating state in which instructions are executed with less privilege than a hypervisor-level privilege (e.g. an operating state associated with operating-system-level privilege).

[0040] In some examples, the TLBI broadcast control circuitry is configured to determine that a local TLBI domain is a target TLBI domain even if the local TLBI domain is determined to be a non-target TLBI domain using the TLBI domain information (the local TLBI domain could be determined to be a non-target TLBI domain either because it is indicated as a non-target TLBI domain by the TLBI domain information itself, if virtualized remapping is not enabled, or if the local TLBI domain is determined to be a physical non-target TLBI domain based on the remapping of virtual to physical TLBI domains described above when virtualized remapping is enabled). The local TLBI domain may comprise a TLBI domain which is associated with at least one TLB of a processor comprising the instruction decoding circuitry which decoded the TLBI instruction. While not essential, this approach can be helpful to reduce risk of unpredictable system operation caused by programming error in setting the TLBI domain information. It would be very unusual for a given processor core to execute a TLBI instruction requesting that TLBs observe a TLB invalidation event which is not relevant to its own local TLBs in its local TLBI domain. In general, the most likely TLBI domains that may be non-target TLBI domains not required to observe the TLB invalidation event may be TLBI domains further from the processor core executing the TLBI instruction. Also, even if the local TLBI domain actually does not need to observe the TLB invalidation event, the local TLBI domain is likely to experience the shortest delay in processing the TLB invalidation event compared to other TLBI domains further from the processor core that executes the TLBI instruction, so the performance cost of over-invalidating in TLBs in the local TLBI domain is negligible. Therefore, when designing the instruction set architecture, given that there is relatively little benefit to indicating a core’s own TLBI domain as a non-target TLBI domain, it may be useful to constrain the architecture so that the local TLBI domain associated with the processor that executes the TLBI instruction should be regarded as a target TLBI domain regardless of whether it is indicated as a non-target TLBI domain using the TLBI domain information. This approach can make the system more robust against programming errors, reducing risk of incorrect system functioning due to out of date address translation information remaining valid in a TLB within a processor’s local TLBI domain after that processor has instructed a TLB invalidation event to be observed.

[0041] Different formats or encodings of the TLBI domain information may be possible. In some examples, it may be possible to encode the TLBI domain information without requiring a separate target / non-target indicator per TLBI domain. For example, a target TLBI domain combination value may take one of a number of encodings, each encoding corresponding to a certain predetermined subset of the TLBI domains being indicated as target or non-target TLBI domains.

[0042] However, in some examples, for at least one variant of the TLBI instruction, the TLBI domain information may comprise a plurality of TLBI domain indicators each associated with a corresponding TLBI domain and indicating whether that TLBI domain is a target TLBI domain or a non-target TLBI domain. For example, when the TLBI domain indicator for a given TLBI domain has a first value the given TLBI domain may be determined to be a target TLBI domain, and when the TLBI domain indicator for the given TLBI domain has a second value the given TLBI domain may be determined to be a non-target TLBI domain. The first value and second value could for example be 0 and 1 , or 1 and 0, respectively. In some examples, for at least one variant of the TLBI instruction, the TLBI domain information specifies at least one TLBI domain number identifying at least one target TLBI domain. For example, the TLBI domain information can be interpreted as an integer identifying the TLBI domain which is to be considered as the target TLBI domain (or if there is sufficient encoding space, there may be support for identifying two or more integers identifying the TLBI domain numbers of corresponding target TLBI domains).

[0043] Some examples may support only one of the variants described above (either using a mask of TLBI domain indicators per TLBI domain, or using one or more TLBI domain number fields, to identify the target TLBI domains). Alternatively, the ISA may allow for both variants, but it may be a specific design choice of a particular hardware implementation which interpretation of the TLBI domain information is used and so a given hardware implementation may only support one variant (in that case, software can identify the particular encoding of the TLBI domain information used for a given hardware implementation based on firmware tables or other control structures used to identify implementation-specific system properties).

[0044] Other examples could support both variants on the same hardware implementation, so that the software developer can choose which encoding is used to identify the target TLBI domains. For example, the variants of the TLBI instruction could be distinguished by their instruction encodings (e.g. by an opcode), by part of the TLBI domain information or another parameter of the instruction, or by a modal control bit stored in a control register which may control whether a given instance of the TLBI instruction has its TLBI domain information interpreted as a mask of TLBI domain indicators or as one or more TLBI domain number fields.

[0045] For backwards compatibility reasons, it can be particularly useful for the first value (indicating a target TLBI domain) to be 0 and the second value to be 1 (indicating a non-target TLBI domain). This may appear to be counter-intuitive as one might expect that the target TLBI domains should be the domains positively identified using bit indicators of 1 . However, encoding the TLBI domain information so that bits of 0 indicate target TLBI domains and indicators of 1 indicate non-target TLBI domains can be better for backwards compatibility with legacy software, in cases where a portion of a register operand is used to indicate the TLBI domain information. The portion of the register operand used for the TLBI domain information may correspond to a portion which in legacy software would have been assumed to have certain reserved values, which are typically 0. If legacy software has not set the TLBI domain information to other values including at least bit set to 1 , then it may be preferable for all the TLBI domains to be regarded as (virtual) target TLBI domains, to avoid accidentally preventing TLBs in a given TLBI domain failing to observe the TLB invalidation event when the legacy software has not explicitly expressed any hint that it would be safe for any TLBI domains to ignore the TLB invalidation event. Hence, it may be preferable for the encodings of the register field used to identify at least one non-target TLBI domain to be encodings which take values other than the default value used for the reserved portion in the legacy architecture. Therefore, as the legacy architecture may assume 0 for the reserved portion of the register, an encoding with non-target TLBI domains identified by indicators of 1 in that portion can better support backwards compatibility to ensure that legacy software continues to function correctly even when executed on a system implementing the new TLBI instruction with support for TLBI domain information.

[0046] In some examples, the plurality of TLBI domains comprise absolute domains for which an association between a given TLB and its associated TLBI domain is the same regardless of which of a plurality of processors executes the TLBI instruction. This differentiates the TLBI domains from an alternative approach where the scope of a TLB invalidation event could be identified using shareability domains defined relative to the processor core that executes the TLBI instruction. With shareability domains identified in relative terms, which TLBs are associated with a given shareability domain corresponding to a particular encoding of a parameter of the TLBI instruction may vary depending on whether the TLBI instruction is executed on one processor core or another. This approach may be effective in relatively small processing systems, but may not scale well to many-core or multi-chiplet systems because once a TLB invalidation event is considered relevant beyond the local confines of a given core, it is difficult for relative shareability domains to express which sub-regions of the wider processing system far from the processor core executing the TLBI instruction should observe the TLBI invalidation event. Hence, a processor architecture which is based on a definition of TLBI domains as absolute domains irrespective of which processor core is executing the TLBI instruction can be much more scalable to the needs of modern processing systems as they extend to greater numbers of processing cores and chiplets, enabling performance improvements for such larger systems.

[0047] In some examples, the plurality of TLBI domains may be non-overlapping so that each TLB is associated with only one TLBI domain. Hence, in this case it is not possible for any given TLB to be associated with more than one TLBI domain.

[0048] In some examples, the TLBI domain information may not be the only information used by the TLBI broadcast control circuitry to limit the scope of the TLBI broadcasting. In some examples, the TLBI instruction also specifies shareability domain information indicating one of a plurality of shareability domains for which TLBs associated with the selected shareability domain are to observe the TLB invalidation event. The shareability domains may be nested overlapping domains for which each larger domain in a hierarchy of shareability domains includes all smaller domains of the hierarchy. The shareability domains may be defined relative to the processor which executes the TLBI instruction, rather than in absolute terms.

[0049] Support for shareability domains, in addition to TLBI domains, may be useful to support legacy software which may have assumed that such shareability domain information is available and which may expect that this shareability domain information can continue to be used to specify the scope of TLB invalidation events. Hence, in some cases the TLBI broadcast control circuitry may use both the shareability domain information (expressing scope of a TLB invalidation event in terms of overlapping domains defined relative to the processor executing the TLBI instruction) and the I LBI domain information (expressing scope of a TLB invalidation event in terms of nonoverlapping domains defined in absolute terms irrespective of which processor executes the TLBI instruction) to determine the scope of the TLBI broadcast, and hence decide which TLBI domains should actually observe the TLB invalidation event.

[0050] In implementations supporting shareability domains, it is not necessary for the TLBI domain information to be considered for limiting the scope of TLBI broadcasts for all settings of the shareability information. In some examples, the TLBI broadcast control circuitry is configured to determine the one or more TLBI request signals based on the TLBI domain information when the shareability domain information indicates a selected shareability domain. For shareability domains other than the selected shareability domain, the TLBI domain information could be ignored. For example, the shareability domains may include a “non-shareable” domain (comprising the local TLBs associated with a processor core executing the TLBI instruction), an “inner shareable domain” which comprises the non-shareable domain and also includes a wider subset of TLBs in a region of the processing system comprising the processor core executing the TLBI instruction, and an “outer shareable domain” which comprises all TLBs in the processing system. In some cases, the TLBI domain information may be used to limit scope of TLBI broadcast for TLBI instructions which specify the inner shareable domain as the selected shareability domain, and the TLBI domain information may be ignored for instructions for which the shareability information specifies the non-shareable domain or the outer shareable domain.

[0051] The shareability domain information could be specified either using an opcode or other parameter directly encoding in the instruction encoding of the TLBI instruction, or using a portion of a register operand of the TLBI instruction, or using a combination of information specified directly in the instruction encoding and information specified in a register operand.

[0052] The TLBI domain information allows software to express to the hardware a hint about which TLBI domains may be non-target TLBI domains which are not required to observe the TLB invalidation event. However, this is merely a hint, and the hardware of the TLBI broadcast control circuitry may permit over-invalidation where at least one non-target TLBI domain not required to observe the TLB invalidation event nevertheless observes the TLB invalidation event. This flexibility to select whether non-target TLBI domains actually observe the TLB invalidation event can be helpful for system designers, as precisely ensuring that only the target TLBI domains observe the TLB invalidation event may require more complex circuit logic for requesting invalidation and tracking invalidation completion, which may not be justified in some scenarios. For example, in a multi-chiplet system a system designer may consider that the main benefit of use of the TLBI domain information may be to avoid TLBI broadcasting crossing a chiplet boundary if there are no target TLBI domains beyond that boundary, but once there is at least one target TLBI domain beyond the chiplet boundary, the latency associated with that target TLBI domain responding to the TLB invalidation requests may be significant enough that it then becomes irrelevant how many TLBI domains beyond the chiplet boundary are instructed to observe the TLB invalidation event, and so to reduce tracking complexity it may be simpler to request that all TLBI domains beyond the chiplet boundary observe the TLB invalidation event. The particular criteria for which TLBI domains are actually instructed to observe the TLB invalidation event may be highly dependent on the particular system implementation, which is not a consideration for the software developer or the implementation of the instruction set architecture (ISA) supporting the TLBI instruction. The provision of TLBI domain information gives flexibility for software to express a hint on the expected scope of a TLB invalidation event, so that hardware has more information available allowing it to make better decisions on how to handle that TLB invalidation event in the most performance-efficient manner. However, the ISA may not constrain any particular way in which decisions are made regarding whether non-target TLBI domains should observe the TLB invalidation event. Hence, this is another reason why the TLBI domains actually instructed to observe the TLB invalidation event may not necessarily correspond exactly to the target TLBI domains indicated by the TLBI domain information.

[0053] In some examples, the one or more TLBI request signals may specify target domain information indicating which TLBI domains are required to observe the TLB invalidation event. This can allow circuitry at a downstream location to make further decisions on routing of TLBI request signals. For example, the TLBI broadcast control circuitry on one chiplet may be responsible for determining whether it is necessary to provide TLBI request signals over a chiplet- to-chiplet bridge, and downstream circuitry on another chiplet may then use the target domain information specified in the TLBI request signals to make further decisions on which TLBI domains on that chiplet or on a further chiplet should be instructed to observe the TLB invalidation event.

[0054] However, other examples may issue one or more TLBI request signals that do not specify target domain information indicating particular TLBI domains as being required to observe the TLB invalidation event. For example, the TLBI broadcast control circuitry on one particular chiplet may have the responsibility of deciding whether to provide TLBI request signals over a chiplet-to- chiplet bridge, and for any TLB invalidation events where it is necessary to pass the TLBI request signals over the bridge, it may be assumed by default that any TLBI domains beyond that bridge should be instructed to observe the TLB invalidation. This means it is not necessary for those TLBI request signals to specify TLBI domain information identifying particular target TLBI domains.

[0055] In some examples (e.g. multi-chiplet systems), it can be useful for the TLBI broadcast control circuitry to determine, based at least on the TLBI domain information, whether it is necessary to issue at least one TLBI request signal to an off-chip TLBI domain located on a separate integrated circuit from an integrated circuit comprising the instruction decoding circuitry and the TLBI broadcast control circuitry. The use of TLBI domain information can help avoid incurring the significant latency cost of exposing the TLB invalidation to the off-chip TLBI domain if the TLBI domain information indicates that the off-chip TLBI domain is not a target TLBI domain. Once a TLBI instruction is executed, the software that included that TLBI instruction may in some examples need to wait for a guarantee that the TLB invalidation event has been observed by the required target TLBI domains, before it can proceed with subsequent operations. However, often a group of TLBI instructions may be executed associated with a given set of updates to page tables, and it may be very inefficient for the processor executing TLBI instructions to have to wait for individual confirmation of each separate TLB invalidation event instructed by the group of TLBI instructions before proceeding with the next instruction in the group. Therefore, some processor architectures may support a synchronisation barrier instruction which can simplify implementation of TLB acknowledgements, by providing a barrier which causes younger instructions than the barrier to be prevented from being executed until any older TLBI instructions prior to the barrier are guaranteed to be completed. The barrier can handle collective acknowledgement of an entire group of TLBI instructions, rather than implementing a separate barrier for each TLBI instruction individually, to avoid needing to serialize the response of TLBs to each individual TLBI instruction of the group. When TLBI domains are supported as discussed above, it can be useful for the response to the synchronisation barrier instruction to consider which TLBI domains have been instructed to observe a given TLB invalidation event. Hence, the apparatus may comprise TLBI tracking circuitry to track, with respect to a given processor, which TLBI domains have outstanding TLB invalidation events instructed based on TLBI instructions executed on the given processor. In response to a synchronisation barrier instruction executed on the given processor, the given processor may prevent instructions younger than the synchronisation barrier instruction from being executed until the TLBI tracking circuitry confirms that all TLBI domains instructed to observe TLB invalidation events based on older TLBI instructions which are older than the synchronisation barrier instruction have acknowledged that the TLB invalidation events instructed by the older TLBI instructions are guaranteed to be completed. By using the TLBI domain information to limit the scope of TLBI broadcasts, the latency associated with the synchronisation barrier instruction completing can be reduced, and hence processing performance can be improved for subsequently executed instructions that are waiting for completion of the synchronisation barrier before they can execute.

[0056] The instruction decoding circuitry, TLBI broadcast control circuitry (and if provided, TLBI tracking circuitry) may be circuitry associated with a particular processor core or cluster of processor cores, but do not necessarily need to be part of the same apparatus as other processor cores in a multi-core system. Circuit designs for a single processor core supporting a given ISA may be manufactured, licensed and / or sold separately from other components of a system into which that single processor core may be integrated. Hence, although the ISA of the single processor core supports features (such as the TLBI domain information) enabling that processor core design to be more scalable to being included in large many-core or multi-ch iplet systems, it is not necessary for that design of processor core to always be included in a many-core or multi- chiplet system. Some system designers may choose to include that processor core design in a system comprising relatively few processor cores or a system implemented as a system-on-chip on a single integrated circuit. Nevertheless, a processor core supporting the TLBI instruction specifying TLBI domain information as discussed above is a better processor core than without support for such a TLBI instruction, because it gives flexibility for hardware system designers to include that processor core in larger many-core or multi-chiplet systems to give better performance for software executing on those systems, which would not be practical if the TLBI instruction with TLBI domain information was not supported. Hence it will be appreciated that while the architectural features of the apparatus supporting the TLBI instruction have a particular benefit when the apparatus is included in a many-core or multi-chiplet system, it is not essential for the apparatus to always be used in that scenario.

[0057] The specific mapping of TLBI domains onto physical regions of a processing system may be a specific design choice for an individual system designer when designing a processing system including the apparatus comprising the instruction decoding circuitry and TLBI broadcast control circuitry. Hence, the definition of the TLBI instruction in the ISA supported by the apparatus may make no assumptions about what is meant by a particular TLBI domain. The software that executes on the apparatus (which is not a feature of the apparatus itself, as software can be installed at the point of use) may have knowledge about how specified TLBI domain numbers map to particular regions of the processing system. For example, the software may read firmware tables expressing, for a particular processing system design, the association between processor cores or TLBs and the TLBI domain numbers of particular TLBI domains, and can use those firmware tables to decide how to set TLBI domain information of the TLBI instruction to correspond with the way in which the software has chosen to allocate workloads to particular processor cores. The TLBI broadcast control circuitry for a given physical system implementation may then make system-dependent decisions on how to handle cases where certain TLBI domains are identified as target TLBI domains, such as considering the layout of inter-chiplet bridges as mentioned above. However, this particular control by the TLBI broadcast control circuitry is not architecturally defined and can vary significantly from one physical system implementation to another depending on the design features of a particular implementation.

[0058] The TLBI domain information may identify the TLBI domains based on a TLBI domain number space. For example, the TLBI domain information may specify a mask of TLBI domain indicators, where each domain indicator corresponds to a given TLBI domain having a given domain number in the TLBI domain number space, or could specify one or more TLBI domain numbers directly. As mentioned above, the ISA supported by the apparatus may make no assumptions about what is meant by a particular TLBI domain number. However, specific system implementations may interpret a given TLBI domain number in a particular way, to identify a corresponding set of TLBs associated with the domain.

[0059] In some examples, it can be useful for the system implementation to support TLBI domain aliasing, in which a first subset of TLBI numbers of the TLBI domain number space correspond to actual I LBI domains; a second subset of TLBI numbers of the TLBI domain number space correspond to aliased TLBI domains; and for a given TLBI instruction for which the TLBI broadcast control circuitry determines that a given aliased TLBI domain corresponding to one of the second subset of TLBI numbers is identified as a target TLBI domain using the TLBI domain information, the TLBI broadcast control circuitry maps the given aliased TLBI domain to a group of two or more actual TLBI domains and determine that said group of two or more actual TLBI domains should be considered target TLBI domains. For example, it may be considered that when a given aliased TLBI domain with a TLBI domain number X is specified as a target TLBI domain using the TLBI domain information (either directly by the TLBI domain information, or if virtualized remapping is performed, after TLBI domain X is identified as a physical TLBI domain identified by remapping the virtual TLBI domain specified using the TLBI domain information), this should cause the TLBI broadcast control circuitry to identify a corresponding set of actual TLBI domains with TLBI domain numbers A, B, C... to be identified as target TLBI domains. Support for TLBI domain aliasing can be useful to simplify the software control of the TLBI domain information as software can simply treat a given group of actual TLBI domains as if it were a single TLBI domain corresponding to one of the aliased TLBI domain numbers. The correspondence between TLBI domain numbers and the combination of one or more actual TLBI domains identified by that TLBI domain number may be identifiable by software based on firmware tables specified for the particular system implementation, so that software can understand which TLBs would correspond to a particular TLBI domain number in the TLBI domain number space.

[0060] The techniques discussed above may be implemented within an apparatus which has hardware circuitry provided for implementing the instruction decoding circuitry and TLBI broadcast circuitry discussed above. However, the same technique can also be implemented within a computer program which executes on a host data processing apparatus to provide an instruction execution environment for execution of target code. Such a computer program may control the host data processing apparatus to simulate the architectural environment which would be provided on a hardware apparatus which actually supports target code according to a certain instruction set architecture, even if the host data processing apparatus itself does not support that architecture. The computer program may have instruction decoding program logic and TLBI control program logic which emulates functions of the instruction decoding circuitry and TLBI broadcast control circuitry discussed above. The instruction decoding program logic may map the claimed TLBI instruction onto a corresponding set of instructions defined according to the host instruction set architecture supported by the host data processing apparatus. The simulation program may include program logic for simulating one or more TLBs. The TLBI control program logic may determine which of the simulated TLBs need to observe the TLB invalidation event requested by the TLBI instruction of the target code. Hence, in the description above, references to TLBs may, for the simulated embodiment, be understood as referring to simulated TLBs. Also, references to registers may be understood as referring to simulated registers of the target instruction set architecture which are mapped by the simulator program onto registers and / or memory provided by the host apparatus. Such a simulation program can be useful, for example, when legacy code written for one instruction set architecture is being executed on a host processor which supports a different instruction set architecture. Also, the simulation can allow software development for a newer version of the instruction set architecture to start before processing hardware supporting that new architecture version is ready, as the execution of the software on the simulated execution environment can enable testing of the software in parallel with ongoing development of the hardware devices supporting the new architecture.

[0061] In a simulated example, the TLBI control program logic determines, based on the TLBI domain information specified by the TLBI instruction decoded by the instruction decoding program logic, one or more simulated TLBs in one or more TLBI domains which should observe the TLB invalidation event. In some examples, this may include simulation of TLBI request signals mirroring the TLBI request signals determined by the TLBI broadcast control circuitry for a hardware-implemented embodiment. For example, this may be useful if the simulation is a cycle- accurate simulation of a specific processing system implementation. However, in other examples of simulation computer programs, it may not be necessary to simulate the actual physical system design, so that the simulation does not actually determine any specific TLBI request signals mirroring those that would be generated in hardware. Hence in some cases the TLBI control program logic may simply identify certain simulated TLBs as being required to observe the invalidation event, without necessarily simulating the TLBI broadcast request signals which would be issued in a hardware-implemented example.

[0062] The simulation program may be stored on a storage medium, which may be a non- transitory storage medium.

[0063] Specific examples are now set out with reference to the drawings.

[0064] Figure 1 illustrates an example of a data processing system 2 comprising a number of processing elements (PEs) 4. The processing elements 4 may be processor cores, such as CPUs (central processing units) or GPUs (graphics processing units). Each processing element 4 may have at least one translation lookaside buffer (TLB) 6 for caching address translation information derived from translation table structures (page tables) stored in a memory system. The processing elements 4 communicate with each other and with memory storage 12 via a system interconnect 10. The memory system may comprise memory storage 12, as well as one or more caches (not shown in Figure 1 ) that may be associated with specific processing elements 4 or with the interconnect 10. The processing elements 4 are capable of execution of program instructions defined according to a particular instruction set architecture (ISA).

[0065] The system 2 may also include a number of devices 16 which do not themselves execute program instructions according to the ISA, but which share access to the memory 12 accessed by the processing elements. For example, the devices 16 may include one or more hardware accelerators providing certain specialised processing functionality (e.g. direct memory access control, cryptographic processing, or matrix / machine learning processing), which can be configured by software executing on a processing element 4 to carry out certain functions on behalf of that processing element 4 in parallel with continued execution of instructions on the processing element 4. The devices 16 could also include one or more input / output (I / O) devices for handling interaction between the system 2 and external devices and users. For example, I / O devices may include, as a non-limiting set of examples, one or more of a network controller, radio communication unit, user input / output interface device, and / or port for communicating with external data storage circuitry. The devices 16 may not have their own address translation mechanisms, and so to allow those devices to have their memory accesses translated onto the same physical address space used to address the memory 12 based on the address translation mappings set in translation table structures specified by software executing on the processing elements 4, at least one input / output memory management unit (IOMMU, also known as a system memory management unit or SMMU) 14 may be provided. Each IOMMU 14 performs address translations and other memory access control functions on behalf of a set of one or more associated devices 16. The IOMMU 14 may have one or more TLBs 6 for caching address translation information derived from the translation table structures, in a similar way to the TLBs 6 in the processing elements 4.

[0066] While Figure 1 shows, for sake of example, three PEs 4 and two lOMMUs 14, it will be appreciated that in practice a system may comprise many processing elements (e.g. 10s, 100s or greater) and may have more than two lOMMUs 14, so may have many associated TLBs 6. Also, while each PE 4 or IOMMU 14 is shown as having a single TLB 6 in the example of Figure 1 , some PEs 4 or IOMMU 14 may be associated with more than one TLB 6.

[0067] As shown in Figure 1 , to support more fine-grained restriction of the scope of a TLB invalidation operation, a number of TLBI domains 20 may be defined, each comprising one or more TLBs 6. The TLBI domains 20 are non-overlapping domains, so that each TLB 6 is a member of only one TLBI domain 20. The particular allocation of TLBs 6 to TLBI domains 20 may be chosen by the system designer depending on the particular system design. For example, the system designer may aim to group TLBs which are nearer to each other (or which are expected to have similar communication latencies between the TLBs and a given processing element 4) into the same TLBI domain 20, while TLBs far apart on an integrated circuit, on separate chiplets, and / or with significantly different expected communication latencies relative to a given reference point may be separated into separate TLBI domains 20. As discussed further below, a TLBI instruction may be supported which takes, as an input parameter, TLBI domain information identifying whether each TLBI domain 20 is a target TLBI domain which should observe a TLB invalidation event requested by that TLBI instruction or a non-target TLBI domain which is not required to observe that TLB invalidation event. This allows the extent to which TLBI request signals need to be broadcast across the entire system 2 to be limited based on which portions of the system have been indicated by software as being relevant to that TLB invalidation event. Figure 2 shows in more detail circuitry provided within a given processing element 4, which may be considered an example of an apparatus comprising instruction decoding circuitry 30 and TLBI broadcast control circuitry 40. The processing element 4 comprises instruction decoding circuitry 30 for decoding program instructions according to the ISA supported by the processing element 30. The instruction decoding circuitry receives a given instruction fetched from the memory system 12 and maps it to one or more micro-operations (decoded instructions) which are passed to processing circuitry 32 for processing. The processing circuitry 32 has various execution units 34, 36, 38, 40 for performing processing operations corresponding to various classes of program instruction decoded by the instruction decoding circuitry 30. For example, the execution units may include one or more arithmetic / logic units (ALUs) 34 for processing arithmetic / logical operations in response to arithmetic / logical instructions, a branch processing unit 36 for processing branch operations in response to branch instructions, a load / store unit 38 for performing load / store operations to access the memory system in response to load / store instructions, and a TLBI broadcast control unit (TLBI broadcast control circuitry) 40 for controlling broadcasting of TLBI requests in response to TLBI instructions. The processing element 4 has a set of registers 42, for storing a set of architectural state structured according to the ISA supported by the instruction decoding circuitry 30 and processing circuitry 32. The registers 42 may include general purpose registers 44 for storing general purpose operands and results of processing operations, and control registers 46 for storing architecturally-defined control state information used to control the operation of the processing circuitry 32. The processing element 4 also has a memory management unit (MMU) 48 for performing address translation and memory access control functions in response to the load / store operations processed by the load / store unit 38, based on address translation information and / or access permissions or attributes defined using translation table structures obtained from the memory system 12. At least one TLB 6 associated with the MMU 48 caches information derived from such translation table structures for faster access than if the information had to be obtained from memory 12 every time. While Figure 2 shows a single MMU 48 associated with processing of load / store operations performed by the load / store unit 38, it will be appreciated that addresses of instruction fetches may also require address translation and memory management functions, so instruction fetch addresses could either be translated using the same MMU 48 as the load / store addresses used for data accesses, or could be translated using a separate instruction-side MMU, separate from the data-side MMU 48 shown in Figure 2. Hence, in some cases the instruction-side MMU may include a further TLB separate from the TLB 6 of the data-side MMU.

[0068] The TLBI broadcast control circuitry 40 is responsible for processing of TLB invalidation operations triggered by TLBI instructions decoded by instruction decoding circuitry 30. In response to a TLBI instruction being decoded by the instruction decoding circuitry 30, the TLBI broadcast control circuitry 40 determines the scope of the TLB invalidation event required by that TLBI instruction, and then, as well as triggering any required invalidation of TLB entries from the I LBs 6 ot the processing element 4 itself, also generates one or more TLBI broadcast request signals 49 for requesting that other TLBs 6 in other parts of the processing system 2 also observe the TLB invalidation event. TLBI tracking circuitry 50 (described in more detail later) is provided to track acknowledgement of requested TLBI operations by other TLBs 6 or components of the system, to determine when subsequent operations processed by the processing circuitry 32 can be executed (those operations should not proceed until it is guaranteed that any stale address translation information can no longer be obtained from the caching structures of any TLBs 6 that may have cached that information).

[0069] Figure 3 shows an example encoding for a TLBI instruction, which specifies an invalidation event type 60, shareability domain information 62, and a register operand 64 (denoted by Xt in this example). The instruction encoding 66 for the TLBI instruction includes a number of opcode bits 67 and a register field 68 which identifies which of the general purpose registers 44 is the register specifying the register operand 64. The opcode bits 67 identify that this instruction is a TLBI instruction (distinguished from other types of instructions that do not trigger TLB invalidations). In this example, the opcode bits 67 also identify the invalidation event type 60 and shareability domain information 64. In other examples, the invalidation event type 60 and / or shareability domain information 62 could be encoded in the register operand 64. The register operand 64 (i.e. the value stored in the register referenced in register field 68) in this example specifies invalidation condition specifying information 71 and TLBI domain information 72 (NTLBID), with remaining bits 73 of the register operand 64 in this example being reserved bits assumed by default to be set to a default value of 0. In other examples, additional information could also be specified in the register operand 64, with some of the reserved bits 73 being used for other purposes. The invalidation event type information 60 and the invalidation condition specifying information 71 together act as information defining the conditions to be satisfied by a given TLB entry in order for that TLB entry to be invalidated in response to the TLBI instruction.

[0070] The invalidation event type 60 distinguishes what type of invalidation event is required. For example, different types of invalidation event may be associated with different filter criteria for determining which entries of a given TLB should be invalidated when the TLBI instruction causes an invalidation request to be sent to the given TLB. For example, supported types of invalidation event could include:

[0071] - invalidate all: invalidate all entries of the given TLB regardless of the properties of each entry;

[0072] - invalidate by address or address range: invalidate entries of the given TLB which correspond to an address corresponding to an address or address range defined by the invalidation condition specifying information 71 of the register operand 64; invalidate by context identifier: invalidate entries of the given TLB which correspond to a given address translation context defined by the invalidation condition specifying information 71 of the register operand 64. Some examples may support more than one type of address translation context identifier. For example, in a system supporting two stages of address translation, with stage 1 based on stage-1 translation tables identified by an address space identifier (ASID) and stage 2 based on stage-2 translation tables identified by a virtual machine identifier (VMID), invalidations by context identifier could be based on the ASID or the VMID associated with the TLB entry, or a combination of both identifiers.

[0073] Some example invalidation event types may combine different criteria (e.g. determining entries to be invalidated based on the combination of a context identifier and address / address range information). It will be appreciated that the list of invalidation event type shown above is not exhaustive and other types of invalidation event could also be defined.

[0074] Hence, in general the invalidation event type 60 and invalidation condition specifying information 71 controls how a given TLB will respond to an invalidation request once that given TLB has received the invalidation request. On the other hand, the shareability domain information 62 and TLBI domain information 72 may be for controlling which TLBs are sent invalidation requests.

[0075] The shareability domain information 62 defines a selected shareability domain, as one of a non-shareable (NSH) domain, an inner shareable (ISH) domain or an outer shareable (OSH) domain. Figure 4 illustrates the concept of shareability domains for a given system example comprising a number of processing elements 4 and lOMMUs 14 similar to Figure 1. The shareability domains NSH, ISH, OSH form a nested set of concentric domains defines relative to the processing element 4 that executes the TLBI instruction. The NSH domain refers to a domain comprising that processing element 4 only. The ISH domain is a wider domain that also includes a number of other processing elements 4 or lOMMUs 14 that are in the local vicinity of the processing element 4 that executed the TLBI instruction (and so are expected to have shorter communication latencies when TLB invalidation events are instructed by that processing element 4, in comparison to other processing elements 4 and lOMMUs 14 in other more remote parts of the system). The OSH domain encompasses the entire system 2 including all processing elements 4 and lOMMUs 14.

[0076] For example, when a TLBI instruction is executed on a first processor (PEx) 4-x, the corresponding non-shareable domain, NSHx, for that processor comprises that processor 4-x itself; the corresponding inner-shareable domain, ISHx, comprises processor 4-x and a number of other local processors 4 and / or IOMMU 14, and the outer shareable domain OSH is the entire system. In contrast, if the TLBI instruction is executed on the second processor (PEy) 4-y, the non-shareable domain NSHy and inner-shareable domain ISHy are different to the corresponding non-shareable and inner-shareable domains NSHx, ISHx identified relative to processor 4-x. As the shareability domains are concentric overlapping domains, and it is not possible to distinguish different portions of the OSH domain that are outside a given processor’s ISH domain, this approach is harder to scale to the many-core and multi-ch iplet systems for which more detailed targeting ot various groups of TLBs remote from the processor executing the TLBI instruction may be desirable to reduce latencies associated with completion of TLB invalidation events.

[0077] In contrast, the TLBI domain information 72 is information for defining which TLBI domains 20 are required to observe the TLB invalidation event. As shown in Figure 4, the TLBI domains are defined as absolute domains, rather than relative domains like the shareability domains. Hence, the identification of which TLB 6 belongs to a particular TLBI domain having a given TLBI domain number is consistent across the system 2 regardless of which processor executes the TLBI instruction. In the example of Figure 4, say, four TLBI domains 0, 1 , 2, 3 are shown, and the TLBs 6 in processor 4-x belong to TLBID 0 regardless of whether a TLBI instruction references TLBI domain identifier 0 when executed on processor PEx or processor PEy. This approach can be much more scalable to different physical system designs having varying numbers of processor cores and varying numbers of chiplets.

[0078] Hence, the TLBI broadcast control circuitry 40 uses the TLBI domain information 72 of the TLBI instruction to determine how to generate the TLBI broadcast signals 49 in response to that instruction. The TLBI broadcast control circuitry 40 can limit, based on the TLBI domain information 72 specified by the register operand 46 of the TLBI instruction, which TLBI domains 20 are instructed to observe the TLB invalidation event.

[0079] It will be appreciated that while Figure 4 shows an example with multiple ISH domains, this is not essential and in other systems all TLBs may be within the same ISH domain.

[0080] Figure 3 shows an example encoding of the TLBI domain information (NTLBID field) 72 in the register operand 64. The TLBI domain information 72 includes a number of TLBI domain indicators 74, each associated with a corresponding one of the TLBI domains 20 and indicating whether that TLBI domain is a target TLBI domain or a non-target TLBI domain. In this example, the following encoding of each TLBI domain indicator 74 is used:

[0081] NTLBI D[i] = 0: if the shareability domain information 62 specifies the ISH domain, the TLBI domain with TLBI domain identifier i is a target domain required to observe the TLB invalidation event triggered by this TLBI instruction;

[0082] NTLBI D[i] = 1 : if the shareability domain information 62 specifies the ISH domain, the TLBI domain with TLBI domain identifier i is a non-target domain which is not required to observe the TLB invalidation event triggered by this TLBI instruction.

[0083] This encoding can be useful for backwards compatibility with legacy code, because when the reserved bits 73 of the register operand 64 of the TLBI instruction are set to zero by default, and the TLBI domain information 72 has been encoded using a portion of the register operand 64 previously regarded as reserved bits, it may be desirable that the default behaviour for TLBI instructions targeting the ISH shareability domain may remain the same as prior to introducing support for the TLBI domain information 72 when the NTLBID field 72 has the previously reserved value of all 0s. Therefore, it is preferable for all TLBI domains to be considered target TLBI domains it the software has specified all Os for the NTLBID field 72. Nevertheless, other encodings are also possible.

[0084] Figure 3 shows an example encoding where the NLBID field 72 specifies a bit mask of TLBI domain indicators, each corresponding to one TLBI number in the TLBI number space. However, another encoding could be that the NTLBID field 72 specifies an integer identifying the TLBI number of a particular TLBI domain to be regarded as a target TLBI domain (or if sufficient encoding space is available, there could be multiple such fields to allow two or more TLBI domains to be identified as target TLBI domains). Some implementations could support only one of these encodings of the NLBID field (either a bit mask of TLBI domain indicators as shown in Figure 3, or an integer-based identification of TLBI domain numbers), while other implementations could support one instruction variant with the bit mask based encoding of the TLBI domain information as shown in Figure 3 and another instruction variant with the integer-based encoding of the TLBI domain information. If both variants are supported, these could be distinguished based on the opcode 67, part of the register operand 64, another parameter of the instruction and / or a modal control bit stored in a control register. Also, in some implementations, the ISA may not prescribe which format is used for the NLBID field 72, but the hardware of a particular system implementation may interpret that field according to one encoding or another depending on the choices of that particular hardware designer (in that case, identification information provided in a control register or in a firmware table or other stored control structure may be queried by software to determine which encoding is used by that particular hardware implementation, so that software can know how to set the NLBID field 72 to indicate the desired combination of target TLBI domains).

[0085] In one example, if the shareability domain information 62 specifies the OSH shareability domain, the TLBI domain information 72 is ignored and the TLBI broadcast control circuitry 40 generates the TLBI broadcast signals 49 to indicate that all TLBs 6 in the OSH domain should observe the TLB invalidation event. In other examples, use of TLBI domains to limit the applicability of TLBI events could also apply to TLBI instructions which specify the OSH shareability domain.

[0086] If the shareability domain information 62 specifies the NSH shareability domain, the TLBI domain information 72 is also ignored, and the TLBI broadcast control circuitry 40 generates signals to trigger invalidation by the local TLB 6 associated with MMUs 48 of that particular processing element 48, but does not broadcast the TLBI signals to TLBs outside the NSH domain. Hence, the TLBI domain information 72 may, in some examples, only be considered for TLBI instructions which specify shareability domain information 62 identifying the ISH shareability domain.

[0087] As mentioned earlier, the particular way in which TLBI domains assignments are controlled, to map TLBs onto particular TLBI domains 20, may be highly system dependent and can be selected by the system designer of a particular processing system. Hence, there may be no architectural state defined in the ISA which mandates a particular way in which TLBI domain assignments are controlled. The provider of a particular hardware system implementations may provide firmware tables accessible to software, which indicate for particular processor cores 4 or lOMMUs 14 the TLBI domain identifier associated with that core 4 or IOMMU 14, so that operating system or hypervisor software may read the firmware tables and use the information specified in the firmware tables to make decisions on how to set the TLBI domain information 72 used for particular TLBI instructions.

[0088] Regardless of whether the TLBI domain information 72 identifies target TLBI domains using a bit mask of indicators 74 as in Figure 3 or using one or more integer fields identifying TLBI domain numbers of target TLBI domains, the TLBI domain information can be regarded as identifying the TLBI domains using a given TLBI domain number space. In some system implementations, it is possible to support TLBI domain aliasing, where a certain portion of the TLBI domain number space does not correspond to actual TLBI domains, but instead can be used to identify, as target TLBI domains, combinations of two or more actual TLBI domains having TLBI domain numbers in another portion of the TLBI domain number space. As an illustrative example, one particular system designer may choose that, in a system that has actual domains 0 to 15, TLBI number 16 is used for an aliased TLBI domain that maps to actual domains 0 to 15 (each of which are to be regarded as target TLBI domains when TLBI number 16 is indicated as being a target TLBI domain). Similarly, TLBI number 17 could map to actual domains 0 to 7; TLBI number 18 could map to actual domains 8 to 15; and TLBI domain 123 (say) could map to actual domains 1 , 3 and 23. The particular aliasing between TLBI numbers and actual domains is a property of a particular system implementation, depending on design choices of the system designer, and can be advertised to software in firmware tables as described above. The ISA can be agnostic to the particular way in which hardware designers choose to interpret the TLBI numbers defined in the TLBI number space supported by the ISA.

[0089] Hence, in some examples, if one of the aliased TLBI numbers is identified as corresponding to a target TLBI domain by the TLBI domain information (e.g. the TLBI domain information identifies that TLBI number 17 is a target domain in the example above), then the TLBI broadcast control circuitry 40 can map that aliased TLBI domain to a set of actual TLBI domains to be regarded as target TLBI domains (e.g. to actual domains 0 to 7 in the example above).

[0090] Such aliasing could be helpful in some examples, particularly if integer-based identification of a target TLBI domain using the TLBI domain information 72 is supported, to allow more flexible identification of particular combinations of TLBI domains as being target TLBI domains, and simplify software setting of the TLBI domain information 72.

[0091] Figures 5 and 6 illustrate, purely by way of example, possible TLBI domain mappings that could be considered for particular system examples. For example, Figure 5 shows an example of a system implemented using four compute chiplets 80, each compute chiplet comprising one or more processing elements 4 for processing instructions according to the ISA supporting the TLBI instruction. Each chiplet 80 is a separate integrated circuit implemented on a separate piece of silicon or other semiconductor substrate. Inter-chiplet bridges 82 may be provided for communication between chiplets 80. In general, communication latency may be greater between components on different chiplets 80, than between components on the same chiplet 80. Recognising this, it may be that the system designer decides to allocate a separate TLBI domain to each chiplet 80, so that all TLBs within processing elements 4 on the same chiplet 80 may be regarded as belonging to the same TLBI domain, but TLBs 6 in processing elements 4 on different compute chiplets 80 may be assigned to different TLBI domains. The software can then use its knowledge of which compute chiplets 80 have been scheduled a given processing workload, together with the firmware tables defining the mapping of TLBI domains for this system implementation, to decide how to set the TLBI domain information which limits which TLBs 6 are required to observe a TLB invalidation event. Some of the compute chiplets 80 in the example of Figure 5 may also include an IOMMU 14. For compute chiplets 80 comprising an IOMMU 14, TLBs 6 in the IOMMU 14 could be assigned to the same TLBI domain as the TLBs 6 of corresponding processing elements 4 on that chiplet 80 (e.g. TLBI domain 3 can be assigned to the IOMMU 14 on the compute chiplet 80 whose processing elements 4 are mapped to TLBI domain 3), or alternatively the IOMMU 14 could be assigned a different TLBI domain compared to the TLBs 6 of the processing elements 4 on that chiplet 80 (e.g. TLBI domain 4 can be assigned to the IOMMU 14 on the compute chiplet 80 whose processing elements 4 are mapped to TLBI domain 3).

[0092] Figure 6 illustrates another example of a possible system topology. In this example, a number of compute chiplets 80 are connected by respective inter-chiplet bridges 82 to an IO (input / output) hub chiplet 84 which manages communication to devices 16 (shown earlier in Figure 1 , but omitted from Figure 6 for conciseness) under control of an IOMMU 14 (the IO hub in this example also manages inter-compute-chiplet communication, as there is no direct bridge 82 from one compute chiplet 80 to another). Hence, in this case the layout of the processing system 2 may lend itself to defining each compute chiplet 80 as a separate TLBI domain and the IOMMU 14 in the IO hub chiplet as a further separate TLBI domain, rather than sharing a TLBI domain between IOMMU 14 and compute chiplet 80.

[0093] It will be appreciated that these are just some examples, but they illustrate that the particular mapping of TLBI domains may depend on the physical system layout.

[0094] Figures 5 and 6 also demonstrate how it can be helpful to allow software to give a hint to the hardware on which TLBs 6 in the system are required to observe a given TLB invalidation event. Considering the example of Figure 5, for example, without support for TLBI domains, if it can be distinguished that, for a TLBI instruction executed on a processing element 4 in TLBI domain 0, say, it is not necessary for TLBs in TLBI domain 3 to observe the invalidation, this can avoid the latency cost of two crossings of chiplet-to-chiplet bridges 82, so it is likely that completion of the TLB invalidation process can be faster and therefore subsequent operations requiring a guarantee that the TLB invalidation is complete (or will be completed) may be performed earlier, improving performance. This would be harder to achieve in a system supporting only relative shareability domains, especially as the number of chiplets and / or processing elements increases.

[0095] Figure 7 illustrates steps for processing a TLBI instruction. At step 100, the instruction decoding circuitry 30 decodes the TLBI instruction. The TLBI instruction specifies (e.g. using the register operand 64) TLBI domain information which identifies, for two or more TLBI domains defined as absolute non-overlapping TLBI domains, whether that TLBI domain is a target TLBI domain required to observe the TLBI invalidation event or a non-target TLBI domain which is not required to observe the TLBI invalidation event. In response to the decoding of the TLBI instruction, at step 102 the TLBI broadcast control circuitry 40 determines one or more TLBI request signals 49 based on the TLBI information specified by the TLBI instruction, to limit which TLBI domains 20 are instructed to observe the TLBI invalidation event based on the TLBI domain information.

[0096] It is not essential that the TLBI domains indicated by the TLBI request signals 49 are necessarily the same as the target TLBI domains indicated by the TLBI information 72 specified by the TLBI instruction. For example, as shown in the flowchart of Figure 8, to support virtualisation where a hypervisor running on the processing system 2 may virtualise physical resources of the system 2 so that an operating system managed by a hypervisor can see a different view of the system than is actually provided in hardware, it can be useful to provide support for virtualized remapping of target TLBI domains.

[0097] Figure 8 illustrates steps performed by the TLBI broadcast control circuitry 40 at step 102 of Figure 7 in an implementation supporting such virtualized remapping. At step 120 of Figure 8, the TLBI broadcast control circuitry 40 determines whether the processing element 4 executing the TLBI instruction is currently in at least one operating state for which virtualized remapping of target TLBI domains is enabled. As discussed further below with reference to Figure 12, for some implementations, virtualized remapping might be supported only for a subset of possible operating states of the processing circuitry. Also, in some examples, whether virtualized remapping is enabled for a given operating state may be controlled based on an enable control parameter stored in a control register 46.

[0098] If the current operating state of the processing element 4 that executes the TLBI instruction is not an operating state for which virtualised remapping is supported, or if an enable control parameter is set to indicate that virtualised remapping is currently disabled, then at step 122 the TLBI broadcast control circuitry 40 generates TLBI request signals 49 specifying that the TLB invalidation event is to be observed by at least the target TLBI domains specified by the TLBI domain information 72. Optionally, the TLBI request signals could also specify that, or imply that, one or more non-target TLBI domains can also observe the TLB invalidation event. When virtualised remapping is currently disabled or is not supported for the current operating states, the target TLBI domains indicated by the TLBI domain information 72 may (at least if corresponding to actual TLBI domains in a system supporting TLBI domain aliasing) be regarded as physical TLBI domains corresponding to the actual TLBI domains 20 designated for the system 2. As noted above, in some systems, the specific hardware implementation may use TLBI domain aliasing and so a given aliased TLBI domain number identified as a target TLBI domain may be mapped to a set of multiple actual TLBI domains to be regarded as target TLBI domains. This aliased mapping is a specific function of the TLBI broadcast control circuitry 40 in a given system implementation, and may take place even if virtualized remapping is currently disabled or not supported. Such aliasing corresponds to hardware-implemented remapping of TLBI domains to account for specifics of a given system implementation, which is different to the virtualized remapping which is based on architecturally-defined control state information programmed by the software running on the processing system 2.

[0099] On the other hand, if the current operating state of the processing element 4 that executes the TLBI instruction is an operating state for which virtualised remapping of target TLBI domains is enabled, then at step 124, the TLBI broadcast control circuitry 40 maps, based on target domain remapping information 156 (see Figure 9), a set of virtual target TLBI domains specified by the TLBI domain information 72 of the TLBI instruction to a set of physical target TLBI domains which correspond to the actual TLBI domains 20 (as designated for the physical system 2) which are to observe the TLB invalidation event. At step 126, the TLBI broadcast control circuitry 40 generates the TLBI request signals 49 specifying that the TLBI invalidation event is to be observed by at least the set of physical target TLBI domains identified based on the target domain remapping information 156. Again, optionally some of the physical non-target TLBI domains, which are not designated as being required to observe the TLB invalidation event based on the target domain remapping information 156, can nevertheless be among the TLBI domains instructed to observe the TLB invalidation event by the TLBI request signals 49. If the hardware supports aliased TLBI domains, then such aliasing is applied to the physical non-target TLBI domains identified based on the virtualized remapping. Hence, if one of the physical target TLBI domains identified using the target domain remapping information 156 corresponds to an aliased TLBI domain number, the TLBI broadcast control circuitry 40 can map that to a set of two or more actual TLBI domain numbers identifying actual TLBI domains for which the corresponding TLBs are required to observe the TLB invalidation event.

[0100] Figure 9 shows an example of control state information that may be stored in the control registers 46 for controlling virtualized remapping of TLBI domains. In this example, the control information 150, 152, 154, 156 is shown as designated within particular control registers, but it will be appreciated that the same information could be stored in a different format or could be grouped into registers in a different manner. I he processing elements 4 may support execution of program instructions at one of a number of exception levels (EL) associated with different levels of privilege. In this example, the exception levels supported include: ELO with application-level privilege (the least privileged operating state), EL1 with operating-system-level privilege, EL2 with hypervisor-level privilege, and EL3 intended for supervisory code with the greatest level of privilege. Hence, exception levels with successively increasing exception level numbers are associated with progressively increasing privileges (such that a more privileged exception level may be associated with rights for software executed in that exception level to access some information or carry out certain operations that are not available to a less privileged exception level). Transitions between exception levels may occur based on exceptions and exception returns, as well as triggered by certain instructions such as supervisor call instructions. A control register denoted with a suffix _ELx is a register for which writes to that register are restricted so that the least privileged exception level in which instructions are allowed to write to that register is exception level ELx. A register with suffix _EL2, for example, is a register restricted to be updated by instructions associated with the hypervisor-level privilege or higher.

[0101] In this example, the control state information includes:

[0102] TLBI domain identification information (TLBID) 150, which in this example is provided within an identification register for which access is restricted to instructions executing in EL1 , EL2 or EL3. The TLBID 150 identifies how many TLBI domains are supported by the system 2. The TLBID 150 field may correspond to reserved bits of the identification register which in legacy systems would be set to zero to indicate that the number of supported TLBI domains is zero. The identification information 150 can be useful to allow software to query the number of TLBI domains supported in a given hardware implementation, so that software can avoid the overhead of maintaining target domain remapping information 156 corresponding to unsupported TLBI domains, for example. Accesses to the identification register could, in some examples, be trapped and emulated by EL2, such that EL2 gives an illusion that the hardware supports fewer domains than advertised in the real ID register. For example, the hardware may support 16 domains but the hypervisor running at EL2 could decide it will only provide the illusion of 4 domains to EL1 (kernel), so when EL1 attempts to read the register, this is trapped to EL2 and EL2 emulates the read response as indicating “4 domains are supported” instead of the real hardware value of “16 domains are supported. virtualized remapping enable control information (VTLBIDEn) 152, which specifies whether, when the current operating state is exception level EL1 , the virtualised remapping discussed with respect to Figure 8 should be regarded as enabled. VTLBIDEn 152 is specified in a hypervisor control register HCRX EL2 for which updates are restricted to instructions executed in exception level EL2 or EL3. I rap control information 154 (VTLBIDEn) provided in a system control register SCR_EL3, for controlling whether hypervisor-triggered requests to update the target domain remapping information 156 should cause an exception to be raised to trap processing to EL3. While often it may be acceptable for instructions at EL2 to write to target domain remapping information 156 directly, some processing systems might include a highly- secure system region (e.g. a dedicated processor for handling cryptographic or other sensitive operations) which might include a TLB 6 for which information about whether that TLB 6 is relevant to a given workload may expose side channel information which it might be undesirable to expose to the hypervisor executing at EL2. In that case, the trap control information 154 might be set by secure software executing at EL3 to indicate that any attempt by the hypervisor at EL2 to set the target domain remapping information 156 should cause an exception to be raised so that software at EL3 can decide how to update the target domain remapping information 156 based on information specified by hypervisor, including making any adjustments for the TLBI domain associated with the secure system region.

[0103] - Target domain remapping information 156, in this example provided in a set of n control registers VTLBID<0>_EL2 to VTLBID<n-1 >_EL2 which are restricted to being updated by instructions executed at EL2 or EL3 (subject to the trap control implemented using trap control information 154). The target domain remapping information 156 is implemented across multiple registers in this example because the target domain remapping information 156 is too large to specify in a single register. In general, the target domain remapping information 156 specifies how to map a given set of virtual target TLBI domains (e.g. identified by indicators 74 set to 0 in the TLBI domain information 72 specified by a given TLBI instruction) to a set of physical target TLBI domains for which the TLBI broadcast control circuitry 40 should generate TLBI broadcast request signals 49 instructing that TLBs 6 in those physical target TLBI domains should observe the TLB invalidation event.

[0104] Figure 10 illustrates an example format for the target domain remapping information 156. Each remapping register corresponding to the target domain remapping information 156 specifies a number of remapping fields 158 each corresponding to one specified TLBI domain. The example of Figure 10 supports a maximum of 16 TLBI domains, with each remapping register specifying remapping fields 158 for four of the TLBI domains. Depending on the number of TLBI domains supported and the size of each remapping field 158 compared to the number of bits per register, other examples may use a different number of remapping registers and / or different number of fields 158 per register.

[0105] The remapping field 158 corresponding to a given virtual TLBI domain comprises a set of physical TLBI domain indicators 159, each corresponding to a given physical TLBI domain implemented in the system 2. The physical TLBI domain indicator 159 for a given physical TLBI domain indicates whether, when the given virtual TLBI domain associated with that remapping field 158 is indicated as a virtual target TLBI domain (by setting the corresponding indicator 74 of the TLBI domain information to the value (e.g. 0) indicating a target TLBI domain), that given physical TLBI domain should be regarded as a target TLBI domain for which the TLBI broadcast control circuitry 40 should instruct TLBs in that domain to observe the TLB invalidation event.

[0106] In this example, an encoding of field 158 is used in which, when the indicator 159 for a given physical TLBI domain has a first value (e.g. 1 ), the corresponding physical TLBI domain is regarded as a target TLBI domain when the virtual TLBI domain associated with that remapping field 158 is also indicated as a target TLBI domain, and when the indicator 159 for a given physical TLBI domain has a second value (e.g. 0), the corresponding physical TLBI domain does not necessarily need to be regarded as a target TLBI domain (although that physical TLBI domain could still be a target TLBI domain if the corresponding indicator 159 is set to the first value in one of the other remapping fields 158 which corresponds to a virtual target TLBI domain).

[0107] Figure 11 shows an example of a remapping function (implemented using hardware circuit logic of the TLBI broadcast control circuitry 40) which could be used to carry out the virtualised remapping of TLBI domains based on the target domain remapping information 156. The indicators 159 of each remapping field 158 are combined with the corresponding indicator 74 of the TLBI domain information so as to clear to 0 all the target domain indicators 159 of a given remapping field 158 if the corresponding indicator 74 indicates that the virtual TLBI domain corresponding to that remapping field 158 is not a target TLBI domain. For example, this combination may be based on a logical AND combination of each indicator 159 of that remapping field 158 with the inverse of the corresponding indicator 74 of TLBI domain information 72 indicating whether the virtual TLBI domain corresponding to that remapping field 158 is a target TLBI domain (for an encoding where a value of 0 for the indicator 74 indicates a target domain - for other encodings the Boolean function used to combine indicator 74 with the remapping field 158 may be different). Hence, each remapping field 158 is qualified based on the corresponding virtual TLBI domain indicator 74 from TLBI domain information. The qualified values of each remapping field 158 are combined in a Boolean OR operation to generate a final set of target domain indications 160 which indicate the set of physical TLBI domains which define the minimum set of TLBI domains 20 required to be instructed to observe the TLB invalidation event. It will be appreciated that, while Figure 1 1 shows one example of remapping circuitry, other examples could generate the same result in a different way.

[0108] For example, consider an example where the TLBI domain information 72 is set to 0b11 11_1 11 1_1101_1 110 (indicating that only TLBI domains 0 and 5 are virtual target TLBI domains and the other TLBI domains are virtual non-target TLBI domains). The bits set to 1 in the TLBI domain information 72 cause the outputs of the AND gates corresponding to the virtual non-target TLBI domains to all be set to 0, so that these outputs do not influence the selection of physical target TLBI domains 160. The selected physical target TLBI domains therefore depend only on the remapping fields 158 VLTBID<0>_EL2.TD[0] and VLTBID<1 >_EL2.TD[1] corresponding to virtual TLBI domain identifiers 0 and 5 respectively (see mapping between virtual TLBI domain identifiers and remapping fields 158 in Figure 10 for this example).

[0109] Assuming those remapping fields 158 have been set to the following values: VLTBID<0>_EL2.TD[0] = 0b0000_1101_0010_0001 VLTBID<1 >_EL2.TD[1] = 0b0000_0100_01 10_0101 the set of physical TLBI domains 160 will be indicated as the bitwise OR of VLTBID<0>_EL2.TD[0] and VLTBID<1>_EL2.TD[1], i.e.: 0b0000_1101_ 0110_0101 .

[0110] Hence, in this particular example the physical TLBI domains to be instructed with TLB invalidation requests would be at least physical TLBI domains 0, 2, 5, 6, 8, 10 and 11 , and it would not be necessary (although still possible) to instruct physical TLBI domains 1 , 3, 4, 7, 9 and 12- 15 to observe the TLB invalidation event. In this example, the virtualised remapping results in additional TLBI domains being required to observe the TLB invalidation event (compared to the number of virtual target TLBI domains indicated by the TLBI domain information 72), but in other examples, depending on the setting of the remapping fields 158, the remapping might result in the set of physical TLBI domains comprising fewer TLBI domains than the set of virtual TLBI domains.

[0111] Figures 10 and 1 1 show an example of the target domain remapping information 156 where each domain remapping field 158 specifies a bitmask of target domain indicators 159. However, similar to the alternative encoding for the TLBI domain information 72 mentioned earlier, it is also possible for each domain remapping field 158 to specify a physical target TLBI domain using a TLBI domain number expressed as an integer. For example, the domain remapping field 158 corresponding to a given virtual TLBI domain number X could store a given physical TLBI domain number Y, indicating that if virtual TLBI domain number X is selected as a target virtual TLBI domain, then physical TLBI domain number Y should be regarded as a target physical TLBI domain. In some system implementations, physical TLBI domain number Y could itself be subject to TLBI domain aliasing so as to be mapped by the hardware to a corresponding set of one or more actual TLBI domain numbers.

[0112] Hence, it will be appreciated that the particular format of the target domain remapping information 156 shown in Figures 10 and 11 is just one example, and other encodings can also be used.

[0113] Figure 12 illustrates a more detailed example for controlling processing of a TLBI instruction in an example supporting virtualized remapping of target TLBI domains. At step 200, the TLBI instruction is decoded by the instruction decoding circuitry 30 of a given processor 4. In response, at step 202, the TLBI broadcast control circuitry 40 determines which shareability domain is specified by the shareability domain information 62 of the TLBI instruction. If the shareability domain selected by the shareability domain information 62 is the non-shareable (NSH) domain, then at step 204 the TLBI broadcast control circuitry 40 determines that the TLBs 6 in the given processor’s non-shareable domain (determined relative to that processor 4) should observe the TLB invalidation event. If the shareability domain selected by the shareability domain information 62 is the outer shareable (OSH) domain, then at step 206 the TLBI broadcast control circuitry 40 determines that the TLBs 6 in the outer shareable domain should observe the TLB invalidation event. For steps 204, 206, the scope of the TLB invalidation does not depend on either the TLBI domain information 72 or the target domain remapping information 156.

[0114] If the shareability domain information 62 specifies the inner shareable (ISH) domain, then at step 208 the TLBI broadcast control circuitry 40 determines the current operating state of the given processor 4 that is executing the TLBI instruction. If the current operating state is EL3 (the most privileged state), then at step 210 the TLB invalidation is broadcast at least to all TLBs 6 in the given processor’s inner shareable (ISH) domain, instructing those TLBs to observe the TLB invalidation event.

[0115] If the current operating state is EL2 (a hypervisor-level operating state), then at step 212, the TLBI domains which should be instructed to observe the TLB invalidation include at least those target TLBI domains identified based on the TLBI domain information 72 specified by the instruction. At step 214, the TLBI broadcast control circuitry 40 determines that, additionally, even if not actually specified as target TLBI domain by the TLBI domain information 72, the local TLBI domain including the TLB(s) 6 associated with the given processor 4 that is executing the TLBI instruction should also be regarded as a target TLBI domain. This reduces the chance of unpredictable system operation which might occur as a consequence of a programming error where software mistakenly sets the TLBI domain indicator to 1 for its own TLBI domain, when it is very likely that actually the translation table information cached in that processor’s own TLBs 6 may be likely to be subject to the same requirement to be invalidated as TLBs in remote domains. In any case, even if it is actually correct that there is no need for the local TLBs 6 in the processor’s own local TLBI domain to observe the TLB invalidation, in practice sending requests for the local TLBI domain to invalidate in response to the TLBI instruction is unlikely to greatly increase latency associated with completing the TLBI invalidations, since the local TLBI domain has much lower communication round trip latency relative to the source of the TLBI invalidation instruction than other more remote TLBI domains. Hence, in some examples, it can be useful to assume, by default, that the local TLBI domain of the processor 4 executing the TLBI instruction should always be regarded as a target TLBI domain. In other examples, this requirement may not be imposed and step 214 could be omitted.

[0116] On the other hand, if at step 208 the current operating state was determined to be EL1 (an operating state with operating-system-level privilege), then at step 218 the TLBI broadcast control circuitry 40 determines, based on the virtualised remapping enable control information 152, whether virtualized remapping of target TLBI domains is currently enabled. If virtualized remapping of target TLBI domains is currently disabled (or is not supported at all for EL1 ), then the method proceeds to step 212 to generate the set of target TLBI domains in the same way as discussed above for EL2. If virtualized remapping of target TLBI domains is enabled, then at step 220 any virtual target TLBI domains specified by the TLBI domain information 72 are mapped to a set of physical target TLBI domains based on the target domain remapping information 156 (e.g. using a set of fields 158 which specify for each virtual TLBI domain which physical TLBI domains are required to be set as target domains if the corresponding virtual TLBI domain is identified as a target domain, as discussed above with respect to Figure 11 ). The method then proceeds to step 214 as explained above, to allow for the given physical processor’s own local TLBI domain to be set as an additional physical target TLBI domain if not already identified as a physical target TLBI domain.

[0117] Regardless of whether or not virtualized remapping of target domains is performed, at step 216, having identified a set of (physical) target TLBI domains, the TLBI broadcast control circuitry 40 determines the TLBI request signals 49 requesting that at least any physical target TLBI domain identified in the earlier steps 212 or 220 is required to observe the TLB invalidation event (in some system implementations, it is possible some of those target TLBI domains could be aliased TLBI domains which cause a corresponding set of one or more actual TLBI domains to be instructed to observe the TLB invalidation event). Optionally, the TLBI request signals 49 could also control other non-target TLBI domains, not identified as a target TLBI domain in steps 212, 220 to also observe the TLB invalidation event. Hence, the set of target TLBI domains is a minimum subset of TLBI domains required to observe the TLB invalidation, but there is flexibility for hardware to over-invalidate, and cause the TLB invalidation to be observed by other TLBI domains as well.

[0118] If at step 208 the current operating state was identified to be ELO (an operating state with application-level privilege), then a fault is signalled, since application-level code should not be managing translation tables and should not cause TLB invalidations.

[0119] Figure 13 illustrates an example of how different instances of the TLBI broadcast control circuitry 40 might respond differently when determining the TLBI request signals 49, even when faced with essentially similar settings for the TLBI domain information 72 and / or target domain remapping information. This illustrates an example where some system designers may choose to permit over-invalidation of entries in TLBs belonging to additional non-target TLBI domains, to simplify the circuit implementation.

[0120] As shown in the left-hand example of Figure 13, for a compute chiplet assigned TLBI domain 0, TLBI broadcast control circuitry 40 might determine that any TLBI instruction requiring only TLBI domain 0 as a target TLBI domain is restricted to being observed by TLBs 6 within that compute chiplet itself, but any TLBI instruction that specifies any one or more of TLBI domains 1 , 2 and 3 as target TLBI domains (hence requiring signals to leave that chiplet) may cause the TLBI event to be broadcast to all other domains, even if one of TLBI domains 1 , 2 and 3 is not actually indicated as a target TLBI domain. This may recognise that once the scope of the invalidation extends beyond the local chiplet, an extremely large delay is likely, and from a software point of view there is little difference whether that delay extends to one inter-chiplet communication round trips or multiple inter-chiplet communication round trips, so to reduce complexity in the formatting of TLBI request signals sent to other chips and reduce overhead in managing responses from other TLBI domains, it can be simpler to broadcast everywhere.

[0121] On the other hand, another chiplet implementation may use an approach as shown in the right hand example of Figure 13, where the TLBI request signals 49 sent from the compute chiplet associated with TLBI domain 2 may specify, in cases where it is necessary to broadcast the invalidation outside that chiplet because at least one of TLBI domains 0, 1 and 3 is specified as a target TLBI domain, target domain information identifying which specific other TLBI domains are target TLBI domains. The target domain information allows circuitry on another chiplet (e.g. the IO hub chiplet in this example) to determine whether or not to route those TLBI requests to other chiplets, allowing for elimination of some latency associated with certain inter-chiplet communication round trips if there are no target TLBI domains across the corresponding inter- chiplet bridges.

[0122] Hence, the particular approach taken may depend on the system designer’s goals and whether it is preferred to reduce circuit complexity or improve processing performance. The architectural definition of the TLBI instruction in the ISA can support either approach by defining the target TLBI domains as a minimum subset of TLBI domains to observe the invalidation.

[0123] As shown in Figure 2, the apparatus may also comprise TLBI tracking circuitry 50 for tracking completion of invalidation operations associated with a given TLBI instruction. Figure 14 illustrates steps for tracking completion of the invalidation. At step 230, the TLBI tracking circuitry 50 tracks, for a given processor 4, which of the TLBI domains 20 have outstanding TLB invalidation events which have not yet been acknowledged as being guaranteed to be completed. When TLBs 6 in a given TLBI domain 20 are instructed to observe the invalidation, then at a certain point of the processing flow for responding to the TLBI request (beyond which it is guaranteed that the invalidation operation will be completed and any later requests received after that point will not receive responses based on TLB entries that meet the invalidation conditions defining which TLB entries should be invalidated in the TLBs 6 within target TLBI domains), circuitry in that TLBI domain 20 may respond with an acknowledgement signal. It is not necessary to wait until the storage cache of a given TLB 6 has actually invalidated a given TLB entry before sending the acknowledgement signal, provided it is guaranteed that no later request can hit on the TLB entry awaiting invalidation and that all accesses to memory that were using the translation information that was invalidated by the TLBI operation have become observable to other observers. The TLBI tracking circuitry 50 may maintain a data structure for tracking outgoing TLBI requests and the corresponding acknowledgement signals, so that it can be identified whether there are any outstanding TLBI invalidation events at a given time. At step 232, in response to a data synchronisation barrier (DSB) instruction being executed at the given processor, the TLBI tracking circuitry 50 prevents any instruction that is younger than the DSB instruction (instructions which are at a later position than the DSB instruction in the program order of the instructions being executed) from being executed. At step 234, the TLBI tracking circuitry 50 determines whether all TLBI domains which were instructed to observe TLB invalidation events based on older TLBI instructions (instructions at an earlier position in program order than the DSB instruction and later in program order than the previous DSB instruction before the current DSB instruction) are guaranteed to be completed (having received the corresponding acknowledgement signal). If not, then the TLBI tracking circuitry 50 waits for such acknowledgement guarantees and continues to block younger instructions than the DSB instruction from being executed. Once it is determined that all TLBI domains which had been instructed to observe TLB invalidation events based on older TLBI instructions than the DSB instruction have guaranteed completion of their respective invalidation operations in response to those older TLBI instructions, then at step 236 the younger instructions than the DSB instruction can be permitted to execute. This approach helps maintain synchronisation of instructions relative to translation table updates to ensure that a given instruction does not use a stale copy of address translation information, while permitting grouping of multiple TLBI instructions relative to the barrier to avoid each TLBI instruction having to be separately acknowledged before permitting continued execution.

[0124] For the flowcharts discussed above, it will be appreciated that these show one possible sequence of operations that could be performed. However, some of the steps could be reordered to be performed in a different order to the order shown in the diagrams, and / or some steps could be performed at least partially in parallel.

[0125] Concepts described herein may be embodied in computer-readable code for fabrication of an apparatus that embodies the described concepts. For example, the computer-readable code can be used at one or more stages of a semiconductor design and fabrication process, including an electronic design automation (EDA) stage, to fabricate an integrated circuit comprising the apparatus embodying the concepts. The above computer-readable code may additionally or alternatively enable the definition, modelling, simulation, verification and / or testing of an apparatus embodying the concepts described herein.

[0126] For example, the computer-readable code for fabrication of an apparatus embodying the concepts described herein can be embodied in code defining a hardware description language (HDL) representation of the concepts. For example, the code may define a register-transfer-level (RTL) abstraction of one or more logic circuits for defining an apparatus embodying the concepts. The code may define a HDL representation of the one or more logic circuits embodying the apparatus in Verilog, SystemVerilog, Chisel, or VHDL (Very High-Speed Integrated Circuit Hardware Description Language) as well as intermediate representations such as FIRRTL. Computer-readable code may provide definitions embodying the concept using system-level modelling languages such as SystemC and SystemVerilog or other behavioural representations of the concepts that can be interpreted by a computer to enable simulation, functional and / or formal verification, and testing of the concepts. Additionally or alternatively, the computer-readable code may define a low-level description of integrated circuit components that embody concepts described herein, such as one or more netlists or integrated circuit layout definitions, including representations such as GDSIL The one or more netlists or other computer-readable representation of integrated circuit components may be generated by applying one or more logic synthesis processes to an RTL representation to generate definitions for use in fabrication of an apparatus embodying the invention. Alternatively or additionally, the one or more logic synthesis processes can generate from the computer-readable code a bitstream to be loaded into a field programmable gate array (FPGA) to configure the FPGA to embody the described concepts. The FPGA may be deployed for the purposes of verification and test of the concepts prior to fabrication in an integrated circuit or the FPGA may be deployed in a product directly.

[0127] The computer-readable code may comprise a mix of code representations for fabrication of an apparatus, for example including a mix of one or more of an RTL representation, a netlist representation, or another computer-readable definition to be used in a semiconductor design and fabrication process to fabricate an apparatus embodying the invention. Alternatively or additionally, the concept may be defined in a combination of a computer-readable definition to be used in a semiconductor design and fabrication process to fabricate an apparatus and computer- readable code defining instructions which are to be executed by the defined apparatus once fabricated.

[0128] Such computer-readable code can be disposed in any known transitory computer- readable medium (such as wired or wireless transmission of code over a network) or non- transitory computer-readable medium such as semiconductor, magnetic disk, or optical disc. An integrated circuit fabricated using the computer-readable code may comprise components such as one or more of a central processing unit, graphics processing unit, neural processing unit, digital signal processor or other components that individually or collectively embody the concept.

[0129] Figure 15 illustrates a simulator implementation that may be used. Whilst the earlier described embodiments implement the present invention in terms of apparatus and methods for operating specific processing hardware supporting the techniques concerned, it is also possible to provide an instruction execution environment in accordance with the embodiments described herein which is implemented through the use of a computer program. Such computer programs are often referred to as simulators, insofar as they provide a software based implementation of a hardware architecture. Varieties of simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on a host processor 730, optionally running a host operating system 720, supporting the simulator program 710. In some arrangements, there may be multiple layers of simulation between the hardware and the provided instruction execution environment, and / or multiple distinct instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations which execute at a reasonable speed, but such an approach may be justified in certain circumstances, such as when there is a desire to run code native to another processor for compatibility or re-use reasons. For example, the simulator implementation may provide an instruction execution environment with additional functionality which is not supported by the host processor hardware, or provide an instruction execution environment typically associated with a different hardware architecture. An overview of simulation is given in “Some Efficient Architecture Simulation Techniques”, Robert Bedichek, Winter 1990 USENIX Conference, Pages 53 - 63.

[0130] To the extent that embodiments have previously been described with reference to particular hardware constructs or features, in a simulated embodiment, equivalent functionality may be provided by suitable software constructs or features. For example, particular circuitry may be implemented in a simulated embodiment as computer program logic. Similarly, memory hardware, such as a register or cache, may be implemented in a simulated embodiment as a software data structure. In arrangements where one or more of the hardware elements referenced in the previously described embodiments are present on the host hardware (for example, host processor 730), some simulated embodiments may make use of the host hardware, where suitable.

[0131] The simulator program 710 may be stored on a computer-readable storage medium (which may be a non-transitory medium), and provides a program interface (instruction execution environment) to the target code 700 (which may include applications, operating systems and a hypervisor) which is the same as the interface of the hardware architecture being modelled by the simulator program 710. Thus, the program instructions of the target code 700 described above, may be executed from within the instruction execution environment using the simulator program 710, so that a host computer 730 which does not actually have the hardware features of the apparatus 2 discussed above can emulate these features.

[0132] For example, the simulator code 710 may comprise instruction decoding program logic 712 which simulates decoding and processing of instructions in an equivalent manner to the functionality offered by the instruction decoding circuitry 30 and processing circuitry 32 described above. The instruction decoding program logic 712 decodes instructions of the target code 700 and maps these to corresponding sets of instructions in the native instruction set of the host apparatus 730. TLB simulating program logic 716 simulates the behaviour of one or more simulated TLBs 6, e.g. allowing more efficient lookups of address translation information than if a multi-level translation table walk was required to traverse translation table structures every time an address translation is needed. TLBI control program logic 714 determines, in response to a TLBI instruction decoded by instruction decoding program logic 712, which simulated TLBs need to observe the invalidation, based on the corresponding TLBI domain information 72 specified by the TLBI instruction. Host storage mapping program logic 718 maps accesses to simulated registers or simulated memory (requested by the target code 700 according to the target ISA supported by the target code 700) onto host storage resources (e.g. registers and / or memory) provided in hardware in the host apparatus 730. For example, where an operation requested by the target code 700 requires access to a given register 42, the register access may be mapped onto a register simulating data structure maintained in host memory by the host storage mapping program logic 718 of the simulation program 710, while when an operation requested by the target code 700 requires an access to an address in simulated physical memory, this is mapped onto host virtual addresses in the host virtual address space by the host storage mapping program logic 718. These host virtual addresses may themselves be translated into host physical addresses using the address translation mechanisms supported by the host (the translation of host virtual addresses to host physical addresses being outside the scope of what is controlled by the simulator program 710).

[0133] In the present application, the words “configured to...” are used to mean that an element of an apparatus has a configuration able to carry out the defined operation. In this context, a “configuration” means an arrangement or manner of interconnection of hardware or software. For example, the apparatus may have dedicated hardware which provides the defined operation, or a processor or other processing device may be programmed to perform the function. “Configured to” does not imply that the apparatus element needs to be changed in any way in order to provide the defined operation.

[0134] In the present application, lists of features preceded with the phrase “at least one of” mean that any one or more of those features can be provided either individually or in combination. For example, “at least one of: [A], [B] and [C]” encompasses any of the following options: A alone (without B or C), B alone (without A or C), C alone (without A or B), A and B in combination (without C), A and C in combination (without B), B and C in combination (without A), or A, B and C in combination.

[0135] Although illustrative embodiments of the invention have been described in detail herein with reference to the accompanying drawings, it is to be understood that the invention is not limited to those precise embodiments, and that various changes and modifications can be effected therein by one skilled in the art without departing from the scope of the invention as defined by the appended claims.

Claims

CLAIMS1 . An apparatus comprising: instruction decoding circuitry to decode a translation lookaside buffer invalidation, TLBI, instruction for instructing at least one translation lookaside buffer, TLB, to observe a TLB invalidation event, the TLBI instruction specifying TLBI domain information indicating, for each TLBI domain of a plurality of TLBI domains, whether that TLBI domain is a target TLBI domain for which any TLBs in that TLBI domain are required to observe the TLB invalidation event or a nontarget TLBI domain for which any TLBs in that TLBI domain are not required to observe the TLB invalidation event, the plurality of TLBI domains including at least two non-overlapping TLBI domains; andTLBI broadcast control circuitry to determine, based on the TLBI domain information specified by the TLBI instruction decoded by the instruction decoding circuitry, one or more TLBI request signals to be issued for instructing TLBs in one or more TLBI domains to observe the TLB invalidation event; wherein the TLBI broadcast control circuitry is configured to limit which TLBI domains are instructed to observe the TLB invalidation event based on the TLBI domain information.

2. The apparatus according to claim 1 , in which the TLBI instruction specifies a register identifier indicating a register storing a register operand, and the TLBI domain information is encoded within the register operand.

3. The apparatus according to claim 2, in which: the TLB invalidation event comprises invalidation of TLB entries that satisfy at least one invalidation condition; and the at least one invalidation condition depends on invalidation condition specifying information specified in the register operand.

4. The apparatus according to any preceding claim, in which, in at least one operating state where virtualized remapping of target TLBI domains is enabled, the TLBI broadcast control circuitry is configured to determine the one or more TLBI request signals based on the TLBI domain information specified by the TLBI instruction and target domain remapping information specifying a mapping between a set of virtual target TLBI domains specified using the TLBI domain information and a set of physical target TLBI domains that are actually required to observe the TLB invalidation event.

5. The apparatus according to claim 4, in which:the target domain remapping information comprises a plurality of remapping fields, each remapping field associated with a corresponding virtual TLBI domain and specifying which of a plurality of physical TLBI domains are to be regarded as a physical target TLBI domain when the corresponding virtual TLBI domain is indicated as being a virtual target TLBI domain by the TLBI domain information; the set of physical target TLBI domains comprising any of the plurality of physical TLBI domains indicated as being a physical target TLBI domain by any one or more of a subset of remapping fields that are associated with the set of virtual target TLBI domains indicated by the TLBI domain information.

6. The apparatus according to any of claims 4 and 5, in which the TLBI broadcast control circuitry is configured to determine whether virtualized remapping of target TLBI domains is enabled based on virtualized remapping enable control information configurable by software.

7. The apparatus according to any of claims 4 to 6, in which the target domain remapping information comprises information specified in at least one control register.

8. The apparatus according to claim 7, in which the at least one control register comprises one or more registers to which write access is denied for instructions executed with less privilege than a hypervisor-level privilege.

9. The apparatus according to any of claims 4 to 8, in which said at least one operating state comprises an operating state in which instructions are executed with less privilege than a hypervisor-level privilege.

10. The apparatus according to any preceding claim, in which the TLBI broadcast control circuitry is configured to determine that a local TLBI domain is a target TLBI domain even if the local TLBI domain is determined to be a non-target TLBI domain using the TLBI domain information, the local TLBI domain comprising a TLBI domain which is associated with at least one TLB of a processor comprising the instruction decoding circuitry which decoded the TLBI instruction.11 . The apparatus according to any preceding claim, in which, for at least one variant of the TLBI instruction, the TLBI domain information comprises a plurality of TLBI domain indicators each associated with a corresponding TLBI domain and indicating whether that TLBI domain is a target TLBI domain or a non-target TLBI domain.

12. I he apparatus according to any preceding claim, in which, for at least one variant of the TLBI instruction, the TLBI domain information specifies at least one TLBI domain number identifying at least one target TLBI domain.

13. The apparatus according to any preceding claim, in which the plurality of TLBI domains comprise absolute domains for which an association between a given TLB and its associated TLBI domain is the same regardless of which of a plurality of processors executes the TLBI instruction.

14. The apparatus according to any preceding claim, in which the plurality of TLBI domains are non-overlapping so that each TLB is associated with only one TLBI domain.

15. The apparatus according to any preceding claim, in which the TLBI instruction also specifies shareability domain information indicating one of a plurality of shareability domains for which TLBs associated with the selected shareability domain are to observe the TLB invalidation event; and the TLBI broadcast control circuitry is configured to determine the one or more TLBI request signals based on the TLBI domain information when the shareability domain information indicates a selected shareability domain.

16. The apparatus according to any preceding claim, in which the TLBI broadcast control circuitry is configured to permit over-invalidation where at least one non-target TLBI domain not required to observe the TLB invalidation event nevertheless observes the TLB invalidation event.

17. The apparatus according to any preceding claim, in which the one or more TLBI request signals specify target domain information indicating which TLBI domains are required to observe the TLB invalidation event.

18. The apparatus according to any preceding claim, in which the TLBI broadcast control circuitry is configured to determine, based at least on the TLBI domain information, whether it is necessary to issue at least one TLBI request signal to an off-chip TLBI domain located on a separate integrated circuit from an integrated circuit comprising the instruction decoding circuitry and the TLBI broadcast control circuitry.

19. The apparatus according to any preceding claim, comprising TLBI tracking circuitry to track, with respect to a given processor, which TLBI domains have outstanding TLB invalidation events instructed based on TLBI instructions executed on the given processor; andin response to a synchronisation barrier instruction executed on the given processor, the given processor is configured to prevent instructions younger than the synchronisation barrier instruction from being executed until the TLBI tracking circuitry confirms that all TLBI domains instructed to observe TLB invalidation events based on older TLBI instructions which are older than the synchronisation barrier instruction have acknowledged that the TLB invalidation events instructed by the older TLBI instructions are guaranteed to be completed.

20. The apparatus according to any preceding claim, in which: the TLBI domain information identifies the TLBI domains based on a TLBI domain number space; a first subset of TLBI numbers of the TLBI domain number space correspond to actual TLBI domains; a second subset of TLBI numbers of the TLBI domain number space correspond to aliased TLBI domains; and for a given TLBI instruction for which the TLBI broadcast control circuitry determines that a given aliased TLBI domain corresponding to one of the second subset of TLBI numbers is identified as a target TLBI domain using the TLBI domain information, the TLBI broadcast control circuitry is configured to map the given aliased TLBI domain to a group of two or more actual TLBI domains and determine that said group of two or more actual TLBI domains should be considered target TLBI domains.21 . Computer-readable code for fabrication of an apparatus comprising: instruction decoding circuitry to decode a translation lookaside buffer invalidation, TLBI, instruction for instructing at least one translation lookaside buffer, TLB, to observe a TLB invalidation event, the TLBI instruction specifying TLBI domain information indicating, for each TLBI domain of a plurality of TLBI domains, whether that TLBI domain is a target TLBI domain for which any TLBs in that TLBI domain are required to observe the TLB invalidation event or a nontarget TLBI domain for which any TLBs in that TLBI domain are not required to observe the TLB invalidation event, the plurality of TLBI domains including at least two non-overlapping TLBI domains; andTLBI broadcast control circuitry to determine, based on the TLBI domain information specified by the TLBI instruction decoded by the instruction decoding circuitry, one or more TLBI request signals to be issued for instructing TLBs in one or more TLBI domains to observe the TLB invalidation event; wherein the TLBI broadcast control circuitry is configured to limit which TLBI domains are instructed to observe the TLB invalidation event based on the TLBI domain information.

22. A method comprising: decoding a translation lookaside buffer invalidation, TLBI, instruction for instructing at least one translation lookaside buffer, TLB, to observe a TLB invalidation event, the TLBI instruction specifying TLBI domain information indicating, for each TLBI domain of a plurality of TLBI domains, whether that TLBI domain is a target TLBI domain for which any TLBs in that TLBI domain are required to observe the TLB invalidation event or a non-target TLBI domain for which any TLBs in that TLBI domain are not required to observe the TLB invalidation event, the plurality of TLBI domains including at least two non-overlapping TLBI domains; and determining, based on the TLBI domain information specified by the TLBI instruction decoded by the instruction decoding circuitry, one or more TLBI request signals to be issued for instructing TLBs in one or more TLBI domains to observe the TLB invalidation event, to limit which TLBI domains are instructed to observe the TLB invalidation event based on the TLBI domain information.

23. A computer program for controlling a host data processing apparatus to provide an instruction execution environment for execution of target program code, the computer program comprising: instruction decoding program logic to decode a translation lookaside buffer invalidation, TLBI, instruction for instructing at least one simulated translation lookaside buffer, TLB, to observe a TLB invalidation event, the TLBI instruction specifying TLBI domain information indicating, for each TLBI domain of a plurality of TLBI domains, whether that TLBI domain is a target TLBI domain for which any simulated TLBs in that TLBI domain are required to observe the TLB invalidation event or a non-target TLBI domain for which any simulated TLBs in that TLBI domain are not required to observe the TLB invalidation event, the plurality of TLBI domains including at least two non-overlapping TLBI domains; andTLBI control program logic to determine, based on the TLBI domain information specified by the TLBI instruction decoded by the instruction decoding program logic, one or more simulated TLBs in one or more TLBI domains which should observe the TLB invalidation event; wherein the TLBI broadcast control program logic is configured to limit which TLBI domains should observe the TLB invalidation event based on the TLBI domain information.

24. A storage medium storing the computer-readable code of claim 21 or the computer program of claim 23.

Citation Information

Patent Citations

  • Realm identifier comparison for translation cache lookup

    US20200159677A1

  • Realm execution context masking and saving

    US20200192585A1