Masking control for ancillary event interrupt
Patent Information
- Application Number
- PCT/GB2026/050333
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-20
- Filing Date
- 2026-03-05
- Publication Date
- 2026-09-24
Smart Images

Figure GB2026050333_24092026_PF_FP_ABST
Abstract
Description
[0001] MASKING CONTROL FOR ANCILLARY EVENT INTERRUPT
[0002] The present technique relates to the field of data processing.
[0003] Processing circuitry for processing instructions can be provided with ancillary circuitry for performing an offloaded function on behalf of the processing circuitry. A control interface may be provided to enable the processing circuitry to offload functionality to the ancillary circuitry.
[0004] At least some examples of the present technique provide an apparatus comprising: processing circuitry configured to process instructions of a current processing context identified by at least one processing context identifier stored in processing context identifier storage circuitry; control interface circuitry configured to exchange control signals with ancillary circuitry configurable to perform an offloaded function on behalf of the processing circuitry, where the ancillary circuitry is capable of performing, in parallel with processing of the current processing context on the processing circuitry, the offloaded function for a non-resident execution context that is different to the current processing context; ancillary context identifier storage circuitry configured to store at least one ancillary context identifier indicative of a current ancillary context associated with the offloaded function allocated to the ancillary circuitry; and wherein the control interface circuitry is configured to: determine, based on a comparison between the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one ancillary context identifier stored in the ancillary context identifier storage circuitry, whether an ancillary event interrupt is masked, the ancillary event interrupt comprising an interrupt for reporting to the processing circuitry an event associated with the ancillary circuitry; and in response to detecting the event associated with the ancillary circuitry when the ancillary event interrupt is not masked, signal the ancillary event interrupt to the processing circuitry.
[0005] At least some examples of the present technique provide a system comprising: the apparatus described above, implemented in at least one packaged chip; at least one system component; and a board, wherein the at least one packaged chip and the at least one system component are assembled on the board.
[0006] At least some examples of the present technique provide a chip-containing product comprising the system described above, wherein the system is assembled on a further board with at least one other product component.
[0007] At least some examples of the present technique provide computer-readable code for fabrication of an apparatus as described above. The computer-readable code can be stored on a storage medium. The storage medium may be a non-transitory storage medium.
[0008] At least some examples of the present technique provide a method comprising: storing, in processing context identifier storage circuitry, at least one processing context identifier indicative of a current processing context for which instructions are processed by processing circuitry; storing, in ancillary context identifier storage circuitry, at least one ancillary context identifier indicative of a current ancillary context associated with ancillary circuitry configurable to perform an offloaded function on behalf of the processing circuitry; depending on a comparisonbetween the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one ancillary context identifier stored in the ancillary context identifier storage circuitry, determining whether an ancillary event interrupt is masked, the ancillary event interrupt comprising an interrupt for reporting to the processing circuitry an event associated with the ancillary circuitry; and in response to detecting the event associated with the ancillary circuitry when the ancillary event interrupt is not masked, signalling the ancillary event interrupt to the processing circuitry.
[0009] Further aspects, features and advantages of the present technique will be apparent from the following description of examples, which is to be read in conjunction with the accompanying drawings, in which:
[0010] Figure 1 illustrates an example of an apparatus comprising processing circuitry and control interface circuitry for exchanging control signals with ancillary circuitry;
[0011] Figure 2 illustrates steps for determining whether the control interface circuitry is available to a current processing context associated with the processing circuitry;
[0012] Figure 3 illustrates more detail for determining whether the control interface circuitry is available;
[0013] Figure 4 illustrates steps for determining whether an ancillary event interrupt is masked; Figure 5 illustrates more detail for determining whether the ancillary event interrupt is masked;
[0014] Figure 6 illustrates an example where the ancillary circuitry comprises a hardware accelerator;
[0015] Figure 7 illustrates an example of control interface registers for the control interface circuitry, including ancillary context identifier registers;
[0016] Figure 8 illustrates an example of signalling on a communication link between the control interface circuitry of a processor and the accelerator; and
[0017] Figure 9 illustrates a system and a chip-containing product.
[0018] An apparatus comprises processing circuitry configured to process instructions of a current processing context identified by at least one processing context identifier stored in processing context identifier storage circuitry; and control interface circuitry configured to exchange control signals with ancillary circuitry configurable to perform an offloaded function on behalf of the processing circuitry. The ancillary circuitry is capable of performing the offloaded function for a non-resident execution context that is different to the current processing context. Ancillary context identifier storage circuitry stores at least one ancillary context identifier indicative of a current ancillary context associated with the offloaded function allocated to the ancillary circuitry. Whether the control interface circuitry is available to the current processing context is controlled by interface availability control circuitry depending on a comparison between the at least one processing context identifier and the at least one ancillary context identifier.This approach can be helpful for supporting a thread-asynchronous model for usage of the ancillary circuitry, where the ancillary circuitry can continue performing offloaded function on behalf of a current ancillary context corresponding to a non-resident execution context while the processing circuitry is processing a different current execution context as the current processing context. This can help avoid the energy and latency cost that would be associated with context switching the ancillary circuitry each time there is a change in current processing context at the processing circuitry (as would be incurred if a thread-synchronous usage model was used for the ancillary circuitry, where it is assumed that at all times the ancillary circuitry is associated with the same execution context as the processing circuitry).
[0019] However, when a thread-asynchronous approach is used, this means that at times when the ancillary circuitry is allocated for a different execution context from the current execution context whose instructions are processed on the processing circuitry, the control interface circuitry via which the ancillary circuitry can be configured to perform offloaded functions on behalf of the processing circuitry may be unavailable to the current execution context. In examples below, the availability of the control interface circuitry is controlled based on a comparison between at least one processing context identifier stored in processing context identifier storage circuitry and at least one ancillary context identifier stored in ancillary context identifier storage circuitry. This provides an interface which is much easier for software to use to control whether the control interface circuitry is available. It can be much simpler for software on the processing circuitry to maintain the context identifiers stored in the processing / ancillary context identifier storage circuitry respectively in response to changes in context associated with the processing circuitry and ancillary circuitry, than to maintain an explicit indication of whether the control interface circuitry is currently available to the current processing context. Software on the processing circuitry can simply update the processing context identifier(s) when switching the current execution context of the processing circuitry, and update the ancillary context identifier(s) when switching the current ancillary context for a function allocated to the ancillary circuitry, rather than needing to update an availability register to specify an availability parameter determined directly by software. By providing hardware of the interface availability control circuitry to generate an indication of interface availability automatically based on a comparison between the respective context identifiers, this makes it simpler to develop software for using the ancillary circuitry in a thread-asynchronous usage model.
[0020] The interface availability control circuitry may determine that the control interface circuitry is unavailable to the current processing context in response to the comparison detecting a mismatch between the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one ancillary context identifier stored in the ancillary context identifier storage circuitry. Hence, when the ancillary processing context differs from the current processing context, the interface is determined to be unavailable to the current processing context.When the control interface circuitry is determined to be unavailable to the current processing context, the interface availability control circuitry may prevent an instruction executed by the processing circuitry being able to cause the control interface circuitry to perform at least one control action for the ancillary circuitry. Hence, instructions of the current processing context executed by the processing circuitry may not be able to use various interface functions of the control interface circuitry which would interfere with use of the ancillary circuitry by the ancillary processing context and / or expose information about the operation of the ancillary circuitry to the current processing context. The control action that is prohibited when the control interface circuitry is determined to be unavailable could, for example, including any one or more of the following:
[0021] • launching a command to the ancillary circuitry;
[0022] • updating control registers of the ancillary circuitry
[0023] • querying status of the ancillary circuitry;
[0024] • obtaining information about progress of previously launched commands to the ancillary circuitry; and
[0025] • obtaining information about faults associated with the ancillary circuitry.
[0026] In response to an attempt by an instruction executed by the processing circuitry to cause the control interface circuitry to perform a control action for the ancillary circuitry when the control interface circuitry is determined to be unavailable to the current processing context, the interface availability control circuitry may cause an unavailability response to be returned to the processing circuitry. Hence, instructions of the current execution context can determine from the unavailability response that the control interface circuitry is not currently available as the ancillary circuitry is already in use for a different execution context. This can prompt the software of the current execution context to, if necessary, request a supervisory process (such as an operating system or hypervisor) to reallocate which execution context is the current ancillary context so that the current execution context on the processing circuitry can then make use of the ancillary circuitry.
[0027] The unavailability response could be implemented in different ways. In some examples, the unavailability response comprises setting an unavailable indication in a storage location pollable by instructions executed by the processing circuitry. In this case, the unavailability response does not need to be signalled to the software of the current execution context directly in response to the instruction that tried to access the control interface circuitry, but could be detected later by a subsequent instruction of the current execution context reading the storage location comprising the unavailable indication.
[0028] In some examples, the unavailability response comprises an exception being signalled to the processing circuitry (e.g. an exception denoting an access violation). For example, an attempt to read or write a given register of the control interface circuitry when the control interface circuitry is determined to be unavailable may trigger generation of an access violation exception. Theexception handler software executed in response to the exception could then detect the cause of the exception and decide to call supervisory software to request a switch of ancillary context.
[0029] Another function of the control interface circuitry which may depend on whether the current ancillary context of the ancillary circuitry matches the current processing context of the processing circuitry may be generation of ancillary event interrupt used to report an event associated to with the ancillary circuitry to the processing circuitry. For example, the ancillary event interrupt could be used to report events such as completion of a command by the ancillary circuitry, occurrence of a fault or error, etc. In a system supporting thread-asynchronous use of the ancillary circuitry, whether the ancillary event interrupt is useful at a given time may depend on whether the current ancillary context matches the current processing context. It can be helpful to support masking of the ancillary event interrupt, so that generation of the ancillary event interrupt is suppressed when the interrupt is not considered useful. However, it may be less convenient for software to have to explicitly set an interrupt mask control value depending on whether the interrupt is currently considered useful. In the examples discussed below, the control interface circuitry has hardware for comparing the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one ancillary context identifier stored in the ancillary context identifier storage circuitry to control whether the ancillary event interrupt is masked, so that this comparison does not need to be done by software. In response to detecting the event associated with the ancillary circuitry when the ancillary event interrupt is determined, based on the comparison of context identifiers, not to be masked, the ancillary event interrupt can be signalled to the processing circuitry. Similar to the control of availability of the control interface circuitry, this approach to intron masking control as much simpler for software to control since simply by setting the relevant context identifiers identifying the respective contexts that are current at the processing circuitry and ancillary circuitry respectively, this can automatically control whether or not the ancillary event interrupt is masked, without requiring any explicit comparison of context identifiers to be performed in software to determine the masked status for a given interrupt.
[0030] For some types of ancillary event interrupt, the interrupt may be considered relevant when the current processing context matches the current ancillary context, so that the ancillary event interrupt is determined not to be masked when the current processing context matches the current ancillary context, and determined to be masked when the current processing context mismatches the current ancillary context.
[0031] However, in many cases, when the current processing context is the same as the current ancillary context, it may be undesirable to interrupt the current processing context when an event occurs at the ancillary circuitry, because it may be more efficient for the software of the current processing context to wait until the point of program flow at which the outcome of the offloaded function from the ancillary circuitry is needed, before polling status of the ancillary circuitry at a time that suits the current processing context (this can improve processing performance byreducing the extent to which other tasks not related to the ancillary circuitry are interrupted part way through). The ancillary event interrupt may be more useful in cases where the current processing context differs from the current ancillary context, so that the interrupt may act as a prompt (hint) for an operating system or hypervisor to context switch the current processing context of the processing circuitry so that the context matching the current ancillary context can resume processing at the processing circuitry and deal with the event that occurred at the ancillary circuitry. Hence, in some examples, the control interface circuitry may determine that the ancillary event interrupt is masked, in response to detecting a match between the at least one processing context identifier and the at least one ancillary context identifier. The control interface circuitry may determine that the ancillary event interrupt is not masked in response to detecting a mismatch between the at least one processing context identifier and the at least one ancillary context identifier.
[0032] The control interface circuitry may detect the event associated with the ancillary circuitry based on an event signal (or set of event signals) received from the ancillary circuitry. For example, the event associated with the ancillary circuitry could indicate any one or more of the following:
[0033] • the ancillary circuitry completing an instance of the offloaded function;
[0034] • the ancillary circuitry encountering an error;
[0035] • the ancillary circuitry reporting occurrence of an event that the ancillary circuitry was configured to monitor.
[0036] The ancillary event interrupt may comprise a hint to software that an event associated with the ancillary circuitry has occurred which relates to a non-resident processing context which differs from the current processing context for which instructions are being processed by the processing circuitry. Software on the processing circuitry can use the hint provided by the interrupt to determine whether to trigger a context switch at the processing circuitry.
[0037] The processing context identifier storage circuitry and the ancillary context identifier storage circuitry may each be programmable in response to instructions executed by the processing circuitry. In some examples, different parts of the processing context identifier storage circuitry and the ancillary context identifier storage circuitry may be associated with different access permissions controlling the right for the processing circuitry to update that part of the storage circuitry when operating in a given execution state. For example, some identifiers may be restricted to being updated by software executing in an operating state with hypervisor privilege, while other identifiers may be able to be updated by software executing in an operating state with operating-system privilege.
[0038] The programming interface for the processing context identifier storage circuitry and the ancillary context identifier storage circuitry may be implemented in different ways.
[0039] In some examples, at least part of the ancillary context identifier storage circuitry may comprise at least one memory-mapped register, which is programmable by the processingcircuitry executing a store instruction specifying a memory address mapped to that memorymapped register. Similarly, at least part of the processing context identifier storage circuitry may comprise at least one memory-mapped register. A memory-mapped control interface for accessing at least some of the processing / ancillary context identifier storage can avoid needing to use up encoding space in an instruction set architecture for dedicated instructions for setting the relevant part of the processing / ancillary context identifier storage.
[0040] On the other hand, for some examples, at least part of the ancillary context identifier storage circuitry may comprise a system register updatable in response to a system register update instruction executed by the processing circuitry. Similarly, at least part of the processing context identifier storage circuitry may comprise a system register updatable in response to a system register update instruction executed by the processing circuitry. Providing a dedicated instruction able to update part of the processing / ancillary can help improve processing performance for software interacting with the ancillary circuitry, as using a system register to update the context identifier may be less prone to memory address translation delays associated with page table walks for obtaining mapping information for determining the physical address mapped to the memory-mapped register.
[0041] Some examples may use a combination of one or more memory-mapped registers and one or more system registers for at least one of the ancillary context identifier storage circuitry and the processing context identifier storage circuitry. For example, in examples where the current processing context is identified by a combination of multiple identifiers, it may be that some system registers already exist in an instruction set architecture for indicating a subset of those identifiers, but additional memory-mapped control registers are then also provided for indicating the remaining identifiers used to fully identify the current processing context for the purpose of comparing against the current ancillary context. Hence, it may be that a combination of a memory-mapped access control mechanism and a system register instruction access control mechanism can be used for updating different parts of the context identifiers.
[0042] The particular identifier or identifiers used to identify the current processing context and current ancillary context may vary from one implementation to another, or between one operating state of an apparatus and another. In many cases, a combination of multiple identifiers may be used to uniquely identify the current processing context and current ancillary context.
[0043] In some cases, the at least one processing context identifier and the at least one ancillary context identifier may each comprise a process identifier identifying a given software process.
[0044] If the process can support execution of multiple independent threads which may independently be allocated use of the ancillary circuitry, then the at least one processing context identifier and the at least one ancillary context identifier may each comprise a combination of the process identifier and a thread identifier.In examples supporting virtualisation, the at least one processing context identifier and the at least one ancillary context identifier may each comprise a combination of at least the process identifier and a virtual machine identifier (and, if applicable, the thread identifier).
[0045] In examples where the processing circuitry supports execution in different security states, the at least one processing context identifier and the at least one ancillary context identifier may each comprise a combination of at least the process identifier and a security state identifier (and, if applicable, the thread identifier). If the processing circuitry supports both virtualisation and use of multiple security states, the at least one processing context identifier and the at least one ancillary context identifier may each comprise a combination of at least the process identifier, the virtual machine identifier and the security state identifier (and again, if applicable, the thread identifier).
[0046] Hence, it will be appreciated that there can be a wide variety of ways in which the at least one processing context identifier and the at least one ancillary context identifier can be implemented.
[0047] The techniques discussed above can be used with a wide variety of types of ancillary circuitry. In general, the ancillary circuitry can be any ancillary circuit unit capable of performing an offloaded function on behalf of processing circuitry. The techniques discussed above can be particularly useful where the ancillary circuitry is private to a particular processor (rather than being shared across multiple processors), so that the processor has its own dedicated hardware interface to the ancillary circuitry (rather than interacting with the ancillary circuitry via control data structures stored in memory).
[0048] For example, the ancillary circuitry could be an input / output (peripheral) device such as a network controller or display controller.
[0049] Another example of the ancillary circuitry could be an internal system component within a processing system, such as a cryptographic unit for performing cryptographic operations.
[0050] Another example of the ancillary circuitry could be monitoring circuitry which can be configured by the processing circuitry to monitor for occurrences of a given event and notify the processing circuitry when the event occurs. For example, the monitoring circuitry could be used to notify the processing circuitry when data for a given memory address is updated by an external memory access requester.
[0051] However, the techniques discussed above can be particularly useful in cases where the ancillary circuitry comprises a hardware accelerator configured to perform a delegated task asynchronously with respect to operations performed by the processing circuitry (the delegated task being an example of the offloaded function mentioned earlier). The hardware accelerator may comprise circuit logic which is more specialized than the general purpose logic of a processing circuitry performing synchronous operations in response to instructions processed synchronously by a processing pipeline. It can be particularly useful to support a thread asynchronous usage model for a hardware accelerator because the hardware accelerator maycomprise a relatively large amount of internal state storage for storing internal state information associated with its delegated task, and so the energy cost, memory bandwidth and latency incurred in performing a context switch of the hardware accelerator part way through performing the task may be considerable, so it may be preferable to allow the accelerator to continue processing a task on behalf of a first context even after the processing circuitry has switched away from processing the first context and is now processing a second context as the current processing context. Hence, by controlling availability of the control interface for the hardware accelerator and / or masking of the ancillary event interrupt based on the comparison between context identifiers stored in the processing context identifier storage circuitry and the ancillary context identifier storage circuitry, this can support a more efficient thread-asynchronous usage model for a hardware accelerator accessed by processing circuitry via the control interface circuitry.
[0052] These techniques can be particularly useful for cases where the hardware accelerator comprises accelerator circuitry configured to accelerate operations for one or more machine learning workloads. For example, the hardware accelerator may comprise a machine learning accelerator, also known as an artificial intelligence (Al) accelerator, neural processing unit (NPU), or neural engine. Such accelerators for supporting machine learning workloads may have a large amount of internal state and so the cost of a context switch may be particularly high, so avoiding the context switch by supporting thread asynchronous usage can be particularly useful for machine learning accelerators.
[0053] In some examples, the hardware accelerator comprises a processor-local accelerator associated with the processing circuitry. This may differ from a remote accelerator that is accessed via the memory system, is shared between multiple processors and is configured via an I / O mechanism by setting control data structures stored in memory. Use of a processor-local accelerator can improve processing performance as the latency of control of offloading functions to the accelerator can be greatly reduced, making it feasible to use the accelerator even for relatively short delegated tasks for which the control overhead would have been prohibitive with a remote accelerator usage model.
[0054] In some examples, the hardware accelerator is private to the processing circuitry (so cannot be accessed by another processor other than the processor having the processing circuitry). Providing dedicated processor-local accelerator functionality private to a particular processor can help improve performance for workloads such as machine learning.
[0055] In some examples, the offloading of the offloaded function to the ancillary circuitry via the control interface circuitry may at least partly be performed using a communication signal path separate from a path by which the processing circuitry issues memory access requests to read / write data stored in memory. For example, a dedicated control interface bus may be provided to connect the processing circuitry and the ancillary circuitry (e.g. the accelerator), separate from the memory system bus / interconnect via which the processing circuitry accesses memory. Thiscan speed up configuration of the ancillary circuitry compared to an approach where the configuration actions for configuring the ancillary circuitry have to contend for bandwidth on the general purpose memory system interconnect.
[0056] In some examples, the processing circuitry and the hardware accelerator share translation table walk circuitry configured to control translation table walk operations for obtaining address translation table entries from memory. By reusing the translation table walk circuitry of the processing circuitry for memory accesses triggered by the accelerator which miss in a translation lookaside buffer, this saves circuit area by avoiding the need to duplicate the translation table walk circuitry at both processing circuitry and accelerator. In some examples, the processing circuitry and accelerator could also share at least one translation lookaside buffer (TLB) for caching address translation information obtained in a translation table walk operation. However, it is also possible for the accelerator to have its own dedicated TLB looked up for memory accesses triggered by the accelerator to identify address translation information. Nevertheless, if the TLB of the accelerator detects a miss for a given address to be accessed by the accelerator, the shared translation table walk circuitry associated with the processing circuitry can be used to perform the translation table walk operation to find the missing address translation information.
[0057] In some examples, the processing circuitry and the hardware accelerator share at least one private cache (a cache which is private to a processor comprising the processing circuitry and the hardware accelerator, and not directly accessible by other processors of a multi-processor system). For example, the shared private cache could be a level 2 cache. By sharing at least one level of private cache, this reduces total circuit area, as well as enabling the processing circuitry and accelerator to exchange configuration information, status information and / or results of delegated tasks faster than if communication of this information had to be performed via main memory without the accelerator having direct access into a private cache of the processing circuitry.
[0058] In some examples, the accelerator control interface circuitry is configurable to launch a new delegated task to the hardware accelerator in response to one or more instructions executed by a processing circuitry in an operating state with user-level privilege. For example, the userlevel privilege may be the least privileged level of privilege supported by the processing circuitry (e.g. a state less privileged than an operating system level of privilege). Direct configuration of the accelerator from user-level software can be beneficial in reducing performance overhead associated with configuration of the accelerator, since it removes the need for user-level application software to call more privileged software (e.g. an operating system or hypervisor) to request access to the accelerator, which would cause a significant delay associated with exception entry / exit processes.
[0059] In some examples, the control interface circuitry comprises memory-mapped register storage accessible to the processing circuitry in response to load / store instructions executed by the processing circuitry specifying target addresses mapped to the memory-mapped registerstorage. For example, the memory-mapped register storage could include interface registers used for various purposes, such as launching commands to the ancillary circuitry, checking status of the ancillary circuitry or of individual commands, controlling register updates of internal registers of the ancillary circuitry, and / or performing context switch operations when required to change context of the ancillary circuitry (even though the mechanisms for availability control and interrupt masking control described above help reduce the need for ancillary context switching by supporting a thread-asynchronous usage model, sometimes it may still be desired to perform an ancillary context switch part way through a task allocated to the ancillary circuitry, so the interface may still support the ability to trigger ancillary context switches). The memory-mapped register storage can also include at least part of the processing context identifier storage circuitry and / or ancillary context identifier storage circuitry described earlier. Use of a memory-mapped interface for accessing registers of the control interface circuitry can reduce the amount of encoding space in an instruction set needed for control of the control interface, since a standard load / store instruction can be used to access the interface, with the address specified by the load / store instruction identifying which interface register is to be read / written.
[0060] In some examples, the apparatus may be implemented as a chiplet which implements a subset of functionality of a wider processing system distributed between multiple chiplets. Each chiplet may be manufactured separately and then packaged together at a later stage of a manufacturing flow. Hence, it is not essential that the apparatus comprising the processing circuitry also comprises the ancillary circuitry itself, as the ancillary circuitry could be on a different chiplet or other component external to the apparatus comprising the processing circuitry.
[0061] However, in some examples, the apparatus also comprises the ancillary circuitry.
[0062] Specific examples are now described with reference to the drawings.
[0063] Figure 1 illustrates an example of an apparatus 2 having processing circuitry 6 for performing data processing operations in response to instructions fetched from an instruction cache or memory. The processing circuitry 6 may include instruction decoding circuitry (not shown in Figure 1) for decoding the instructions and generating control signals for controlling the processing circuitry 6 to perform the corresponding processing operations represented by the instructions. The processing circuitry may support a given instruction set architecture (ISA) which defines the encoding of each supported instruction and the corresponding operations which the processing circuitry 6 is expected to perform in response to the instruction. The processing circuitry 6 can access a memory system (comprising one or more caches and random access memory storage) via a load / store interface. When the processing circuitry 6 executes a load instruction specifying one or more operands defining a load target address, a load operation is performed to load data value from a memory system location corresponding to the load target address to a register of the processing circuitry 6. When the processing circuitry 6 executes the store instruction specifying one or more operands defining a store target address, a storeoperation is performed to store the data value obtained from a register of the processing circuitry 6 to a memory system location corresponding to the store target address.
[0064] Control interface circuitry 14 is provided to exchange signals with ancillary circuitry 116 via an ancillary communication path 117 (e.g. a bus supporting a given bus protocol), for configuring the ancillary circuitry 116 to perform ancillary functions on behalf of the processing circuitry 6. For example, the ancillary circuitry 116 could be an input / output device, a cryptographic unit for performing encryption / decryption operations, a monitoring circuit for monitoring events on behalf of the processing circuitry and notifying the processing circuitry when a relevant event occurs, etc. In one particular example described in more detail with respect to Figure 6 below, the ancillary circuitry 116 comprises a hardware accelerator. The control interface circuitry 14 provides a configuration interface by which the processing circuitry 6 can request that a task or function is offloaded to the ancillary circuitry 116 and check status of the ancillary circuitry 116 (e.g. checking whether a previously offloaded task is complete or whether any fault or error has occurred at the ancillary circuitry 116). In some examples, the ISA supported by the processing circuitry may support at least one ancillary control instruction which when decoded and executed by the processing circuitry 6 causes the control interface circuitry 14 to perform an associated ancillary control action such as a launching a command to the ancillary circuitry 116 or returning an indication of current status of the ancillary circuitry 116 (that instruction being a different instruction to the load / store instructions used to perform memory accesses). In other examples, the control interface circuitry 14 could be implemented as a memory-mapped configuration interface comprising a set of memory-mapped registers which are exposed to the software executed on the processing circuitry 6 as memory system locations addressable via addresses in a particular address region of the address space. If a memory-mapped interface is used, then the processing circuitry 6 can cause the control interface circuitry 14 to perform ancillary control actions by executing a load / store instruction specifying, as the load / store target address, one of the addresses mapped to a given control register of the control interface circuitry 14. The processing circuitry 6 may have circuit logic to detect, from the load / store target address of a load / store instruction, whether the instruction is attempting to access the control interface circuitry 14, and if so direct a request to the control interface circuitry 14 rather than the regular path by which load / store operations are performed to memory. Circuitry within the control interface circuitry 14 may detect, based on the target address of a load / store to a given interface control register, what type of control action is needed and then communicate with the ancillary circuitry 116 and / or provide a response to the processing circuitry 6 depending on the type of control action requested.
[0065] As shown in Figure 1, the ancillary circuitry 116 may have its own path by which it can access the memory system shared with the processing circuitry 6. Alternatively, the memory access requests from the ancillary circuitry could be routed via the processing circuitry 6 andshare the same memory access path that is used for load / store operations triggered by instructions executed on the processing circuitry 6.
[0066] The ancillary circuitry 116 may have internal storage for storing internal state information 118. For example, the internal storage could comprise control registers and / or one or more caches. The amount of internal state 118 could be relatively large for some examples, so it may be undesirable to operate the ancillary circuitry 116 in a thread-synchronous manner for which the ancillary circuitry 116 is context switched (by storing the internal state 118 for an outgoing context to memory and reading in internal state 118 for an incoming context) each time there is a corresponding change of processing context at the processing circuitry 6. Instead, a thread-asynchronous approach may be used where the ancillary circuitry 116 can be permitted to continue performing an offloaded function on behalf of a current ancillary context which differs from the current processor context which is active at the processing circuitry 6. This can improve energy efficiency and performance by reducing the frequency with which costly ancillary context switches are needed. However, this means that a given processing context being processed at the processing circuitry 6 cannot assume that the interface to the ancillary circuitry 116 is always available, as if the ancillary circuitry 116 is performing functions for a different ancillary context then it may be undesirable to accept commands from, or expose status information to, instructions of the current processing context. Also, the control interface circuitry 14 may support an ancillary event interrupt which can be supplied to the processing circuitry 6 in response to detecting an ancillary event occurring for the ancillary circuitry 116, for interrupting software running on the processing circuitry 6 so that an exception handler can direct software to deal with the cause of the ancillary event. The ancillary event could be a notification that results of an offloaded function are complete, or notification that an error or fault has occurred. Whether or not it is useful to generate the ancillary event interrupt may depend on whether the current ancillary context at the ancillary circuitry 116 matches the current processing context at the processing circuitry 6.
[0067] Hence, to support control of availability of the control interface circuitry 14 and control over whether the ancillary event interrupt is masked or not masked, the apparatus has a set of one or more processing context identifier registers (an example of processing context identifier storage circuitry) 7 for storing at least one processing context identifier identifying a current processing context assigned to the processing circuitry, and a set of one or more ancillary context identifier registers (an example of ancillary context identifier storage circuitry) 23 for storing at least one ancillary context identifier identifying a current ancillary context allocated to the ancillary circuitry 116.
[0068] A comparator (an example of interface availability control circuitry) 4 compares the identifiers in the processing context identifier registers 7 and the ancillary context identifier registers 23 and controls availability of the control interface circuitry 14 depending on the comparison. While the comparator 4 is shown separate from the control interface circuitry 14 in Figure 1, in other examples the comparator 4 could be regarded as part of the control interfacecircuitry 14. If there is a mismatch between the identifiers in corresponding registers 7, 23, the interface is indicated as unavailable and an attempt by the processing circuitry 6 to interact with the interface may cause an unavailability response to be generated (e.g. the unavailability response could simply be setting an unavailability status in a register or other location that can be polled by the processing circuitry 6, or could comprise raising an exception).
[0069] The comparator 4 also controls the determination by the control interface circuitry 14 of whether the ancillary event interrupt is masked. The ancillary event interrupt can be masked when the processing context identifiers in registers 7 match the corresponding ancillary context identifiers in registers 23, and the ancillary event interrupt is enabled (not masked) when there is a mismatch between the processing context identifiers in registers 7 and the corresponding ancillary context identifiers in registers 23. This means that when the current processing context corresponds to the current ancillary processing context, the current processing context does not need to be interrupted to notify it of ancillary events, as the software of the current processing context can check for such events in its own time when it reaches the point of program flow at which use of the ancillary circuitry 116 is relevant. On the other hand, when an event occurs at the ancillary circuitry 116 for a current ancillary context different from the current processing context, the ancillary event interrupt can be useful for prompt software on the processing circuitry 6 to switch the current processing context (without that interrupt, there is a risk that the processing context corresponding to the current ancillary context never gets scheduled, and so the ancillary circuitry 116 may, even if it has completed its task, not be able to be reallocated for other contexts, so opportunities to use the ancillary circuitry 116 for new tasks may be lost).
[0070] The particular set of identifiers used to identify a given context in registers 7, 23 may vary depending on the ISA supported by the processing circuitry 6. For example, each of the processing context registers 7 and ancillary context registers 23 may comprise storage to identify a combination of one or more identifiers that can uniquely identifying a given thread of processing on the processing circuitry 6.
[0071] For example, in some examples, a given context may be identified by a process identifier. In some examples, a given context may be identified by the combination of a thread identifier identifying a particular thread of a process and a process identifier identifying the process.
[0072] In some examples supporting virtualisation, a given context may be identified by the combination of the process identifier (and thread identifier if applicable) and virtualisation identifier identifying the virtual environment (e.g. virtual machine) in which the process executes.
[0073] In some examples supporting execution of instructions in different security states, a given context may be identified by the combination of a process identifier (and thread identifier if applicable) and a security state in which the given context operates (and if virtualisation is supported then also the virtualisation identifier).Hence, where the context is identified by a combination of multiple identifiers, the context identifier storage 7, 23 may comprise multiple registers each storing a corresponding identifier (alternatively, more than one identifier could be stored within the same register). Different identifiers in the set of context identifiers may be associated with different access permissions controlling the execution states of the processing circuitry 6 for which instructions are allowed to update that context identifier. For example, a thread identifier or process identifier may be restricted to being updated by instructions executing in an operating state with at least an operating-system-level of privilege, so that an operating system or kernel can update the thread / process identifier but application-level (user-level) code cannot update the thread identifier or process identifier. A virtualisation identifier such as a virtual machine identifier may be restricted to being updated by instructions executing in an operating state with at least a hypervisor-level of privilege (more privileged than the operating-system-level of privilege), so that the application-level code and the operating system (kernel) cannot update the virtualisation identifier but the hypervisor can update the virtualisation identifier. A security state identifier may be restricted to being updated by instructions executing in an operating state with a level of privilege sufficient to control a change of security states (e.g. this may be restricted to the most privileged operating state with a level of privilege even higher than the hypervisor-level-privileged operating state). The access control permissions can be enforced in different ways. For any of the identifier registers 7, 23 which are implemented as system registers updatable in response to a system register updating instruction executed by the processing circuitry 6, the processing circuitry 6 may detect attempts to execute a system register updating instruction which tries to update a given identifier register 7, 23 when the current operating state has insufficient privilege to update that register, and can generate an exception if the access should be denied. In this case, the access permissions for a given register are represented using hardwired circuit logic of the processing circuitry 6. On the other hand, for any of the identifier registers 7, 23 which are implemented as memory-mapped registers updated in response to a store instruction specifying a target address mapped to that register, the access permissions controlling whether the update is permitted may be defined in translation table structures (page table structures) set by software which are checked by memory management hardware of the processing circuitry 6 to determine whether a given memory access operation is permitted. In this case, while the hardware comprises memory access checking circuitry for enforcing the rules defined by software using the translation table structures, the specific access permissions associated with the context identifier registers 7, 23 may not be pre-defined in circuit hardware, but can be set at runtime by software programming the translation table structures to implement the required access controls for the addresses mapped to the identifier registers 7, 23.
[0074] Figure 1 shows an example with a single instance of ancillary circuitry 116 corresponding to the processing circuitry 6. It is also possible to have multiple instances of ancillary circuitry 116 associated with the same processing circuitry 6 (the two or more instances of ancillary circuitry116 could be multiple instances of the same type of ancillary circuitry, e.g. two or more accelerators, or could include instances of different types of ancillary circuitry, e.g. one network interface controller and one accelerator). If multiple instances of ancillary circuitry 116 are accessible via the control interface circuitry 14, then there may be multiple instances of the set of one or more ancillary context identifier registers 23, one set for each respective instance of the ancillary circuitry 116, so that each ancillary circuitry 116 can have its associated current ancillary context identified by the corresponding ancillary context identifier registers 23. In that case, the particular ancillary context identifier registers 23 used in a comparison by the comparator 4 depends on which ancillary circuitry 116 is the target ancillary circuitry for which a control action is requested by the processing circuitry 6 (e.g. an identifier specified in the request sent to the control interface circuitry 14 by the processing circuitry 6 may identify which ancillary circuit instance is the target ancillary circuitry).
[0075] Figure 2 illustrates a method for controlling availability of the control interface circuitry 14. At step 200, at least one processing context identifier is stored in the processing context identifier storage circuitry 7, to identify a current processing context assigned to the processing circuitry 6. At step 202, at least one ancillary context identifier is stored in the ancillary context identifier storage circuitry 23, to identify a current ancillary context assigned to the ancillary circuitry 116. At step 204, the interface availability control circuitry (comparator) 4 compares the at least one processing context identifier and the at least one ancillary context identifier, and at step 206, controls availability of the control interface circuitry to the current processing context associated with the processing circuitry depending on the comparison.
[0076] Figure 3 shows more detail for the availability control. At step 220, the processing circuitry executes an instruction to request that the control interface circuitry 14 performs a control action for the ancillary circuitry 116. For example, the control action could be launching a command to the ancillary circuitry 116 (e.g. a command for offloading a new task to the ancillary circuitry 116), or returning information on current status of the ancillary circuitry 116. At step 222, the interface availability control circuitry 4 compares the processing context identifier(s) stored in registers 7 with the corresponding ancillary context identifier(s) stored in registers 23, and at step 224 determines whether the comparison detects a match or mismatch between the compared identifiers. If a match is detected, then at step 226 the interface availability control circuitry 4 determines that the control interface circuitry 14 is available, and allows the control action to be performed. If a mismatch is detected, then at step 228 interface availability control circuitry for determines that the control interface circuitry 14 is not available, does not allow the control action to be performed, and returns an unavailability response to the processing circuitry (e.g. by setting an unavailability status value in a control register or by raising an exception). By preventing the control action being performed, this prevents status of the ancillary circuitry associated with one context being leaked to a different context which is active on the processing circuitry, and prevents one context being able to interfere with the activity of the ancillary circuitry associated with anothercontext, e.g. by launching commands to the ancillary circuitry when the ancillary circuitry is part way through a task for the other context. The unavailability response may, if necessary, prompt software (e.g. a supervisory process) to reallocate the ancillary circuitry 116 for use by the context which tried to perform the control action but received the unavailability response. Alternatively, the supervisory process may choose to allow the ancillary circuitry 116 to continue with its task on behalf of the other context, and then only switch the current ancillary context of the ancillary circuitry 116 later once that task is complete. The particular response taken to the unavailability response is a choice for the software developer and is not a required feature of the hardware apparatus 2.
[0077] Figure 4 illustrates a method for controlling whether the ancillary implement interrupt is masked. Steps 200 to 204 are the same as in Figure 2. At step 240, based on the comparison of processing / ancillary context identifiers performed at step 204, the control interface circuitry 14 determines whether the ancillary event interrupt is masked. At step 242, in response to detecting occurrence of an event associated with the ancillary circuitry 116 (e.g. the ancillary circuitry signalling a fault or indicating its task is complete) when the ancillary event interrupt is not masked, the control interface circuitry 14 signals the ancillary event interrupt to the processing circuitry 6. Signalling of the ancillary event interrupt to the processing circuitry 6 is suppressed (even if the ancillary event is detected) if the ancillary event interrupt is currently masked.
[0078] Figure 5 illustrates more detail for the ancillary event interrupt masking. At step 250, an event associated with the ancillary circuitry 116 occurs. At step 252, the comparator 4 and / or control interface circuitry 14 compares the corresponding processing context identifier(s) and ancillary context identifier(s) stored in the registers 7, 23. At step 254, the control interface circuitry 14 determines whether a match or mismatch is detected in the comparison of context identifiers. At step 256, if matching processing / ancillary context identifiers are detected in the comparison, the ancillary event interrupt is masked and so no ancillary event interrupt is signalled the processing circuitry 6 in response to the event associated with the ancillary circuitry 116. At step 258, if mismatching processing / ancillary context identifiers are detected, the ancillary event interrupt is enabled (not masked) and so the ancillary event interrupt is signalled to the processing circuitry 6 in response to the event detected at step 258.
[0079] Note that in some examples, whether or not the ancillary event interrupt is masked could also depend on further parameters, other than the comparison between context identifiers. For example, the control interface circuitry 14 could also support additional masking parameters that can be set by software to control masking of the ancillary event interrupt. For example, software could override the control based on the comparator 4, by setting a parameter that specifies that the ancillary event interrupt should be masked (disabled) regardless of whether there is a match or mismatch of the context identifiers compared by the comparator 4. Therefore, the hardware is configured to set the masked / non-masked status of the ancillary event interrupt based on the comparison of context identifiers (in the sense that the hardware has circuitry that supports thiscontrol operation), but it is not essential that this comparison always is relevant - whether or not the comparison of context identifiers actually influences masking of the ancillary event interrupt at a given time may depend on other control parameters set by software at runtime.
[0080] Figure 6 shows in more detail a specific example an apparatus 112 in which the ancillary circuitry 116 is a processor-local hardware accelerator 116 associated with a given processor (CPU, central processing unit) 114 comprising the processing circuitry 6 and control interface circuitry 14 described earlier. For example, the accelerator 116 could be a machine learning accelerator, such as an NPU, for accelerating machine learning workloads such as neural network processing. While Figure 6 shows a single accelerator 116 coupled to the CPU 114, other examples could provide more than one accelerator 116 per CPU 114, with multiple accelerators 116 coupled to the same CPU 114 via the accelerator control interface circuitry 14 (in this case, as mentioned above, the interface registers 23 of the control interface circuitry 14 may comprise multiple sets of ancillary context identifier registers).
[0081] The CPU 114 comprises the processing circuitry 6 which decodes and executes instructions defined in an instruction set architecture (ISA) to carry out data processing operations represented by the instructions. The processing circuitry 6 performs operations on data loaded from a memory system, and may store the results back to the memory system. In this example the memory system includes a level one cache 10, a level two cache 20, and memory (not shown in Figure 6, but accessible via a memory system interconnect interface 119 of the CPU 114). However, it will be appreciated that this is just one example of a possible memory hierarchy and other implementations can have further levels of cache or a different arrangement. For example, separate level one caches 10 may be provided for instructions and data. The provision of caches 10, 20 within the CPU 114 enables faster access to data than from memory (which can include on-chip and / or off-chip memory).
[0082] The CPU 114 also comprises a memory management unit 16 (MMU), to perform address translation in response to memory access (load / store) instructions executed by the processing circuitry. The MMU 16 translates virtual addresses specified by the operands of the load / store instructions into physical addresses identifying storage locations of data in the memory system. The MMU 16 has a translation lookaside buffer (TLB) 18 for caching address translation data from page tables stored in the memory system, where the page table entries of the page tables define address translation mappings and may also specify access permissions which govern whether a given process executing on the pipeline is allowed to read data, write data and / or execute instructions associated with a particular memory address.
[0083] The apparatus 112 includes one or more hardware accelerators 116 configurable, based on instructions executed by the processing circuitry 6, to perform a delegated task asynchronously with respect to operations performed by the processing circuitry 6 of the CPU 114 in response to executed instructions.The hardware accelerator 116 is unique (private) to a single CPU (core) 114, and therefore may be referred to as a core local accelerator (CLA). The hardware accelerator 116 is controlled by an associated processor core 114. The CPU 114 comprises accelerator control interface circuitry 14 to exchange control signals with the hardware accelerator 116 via a signal path 117 (e.g. bus) distinct from the memory system interconnect interface 119 by which memory access requests from the CPU 114 are issued to system memory.
[0084] In this example, the accelerator 116 communicates with the memory system via the same path used by the processing circuitry 6. The hardware accelerator 116 can issue accelerator-triggered memory access requests specifying virtual addresses. In response to an accelerator-triggered memory access request received at the accelerator control interface circuitry 14 from the hardware accelerator 114, the MMU 16 of the CPU 114 translates a virtual address specified by the accelerator-triggered memory access request to a physical address of a memory system location to be accessed in response to the accelerator-triggered memory access request. Hence, the hardware accelerator reuses the memory management circuitry 16 of the CPU 114 for address translation. The MMU 16 may translate the virtual address of an accelerator-triggered memory access request according to address mapping information associated with the virtual address and a given address translation context (e.g. an address translation context corresponding to the current ancillary context represented by the ancillary context identifier registers 23). While not shown in Figure 6, the accelerator 116 could have its own internal cache, e.g. a separate level 1 cache separate from the level 1 cache of the processing circuitry 10.
[0085] In this example, the accelerator 116 accesses memory via the CPU’s memory access interface, and does not have a direct path to memory separate from the CPU 114. However, other examples may provide a direct access path from the accelerator 116 to memory, separate from the memory system interconnect interface 119 by which the processing circuitry 6 accesses memory. The accelerator 116 could have some memory management circuitry for translation of virtual to physical addresses, but the memory management circuitry could be more limited than the MMU 16 of the CPU 114. For example, the memory management circuitry of the accelerator 116 could comprise a TLB used to cache translation information from page tables and associated TLB lookup logic which can obtain the physical address corresponding to a virtual of an accelerator-triggered access from the accelerator TLB in cases where the required virtual address hits in the TLB, but which on a miss sends a request to the MMU 16 of the CPU 114 to control the MMU’s page table walk circuitry 19 to perform a page table walk operation to obtain the translation mappings to be cached in the accelerator’s TLB for future use. Hence, even if the accelerator 116 has some address translation capability, it may still share the page table walk circuitry 19 of the CPU 114 (and the page table walk circuitry 19 does not need to be replicated at the accelerator 116).
[0086] Address translation faults arising from translation of accelerator-triggered memory accesses may be reported to the CPU 114, either by raising an exception (e.g. a translation faultcould be an example of an event which may trigger the ancillary event interrupt as discussed earlier) or by reporting the fault more passively, e.g. by setting a status indication in a control register which can be polled later by software executing on the processing circuitry 6.
[0087] The accelerator 116 may have access to one or more private caches (caches which are not shared with any other CPU, e.g. the level 2 cache 20) of the CPU 114. This can allow more efficient sharing of data between the CPU 114 and accelerator 116 compared to sharing of data via memory.
[0088] The accelerator control interface 14 allows the accelerator 116 to be configurable using instructions executed by the CPU 114 in an operating state having user-level privilege (the lowest level of privilege granted to user-level application code), rather than requiring the accelerator to be configured only by more privileged code such as an operating system or hypervisor. This can improve performance by reducing the need for application-level code to call into an operating system or hypervisor when it needs the accelerator to perform delegated tasks.
[0089] In some examples, the processing circuitry 6 may support execution of instructions of an ISA providing a class of accelerator control instructions, separate from load / store instructions, for controlling the accelerator interface circuitry 14 to perform functions such as launching accelerator commands, checking on accelerator status, reading internal accelerator state, writing other accelerator control registers, etc.
[0090] However, in other examples, the CPU 114 may comprise memory-mapped register storage 23 accessible in response to load / store instructions executed by the processing circuitry 6 specifying target addresses mapped to the memory-mapped register storage. Hence, accelerator commands may be triggered by execution of load / store instructions which specify addresses mapped to the memory-mapped register storage 23. These memory-mapped registers can include the ancillary context identifier registers 23 mentioned earlier, as well as including additional registers for storing information other than context identifiers. The CPU 114 (via the accelerator interface circuitry 14) may control operation of the at least one hardware accelerator 116 by writing to and reading from the memory-mapped register storage. Hence, the processing circuitry 6 can control operation of a hardware accelerator 116 using conventional load / store instructions (with the address of the load / store instructions distinguishing accelerator control instructions from other load / store instructions targeting data storage locations in the memory system).
[0091] Figure 7 schematically illustrates an example set of memory-mapped control registers 23 provided in CPU 114 for controlling operation of one or more core local hardware accelerators 116. These include the ancillary context identifier registers 23 mentioned earlier. As shown in Figure 7, there could be multiple sets of such ancillary context identifier registers 23, if there are multiple accelerators 116 associated with the same CPU 114. It will be appreciated that further control registers not illustrated in Figure 7 could also be provided. The physical address of each memory-mapped register can be derived by combining a base address representing the start ofa control structure mapped to the registers 23 with an offset associated with a particular memorymapped control register. The base address may be a programmable parameter of the control interface circuitry 14.
[0092] The set of memory-mapped control registers 23 comprises a set of (e.g. eight) data port registers 400 (DATA). The DATA registers are used to store input and output parameters for control commands communicated between the CPU 114 and the hardware accelerator 116. The DATA registers 400 are provided to enable parameters to be specified by software for control commands for controlling the hardware accelerator 116. Contents of the DATA registers 400 may be transferred to the accelerator 116 alongside launched accelerator commands transmitted via the control request channel, and certain commands may prompt the accelerator 116 to return parameters via the control response channel with the parameters returned over the bus 117 being written by the control interface 14 to the DATA registers 400 from which those parameters can be read by software executing on the CPU 114.
[0093] The set of memory-mapped control registers 23 also comprises a LAUNCH register 402. Processing circuitry 6 can cause accelerator control signals to be issued to a given hardware accelerator 116 by writing to the LAUNCH register 402, which triggers control circuitry 25 of the control interface circuitry 14 to generate the corresponding control signals sent to the accelerator 116 on the accelerator control bus 117. Writing different values to the LAUNCH register indicates that the processing circuitry 6 requests the hardware accelerator control interface circuitry 14 to initiate different operations for performance by the hardware accelerator 116 (e.g. a field within the LAUNCH register 402 may be encoded to represent the type of command being instructed). For example, command encodings may include a command which requests that the accelerator starts a new delegated task, a register read I write command for reading or writing accelerator registers provided within accelerator 116, pause / resume commands to instruct the accelerator to pause its current task or resume after previously pausing, or save / restore commands for instructing the accelerator to save internal state information to memory or restore previously saved internal state from memory. It will be appreciated that the particular commands supported may vary depending on implementation. Also, it will be appreciated that in some cases the command encodings in the launch register 402 may be generic to a wide variety of specific hardware accelerator implementations (e.g. the launch register may support a generic “command launch” command), and the actual implementation-specific commands to a particular implementation of an accelerator may be encoded using the contents of the DATA registers 400 which are to be transmitted on bus 117 as parameters alongside a launch command.
[0094] The set of memory-mapped control registers 23 also comprises a launch response LRESP register 404. The LRESP register 404 is used to indicate a response to a previous write to the LAUNCH register 402. For example, the LRESP register 404 may specify a response pending field used to indicate whether the accelerator is still to respond to a previous command launched via the LAUNCH register 402. A response is pending if an operation has been signalled to agiven hardware accelerator but a response to that signal has not yet been received. If software polls the LRESP register when the pending indication is set, this may indicate that the software should try again later as the contents of the other fields cannot be relied on. Other fields of the register may provide status codes indicating the status of the previously launched command, such as an indication of whether any error occurred, whether a timeout was detected where the accelerator did not respond within a given time period, or whether the accelerator is currently unavailable (e.g. because the accelerator is busy carrying out a task for another software process running on the CPU 114). If the response to an accelerator command indicated by the LRESP register indicates that the command has not been accepted, the software executing on the CPU 114 may retry the accelerator command later. If the accelerator command has successfully been accepted, then the software on the CPU 114 can stop polling the LRESP register and await completion of the task offloaded to the accelerator, which is completed asynchronously by the accelerator (so the instruction which caused the accelerator command to be issued can commit on the processing pipeline of the CPU 114 without waiting for completion of the offloaded task).
[0095] The set of memory-mapped control registers 23 also comprises a set of status reporting registers STATUS[0:7] 414. Unlike the registers 400, 402, 404 which are shared between hardware accelerators, the STATUS registers are each unique to a particular hardware accelerator 116 (hence if there is only one hardware accelerator 116, only a single status register 414 could be provided). Each STATUS register 414 is used to report information about a corresponding hardware accelerator 116 to the CPU 114. For example, the status information may include an indication of whether the accelerator is idle (which can be an indication that a previously offloaded task has completed), whether the accelerator is ready to accept further commands, whether a memory translation fault has been detected during the handling of a previously accepted accelerator command, etc. Hence, the software on the CPU 114 can poll the status register 414 for a given accelerator 116 to identify when the task assigned to that accelerator 116 is complete or to identify errors which have arisen during processing of the task. Once the task is complete, the data processed by the accelerator 116 can be retrieved from memory (either by the CPU 114 itself, or by another CPU of the system).
[0096] As shown in Figure 8, the physical signal paths 117 between the CPU 114 and accelerator 116 may include a number of communication channels, which may include (at least) two groups of channels: control channels and memory interface channels. The memory interface channels comprise a read address channel (RD_AR), a read data channel (RD_R), a write address channel (WR_AW), a write data channel (WR_W), and a write response channel (WR_B). In some examples, multiple read and / or write channels may be supported, and hence for example two or more copies of the RD_AR and RD_R channels may be provided, and so on. For example, the memory interface channels may be implemented according to the AXI protocol provided by Arm® Limited. The memory interface channels may be used to convey accelerator-triggered memory access requests from the accelerator 116 to the CPU 114 and responses to those requests fromthe CPU 114 to the accelerator 116. On the other hand, the control channels are used for launching accelerator commands to the accelerator 116, checking accelerator status, etc., or for any other request / response not related to an accelerator-triggered access to memory.
[0097] Concepts described herein may be embodied in a system comprising at least one packaged chip. The apparatus described earlier is implemented in the at least one packaged chip (either being implemented in one specific chip of the system, or distributed over more than one packaged chip). The at least one packaged chip is assembled on a board with at least one system component. A chip-containing product may comprise the system assembled on a further board with at least one other product component. The system or the chip-containing product may be assembled into a housing or onto a structural support (such as a frame or blade).
[0098] As shown in Figure 9, one or more packaged chips 900, with the apparatus described above implemented on one chip or distributed over two or more of the chips, are manufactured by a semiconductor chip manufacturer. In some examples, the chip product 900 made by the semiconductor chip manufacturer may be provided as a semiconductor package which comprises a protective casing (e.g. made of metal, plastic, glass or ceramic) containing the semiconductor devices implementing the apparatus described above and connectors, such as lands, balls or pins, for connecting the semiconductor devices to an external environment. Where more than one chip 900 is provided, these could be provided as separate integrated circuits (provided as separate packages), or could be packaged by the semiconductor provider into a multi-chip semiconductor package (e.g. using an interposer, or by using three-dimensional integration to provide a multi-layer chip product comprising two or more vertically stacked integrated circuit layers).
[0099] In some examples, a collection of chiplets (i.e. small modular chips with particular functionality) may itself be referred to as a chip. A chiplet may be packaged individually in a semiconductor package and / or together with other chiplets into a multi-chiplet semiconductor package (e.g. using an interposer, or by using three-dimensional integration to provide a multilayer chiplet product comprising two or more vertically stacked integrated circuit layers).
[0100] The one or more packaged chips 900 are assembled on a board 402 together with at least one system component 904 to provide a system 906. For example, the board may comprise a printed circuit board. The board substrate may be made of any of a variety of materials, e.g. plastic, glass, ceramic, or a flexible substrate material such as paper, plastic or textile material. The at least one system component 904 comprise one or more external components which are not part of the one or more packaged chip(s) 900. For example, the at least one system component 904 could include, for example, any one or more of the following: another packaged chip (e.g. provided by a different manufacturer or produced on a different process node), an interface module, a resistor, a capacitor, an inductor, a transformer, a diode, a transistor and / or a sensor.A chip-containing product 916 is manufactured comprising the system 906 (including the board 902, the one or more chips 900 and the at least one system component 904) and one or more product components 912. The product components 912 comprise one or more further components which are not part of the system 906. As a non-exhaustive list of examples, the one or more product components 912 could include a user input / output device such as a keypad, touch screen, microphone, loudspeaker, display screen, haptic device, etc.; a wireless communication transmitter / receiver; a sensor; an actuator for actuating mechanical motion; a thermal control device; a further packaged chip; an interface module; a resistor; a capacitor; an inductor; a transformer; a diode; and / or a transistor. The system 906 and one or more product components 912 may be assembled on to a further board 914.
[0101] The board 902 or the further board 914 may be provided on or within a device housing or other structural support (e.g. a frame or blade) to provide a product which can be handled by a user and / or is intended for operational use by a person or company.
[0102] The system 906 or the chip-containing product 916 may be at least one of: an end-user product, a machine, a medical device, a computing or telecommunications infrastructure product, or an automation control system. For example, as a non-exhaustive list of examples, the chipcontaining product could be any of the following: a telecommunications device, a mobile phone, a tablet, a laptop, a computer, a server (e.g. a rack server or blade server), an infrastructure device, networking equipment, a vehicle or other automotive product, industrial machinery, consumer device, smart card, credit card, smart glasses, avionics device, robotics device, camera, television, smart television, DVD players, set top box, wearable device, domestic appliance, smart meter, medical device, heating / lighting control device, sensor, and / or a control system for controlling public infrastructure equipment such as smart motorway or traffic lights.
[0103] Concepts described herein may be embodied in computer-readable code for fabrication of an apparatus that embodies the described concepts. For example, the computer-readable code can be used at one or more stages of a semiconductor design and fabrication process, including an electronic design automation (EDA) stage, to fabricate an integrated circuit comprising the apparatus embodying the concepts. The above computer-readable code may additionally or alternatively enable the definition, modelling, simulation, verification and / or testing of an apparatus embodying the concepts described herein.
[0104] For example, the computer-readable code for fabrication of an apparatus embodying the concepts described herein can be embodied in code defining a hardware description language (HDL) representation of the concepts. For example, the code may define a register-transfer-level (RTL) abstraction of one or more logic circuits for defining an apparatus embodying the concepts. The code may define a HDL representation of the one or more logic circuits embodying the apparatus in Verilog, SystemVerilog, Chisel, or VHDL (Very High-Speed Integrated Circuit Hardware Description Language) as well as intermediate representations such as FIRRTL. Computer-readable code may provide definitions embodying the concept using system-levelmodelling languages such as SystemC and SystemVerilog or other behavioural representations of the concepts that can be interpreted by a computer to enable simulation, functional and / or formal verification, and testing of the concepts.
[0105] Additionally or alternatively, the computer-readable code may define a low-level description of integrated circuit components that embody concepts described herein, such as one or more netlists or integrated circuit layout definitions, including representations such as GDSII. The one or more netlists or other computer-readable representation of integrated circuit components may be generated by applying one or more logic synthesis processes to an RTL representation to generate definitions for use in fabrication of an apparatus embodying the invention. Alternatively or additionally, the one or more logic synthesis processes can generate from the computer-readable code a bitstream to be loaded into a field programmable gate array (FPGA) to configure the FPGA to embody the described concepts. The FPGA may be deployed for the purposes of verification and test of the concepts prior to fabrication in an integrated circuit or the FPGA may be deployed in a product directly.
[0106] The computer-readable code may comprise a mix of code representations for fabrication of an apparatus, for example including a mix of one or more of an RTL representation, a netlist representation, or another computer-readable definition to be used in a semiconductor design and fabrication process to fabricate an apparatus embodying the invention. Alternatively or additionally, the concept may be defined in a combination of a computer-readable definition to be used in a semiconductor design and fabrication process to fabricate an apparatus and computer-readable code defining instructions which are to be executed by the defined apparatus once fabricated.
[0107] Such computer-readable code can be disposed in any known transitory computer-readable medium (such as wired or wireless transmission of code over a network) or non-transitory computer-readable medium such as semiconductor, magnetic disk, or optical disc. An integrated circuit fabricated using the computer-readable code may comprise components such as one or more of a central processing unit, graphics processing unit, neural processing unit, digital signal processor or other components that individually or collectively embody the concept.
[0108] Examples are set out in the following clauses:
[0109] A1. An apparatus comprising:
[0110] processing circuitry configured to process instructions of a current processing context identified by at least one processing context identifier stored in processing context identifier storage circuitry;
[0111] control interface circuitry configured to exchange control signals with ancillary circuitry configurable to perform an offloaded function on behalf of the processing circuitry, where the ancillary circuitry is capable of performing, in parallel with processing of the current processing context on the processing circuitry, the offloaded function for a non-resident execution context that is different to the current processing context;ancillary context identifier storage circuitry configured to store at least one ancillary context identifier indicative of a current ancillary context associated with the offloaded function allocated to the ancillary circuitry; and
[0112] interface availability control circuitry configured to control whether the control interface circuitry is available to the current processing context depending on a comparison between the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one ancillary context identifier stored in the ancillary context identifier storage circuitry.
[0113] A2. The apparatus according to clause A1, in which the interface availability control circuitry is configured to determine that the control interface circuitry is unavailable to the current processing context in response to the comparison detecting a mismatch between the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one ancillary context identifier stored in the ancillary context identifier storage circuitry. A3. The apparatus according to any of clauses A1 and A2, in which:
[0114] when the control interface circuitry is determined to be unavailable to the current processing context, the interface availability control circuitry is configured to prevent an instruction executed by the processing circuitry being able to cause the control interface circuitry to perform at least one control action for the ancillary circuitry.
[0115] A4. The apparatus according to any of clauses A1 to A3, in which, in response to an attempt by an instruction executed by the processing circuitry to cause the control interface circuitry to perform a control action for the ancillary circuitry when the control interface circuitry is determined to be unavailable to the current processing context, the interface availability control circuitry is configured to cause an unavailability response to be returned to the processing circuitry.
[0116] A5. The apparatus according to clause A4, in which the unavailability response comprises setting an unavailable indication in a storage location pollable by instructions executed by the processing circuitry.
[0117] A6. The apparatus according to clause A4, in which the unavailability response comprises an exception being signalled to the processing circuitry.
[0118] A7. The apparatus according to any of clauses A1 to A6, in which the ancillary context identifier storage circuitry comprises at least one memory-mapped register.
[0119] A8. The apparatus according to any of clauses A1 to A7, in which the processing context identifier storage circuitry comprises at least one memory-mapped register.
[0120] A9. The apparatus according to any of clauses A1 to A8, in which the ancillary circuitry comprises a hardware accelerator configured to perform a delegated task asynchronously with respect to operations performed by the processing circuitry.
[0121] A10. The apparatus according to clause A9, in which the hardware accelerator comprises accelerator circuitry configured to accelerate operations for one or more machine learning workloads.A11. The apparatus according to any of clauses A9 and A10, in which the hardware accelerator comprises a processor-local accelerator associated with the processing circuitry.
[0122] A12. The apparatus according to any of clauses A9 to A11 , in which the hardware accelerator is private to the processing circuitry.
[0123] A13. The apparatus according to any of clauses A9 to A12, in which the processing circuitry and the hardware accelerator share translation table walk circuitry configured to control translation table walk operations for obtaining address translation table entries from memory. A14. The apparatus according to any of clauses A9 to A13, in which the processing circuitry and the hardware accelerator share at least one private cache.
[0124] A15. The apparatus according to any of clauses A9 to A14, in which the accelerator control interface circuitry is configurable to launch a new delegated task to the hardware accelerator in response to one or more instructions executed by a processing circuitry in an operating state with user-level privilege.
[0125] A16. The apparatus according to any of clauses A1 to A15, in which the control interface circuitry comprises memory-mapped register storage accessible to the processing circuitry in response to load / store instructions executed by the processing circuitry specifying target addresses mapped to the memory-mapped register storage.
[0126] A17. The apparatus according to any of clauses A1 to A16, comprising the ancillary circuitry. A18. A system comprising:
[0127] the apparatus according to any of clauses A1 to A17, implemented in at least one packaged chip;
[0128] at least one system component; and
[0129] a board,
[0130] wherein the at least one packaged chip and the at least one system component are assembled on the board.
[0131] A19. A chip-containing product comprising the system of clause A18, wherein the system is assembled on a further board with at least one other product component.
[0132] A20. Computer-readable code for fabrication of an apparatus according to any of clauses A1 to A17.
[0133] A21. A method comprising:
[0134] storing, in processing context identifier storage circuitry, at least one processing context identifier indicative of a current processing context for which instructions are processed by processing circuitry;
[0135] storing, in ancillary context identifier storage circuitry, at least one ancillary context identifier indicative of a current ancillary context associated with ancillary circuitry configurable to perform an offloaded function on behalf of the processing circuitry;
[0136] depending on a comparison between the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one ancillary context identifierstored in the ancillary context identifier storage circuitry, determining whether control interface circuitry for exchanging control signals with the ancillary circuitry is available to the current processing context associated with the processing circuitry.
[0137] A22. An apparatus comprising:
[0138] accelerator control interface circuitry configured to exchange control signals with a hardware accelerator, the accelerator control interface circuitry being configurable, based on instructions executed by a processor, to control the hardware accelerator to perform a delegated task asynchronously with respect to operations performed by the processor;
[0139] processing context identifier storage circuitry configured to store at least one processing context identifier indicative of a current processing context associated with the processor;
[0140] accelerator context identifier storage circuitry configured to store at least one accelerator context identifier indicative of a current accelerator execution context associated with the hardware accelerator; and
[0141] interface availability control circuitry configured to control whether the accelerator control interface circuitry is available to the current processing context depending on a comparison between the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one accelerator context identifier stored in the accelerator context identifier storage circuitry.
[0142] B1. An apparatus comprising:
[0143] processing circuitry configured to process instructions of a current processing context identified by at least one processing context identifier stored in processing context identifier storage circuitry;
[0144] control interface circuitry configured to exchange control signals with ancillary circuitry configurable to perform an offloaded function on behalf of the processing circuitry, where the ancillary circuitry is capable of performing, in parallel with processing of the current processing context on the processing circuitry, the offloaded function for a non-resident execution context that is different to the current processing context;
[0145] ancillary context identifier storage circuitry configured to store at least one ancillary context identifier indicative of a current ancillary context associated with the offloaded function allocated to the ancillary circuitry; and
[0146] wherein the control interface circuitry is configured to:
[0147] determine, based on a comparison between the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one ancillary context identifier stored in the ancillary context identifier storage circuitry, whether an ancillary event interrupt is masked, the ancillary event interrupt comprising an interrupt for reporting to the processing circuitry an event associated with the ancillary circuitry; andin response to detecting the event associated with the ancillary circuitry when the ancillary event interrupt is not masked, signal the ancillary event interrupt to the processing circuitry.
[0148] B2. The apparatus according to clause B1 , in which the control interface circuitry is configured to determine that the ancillary event interrupt is masked, in response to detecting a match between the at least one processing context identifier and the at least one ancillary context identifier.
[0149] B3. The apparatus according to any of clauses B1 and B2, in which the control interface circuitry is configured to determine that the ancillary event interrupt is not masked in response to detecting a mismatch between the at least one processing context identifier and the at least one ancillary context identifier.
[0150] B4. The apparatus according to any of clauses B1 to B3, in which the control interface circuitry is configured to detect the event associated with the ancillary circuitry based on an event signal received from the ancillary circuitry.
[0151] B5. The apparatus according to any of clauses B1 to B4, in which the ancillary event interrupt comprises a hint to software that an event associated with the ancillary circuitry has occurred which relates to a non-resident processing context which differs from the current processing context for which instructions are being processed by the processing circuitry.
[0152] B6. The apparatus according to any of clauses B1 to B5, in which the ancillary context identifier storage circuitry comprises at least one memory-mapped register.
[0153] B7. The apparatus according to any of clauses B1 to B6, in which the processing context identifier storage circuitry comprises at least one memory-mapped register.
[0154] B8. The apparatus according to any of clauses B1 to B7, in which the ancillary circuitry comprises a hardware accelerator configured to perform a delegated task asynchronously with respect to operations performed by the processing circuitry.
[0155] B9. The apparatus according to clause B8, in which the hardware accelerator comprises accelerator circuitry configured to accelerate operations for one or more machine learning workloads.
[0156] B10. The apparatus according to any of clauses B8 and B9, in which the hardware accelerator comprises a processor-local accelerator associated with the processing circuitry.
[0157] B11. The apparatus according to any of clauses B8 to B10, in which the hardware accelerator is private to the processing circuitry.
[0158] B12. The apparatus according to any of clauses B8 to B11, in which the processing circuitry and the hardware accelerator share translation table walk circuitry configured to control translation table walk operations for obtaining address translation table entries from memory. B13. The apparatus according to any of clauses B8 to B12, in which the processing circuitry and the hardware accelerator share at least one private cache.B14. The apparatus according to any of clauses B9 to B14, in which the control interface circuitry is configurable to launch a new delegated task to the hardware accelerator in response to one or more instructions executed by a processing circuitry in an operating state with user-level privilege.
[0159] B15. The apparatus according to any of clauses B1 to B14, in which the control interface circuitry comprises memory-mapped register storage accessible to the processing circuitry in response to load / store instructions executed by the processing circuitry specifying target addresses mapped to the memory-mapped register storage.
[0160] B16. The apparatus according to any preceding clause, comprising the ancillary circuitry. B17. A system comprising:
[0161] the apparatus according to any of clauses B1 to B16, implemented in at least one packaged chip;
[0162] at least one system component; and
[0163] a board,
[0164] wherein the at least one packaged chip and the at least one system component are assembled on the board.
[0165] B18. A chip-containing product comprising the system of clause B17, wherein the system is assembled on a further board with at least one other product component.
[0166] B19. Computer-readable code for fabrication of an apparatus according to any of clauses B1 to B16.
[0167] B20. A method comprising:
[0168] storing, in processing context identifier storage circuitry, at least one processing context identifier indicative of a current processing context for which instructions are processed by processing circuitry;
[0169] storing, in ancillary context identifier storage circuitry, at least one ancillary context identifier indicative of a current ancillary context associated with ancillary circuitry configurable to perform an offloaded function on behalf of the processing circuitry;
[0170] depending on a comparison between the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one ancillary context identifier stored in the ancillary context identifier storage circuitry, determining whether an ancillary event interrupt is masked, the ancillary event interrupt comprising an interrupt for reporting to the processing circuitry an event associated with the ancillary circuitry; and
[0171] in response to detecting the event associated with the ancillary circuitry when the ancillary event interrupt is not masked, signalling the ancillary event interrupt to the processing circuitry. B21. An apparatus comprising:
[0172] accelerator control interface circuitry configured to exchange control signals with a hardware accelerator, the accelerator control interface circuitry being configurable, based oninstructions executed by a processor, to control the hardware accelerator to perform a delegated task asynchronously with respect to operations performed by the processor;
[0173] processing context identifier storage circuitry configured to store at least one processing context identifier indicative of a current processing context associated with the processor;
[0174] accelerator context identifier storage circuitry configured to store at least one accelerator context identifier indicative of a current accelerator execution context associated with the hardware accelerator; and
[0175] wherein the accelerator control interface circuitry is configured to:
[0176] determine, based on a comparison between the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one accelerator context identifier stored in the accelerator context identifier storage circuitry, whether an accelerator event interrupt is masked, the accelerator event interrupt comprising an interrupt for reporting to the processing circuitry an event associated with the hardware accelerator; and
[0177] in response to detecting the event associated with the hardware accelerator when the accelerator event interrupt is not masked, signal the accelerator event interrupt to the processing circuitry.
[0178] In the present application, the words “configured to...” are used to mean that an element of an apparatus has a configuration able to carry out the defined operation. In this context, a “configuration” means an arrangement or manner of interconnection of hardware or software. For example, the apparatus may have dedicated hardware which provides the defined operation, or a processor or other processing device may be programmed to perform the function. “Configured to” does not imply that the apparatus element needs to be changed in any way in order to provide the defined operation.
[0179] In the present application, lists of features preceded with the phrase “at least one of” mean that any one or more of those features can be provided either individually or in combination. For example, “at least one of: [A], [B] and [C]” encompasses any of the following options: A alone (without B or C), B alone (without A or C), C alone (without A or B), A and B in combination (without C), A and C in combination (without B), B and C in combination (without A), or A, B and C in combination.
[0180] Although illustrative embodiments of the invention have been described in detail herein with reference to the accompanying drawings, it is to be understood that the invention is not limited to those precise embodiments, and that various changes and modifications can be effected therein by one skilled in the art without departing from the scope of the invention as defined by the appended claims.
Claims
CLAIMS1. An apparatus comprising:processing circuitry configured to process instructions of a current processing context identified by at least one processing context identifier stored in processing context identifier storage circuitry;control interface circuitry configured to exchange control signals with ancillary circuitry configurable to perform an offloaded function on behalf of the processing circuitry, where the ancillary circuitry is capable of performing, in parallel with processing of the current processing context on the processing circuitry, the offloaded function for a non-resident execution context that is different to the current processing context;ancillary context identifier storage circuitry configured to store at least one ancillary context identifier indicative of a current ancillary context associated with the offloaded function allocated to the ancillary circuitry; andwherein the control interface circuitry is configured to:determine, based on a comparison between the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one ancillary context identifier stored in the ancillary context identifier storage circuitry, whether an ancillary event interrupt is masked, the ancillary event interrupt comprising an interrupt for reporting to the processing circuitry an event associated with the ancillary circuitry; and in response to detecting the event associated with the ancillary circuitry when the ancillary event interrupt is not masked, signal the ancillary event interrupt to the processing circuitry.
2. The apparatus according to claim 1 , in which the control interface circuitry is configured to determine that the ancillary event interrupt is masked, in response to detecting a match between the at least one processing context identifier and the at least one ancillary context identifier.
3. The apparatus according to any of claims 1 and 2, in which the control interface circuitry is configured to determine that the ancillary event interrupt is not masked in response to detecting a mismatch between the at least one processing context identifier and the at least one ancillary context identifier.
4. The apparatus according to any preceding claim, in which the control interface circuitry is configured to detect the event associated with the ancillary circuitry based on an event signal received from the ancillary circuitry.
5. The apparatus according to any preceding claim, in which the ancillary event interrupt comprises a hint to software that an event associated with the ancillary circuitry has occurred which relates to a non-resident processing context which differs from the current processing context for which instructions are being processed by the processing circuitry.
6. The apparatus according to any preceding claim, in which the ancillary context identifier storage circuitry comprises at least one memory-mapped register.
7. The apparatus according to any preceding claim, in which the processing context identifier storage circuitry comprises at least one memory-mapped register.
8. The apparatus according to any preceding claim, in which the ancillary circuitry comprises a hardware accelerator configured to perform a delegated task asynchronously with respect to operations performed by the processing circuitry.
9. The apparatus according to claim 8, in which the hardware accelerator comprises accelerator circuitry configured to accelerate operations for one or more machine learning workloads.
10. The apparatus according to any of claims 8 and 9, in which the hardware accelerator comprises a processor-local accelerator associated with the processing circuitry.
11. The apparatus according to any of claims 8 to 10, in which the hardware accelerator is private to the processing circuitry.
12. The apparatus according to any of claims 8 to 11, in which the processing circuitry and the hardware accelerator share translation table walk circuitry configured to control translation table walk operations for obtaining address translation table entries from memory.
13. The apparatus according to any of claims 8 to 12, in which the processing circuitry and the hardware accelerator share at least one private cache.
14. The apparatus according to any of claims 9 to 14, in which the control interface circuitry is configurable to launch a new delegated task to the hardware accelerator in response to one or more instructions executed by a processing circuitry in an operating state with user-level privilege.
15. The apparatus according to any preceding claim, in which the control interface circuitry comprises memory-mapped register storage accessible to the processing circuitry in response toload / store instructions executed by the processing circuitry specifying target addresses mapped to the memory-mapped register storage.
16. The apparatus according to any preceding claim, comprising the ancillary circuitry.
17. A system comprising:the apparatus according to any preceding claim, implemented in at least one packaged chip;at least one system component; anda board,wherein the at least one packaged chip and the at least one system component are assembled on the board.
18. A chip-containing product comprising the system of claim 17, wherein the system is assembled on a further board with at least one other product component.
19. Computer-readable code for fabrication of an apparatus according to any of claims 1 to 16.
20. A method comprising:storing, in processing context identifier storage circuitry, at least one processing context identifier indicative of a current processing context for which instructions are processed by processing circuitry;storing, in ancillary context identifier storage circuitry, at least one ancillary context identifier indicative of a current ancillary context associated with ancillary circuitry configurable to perform an offloaded function on behalf of the processing circuitry;depending on a comparison between the at least one processing context identifier stored in the processing context identifier storage circuitry and the at least one ancillary context identifier stored in the ancillary context identifier storage circuitry, determining whether an ancillary event interrupt is masked, the ancillary event interrupt comprising an interrupt for reporting to the processing circuitry an event associated with the ancillary circuitry; andin response to detecting the event associated with the ancillary circuitry when the ancillary event interrupt is not masked, signalling the ancillary event interrupt to the processing circuitry.