Performance monitoring

By quantifying memory transaction interference, the performance monitoring circuitry in memory systems allows for optimized resource allocation, addressing the limitations of existing monitoring methods and enhancing system performance through targeted interference reduction.

GB2643579APending Publication Date: 2026-02-25ARM LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
GB2024012437
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2026-02-25

AI Technical Summary

Technical Problem

Existing performance monitoring in memory systems fails to provide resource control agents with sufficient information to determine whether performance issues are inherent to the workload or can be improved by reallocating resources, as utilization metrics fluctuate with workload variations and transaction types.

Method used

Performance monitoring circuitry determines interference-specific performance values that quantify the impact of memory transaction interference on affected transactions, allowing resource control agents to optimize resource allocation and reduce interference.

Benefits of technology

This approach enables more effective resource allocation by identifying and mitigating memory transaction interference, improving system performance by preventing potentially interfering transactions and reallocating resources based on actual interference causes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An apparatus comprises memory system component configured to handle memory transactions for accessing data. The apparatus also comprises performance monitoring circuitry configured to determine an int
Need to check novelty before this filing date? Find Prior Art

Description

The present technique relates to the field of data processing. More particularly, it relates to memory system performance monitoring. A data processing system may have memory system components such as a cache, interconnect or memory controller for handling memory transactions. The memory system components may have limited memory system resource for handling memory transactions. Resource control circuitry may control allocation of memory system resource to improve performance. The resource control circuitry may determine how to allocate memory system resource based on performance monitoring values provided by performance monitoring circuitry. The performance monitoring circuitry can provide useful information for controlling resource usage and / or diagnosing problems caused by excessive demand for limited system resource. At least some examples provide an apparatus, comprising: at least one memory system component configured to handle memory transactions for accessing data; and performance monitoring circuitry configured to determine an interference-specific performance value indicative of a performance impact associated with a given memory system component handling at least one interfering memory transaction; wherein: handling of an interfering memory transaction comprises use of memory system resource which would have otherwise been available for handling at least one affected memory transaction, and the performance monitoring circuitry is configured to make the interferencespecific performance value available to a resource control agent. At least some examples provide a method, comprising: handling memory transactions for accessing data; determining an interference-specific performance value indicative of a performance impact associated with a given memory system component handling at least one interfering memory transaction; and making the interference-specific performance value available to a resource control agent; wherein handling of an interfering memory transaction comprises use of memory system resource which would have otherwise been available for handling at least one affected memory transaction. At least some examples provide a computer program for controlling a host data processing apparatus to provide an instruction execution environment, the computer program comprising: memory system component program logic configured to handle memory transactions for accessing data; and performance monitoring program logic configured to determine an interference-specific performance value indicative of a performance impact associated with given memory system component program logic handling at least one interfering memory transaction; wherein: handling of an interfering memory transaction comprises use of memory system resource which would have otherwise been available for handling at least one affected memory transaction, and the performance monitoring program logic is configured to make the interference-specific performance value available to a resource control agent. Further aspects, features and advantages of the present technique will be apparent from the following description of examples, which is to be read in conjunction with the accompanying drawings, in which: Figure 1 illustrates an example of a processing system comprising performance monitoring circuitry; Figure 2 illustrates an example of a processing unit in a processing system; Figure 3 illustrates an example of determining an interference-specific performance value for a cache; Figure 4 illustrates an example of determining an interference-specific performance value for a queue; Figure 5 and Figure 6 illustrate an example set of task-associated interference-specific performance values; Figure 7 illustrates an example of monitoring circuitry which may be used to calculate a latency for a memory transaction issued by a task; Figure 8 is a flow diagram illustrating a method of using performance monitoring circuitry to calculate an interference-specific performance value; and Figure 9 illustrates a simulation example. An apparatus comprises at least one memory system component configured to handle memory transactions for accessing data. A memory transaction, or memory access request, may comprise, for example, a load request or a store request to load data from or store data to a memory system. The memory system component may comprise any component providing memory system resource involved in at least one stage of processing a memory transaction, such as a cache, interconnect, memory management unit (MMU), queue, buffer, or memory controller, for example. The apparatus also comprises performance monitoring circuitry. In previous techniques, performance monitoring circuitry may have been provided to monitor utilisation of a particular memory system resource, for example indicating how much of the total memory system resource (e.g., a total number of cache entries, or a total amount of bandwidth) is currently available. The inventors have however identified a problem that the information provided by previous performance monitors has had limited value. Simply reporting the utilisation of a particular memory system component may not provide a resource control agent with information which allows it to determine whether any changes can be made to improve performance. For example, the utilisation of a particular memory system component can fluctuate depending on intrinsic factors of a workload such as variations in the size of a working set, and the type and size of values being handled for a given memory transaction. A performance monitor reporting utilisation alone may therefore not enable a resource control agent to determine whether a current performance value is inherent for a workload, or is indicative of an issue which may be improved by controlling allocation of resources. According to the present techniques, performance monitoring circuitry is configured to determine an interference-specific performance value indicative of an impact on performance associated with a given memory system component handling at least one interfering memory transaction, where handling of an interfering memory transaction comprises use of memory system resource which would have otherwise been available for handling at least one affected memory transaction. The performance monitoring circuitry may for example determine that a given memory transaction is an interfering memory transaction when handling of the given memory transaction uses memory system resource in a way which decreases performance of handling at least one affected memory transaction. For example, the interference-specific performance value may indicate a performance impact on the handling of the affected memory transaction caused by the interfering memory transaction. Hence, the interference-specific performance value may provide a metric directly indicating a degree to which performance is affected by memory transactions interfering with each other at the given memory system component. The affected memory transaction may be a memory transaction yet to be issued (e.g., in the case of the interfering transaction causing a cache eviction) or may be a memory transaction which has already been issued, as will be discussed in examples below. The performance monitoring circuitry is configured to make the interference-specific performance value available to a resource control agent. The resource control agent may control usage of memory system resource in a memory system and is not particularly limited, for example being provided by hardware or software. The inventors have realised that when a resource control agent has access to information indicating a performance impact attributable to memory transaction interference, then the resource control agent may be placed in a better position to reallocate resources in a way which may lead to improved performance. For example, the resource control agent may, in response to determining that performance is being reduced as a result of memory transaction interference at a given memory system component, prevent one or more potentially interfering or affected memory transactions from being issued to that memory system component to reduce memory transaction interference. In comparison, if the resource control agent were unaware of the cause of a performance impact, it may be incorrectly assumed that the performance impact is inherent due to the workload and continue issuing interfering memory transactions to the memory system component. Likewise, without the present technique if the resource control agent incorrectly assumed that a performance drop was caused by interference between memory transactions and accordingly reallocated resources, then performance may have been harmed unnecessarily by the reallocation. For example, a utilisation value merely indicating that interconnect bandwidth is being heavily utilised and impacting performance does not allow a resource control agent to determine why performance is being reduced, as for example performance may be reduced merely due to high bandwidth usage of a particular memory transaction. In comparison, an interference-specific performance value indicating a performance impact associated with memory transaction interference at the interconnect allows the resource control agent to determine the cause of the reduction in performance, and if such a value indicates a high degree of memory transaction interference then this can allow the resource control agent to determine that performance may be improved by reducing memory transaction interference, for example by preventing further interfering memory transactions from being issued to that interconnect. The interference-specific performance value may also allow the effects on memory interference to be directly evaluated for a particular change of resource allocation. The location of the performance monitoring circuitry in a processing system is not particularly limited, and the performance monitoring circuitry may be provided in association with the given memory system component, may be provided in association with a memory system component other than the given memory system component, or may be provided elsewhere in a processing system. The performance monitoring circuitry may be configurable via a register, to enable software to have control over the behaviour of the performance monitoring circuitry, such as by controlling which interference-specific performance value may be determined and for which memory system components. The interference-specific performance value may be expressed in various ways. For example, calculation of an interference-specific performance value may take into account factors such as a baseline performance to enable the performance impact to be expressed as the portion of a performance impact directly attributable to memory transaction interference, although such factors may alternatively be taken into account by the resource control agent and hence not need to be determined by the performance monitoring circuitry. In some examples, the at least one memory system component may be configured to handle memory transactions specifying a task identifier identifying a software execution environment associated with said memory transaction. The software execution environment is not particularly limited and may comprise a particular thread, virtual machine, exception level, and so on. The software execution environment may in some cases represent a particular workload. The task identifier may for example identify which software execution environment caused the memory transaction to be issued. The task identifier may propagate through the memory system along with the corresponding memory transaction, and can allow control over allocation of memory system resource to a given software execution environment, for example based on at least one configurable resource control parameter configurable at runtime by the resource control agent. In such examples, the performance monitoring circuitry may be configured to determine at least one task-associated interference-specific performance value for which at least one of the affected memory transaction and the interfering memory transaction specifies a task identifier within a given set of one or more task identifiers. By providing a task-associated interference-specific performance value, the performance monitoring circuitry can provide more fine grained monitoring of memory transaction interference, for example identifying which tasks are more likely to issue memory transactions which impact or are affected by interference with other memory transactions. For example, this can allow a resource control agent to discern whether memory performance degradation, such as a reduction in the number of cache lines used by a particular task, is caused by a task’s natural memory utilisation behaviour or by memory interference. This can allow resources to be allocated in a more specific manner, for example to particular tasks, to improve performance. For example, a task-associated interference-specific performance value may represent a performance impact attributable to interference by interfering memory transactions issued by a particular group of one or more tasks. Such information may for example allow the resource control agent to allocate more memory system resource to memory transactions issued by the tasks having the largest performance impact, or may for example alter scheduling of the tasks having the largest impact on other memory transactions (e.g., by cancelling those tasks if it is determined that the usefulness of the task does not justify the high performance impact on memory transactions issued by other tasks). Similarly, a task-associated interference-specific performance value may represent a performance impact attributable to interference of affected memory transactions issued by a particular group of one or more tasks. Such information may for example allow the resource control agent to allocate more memory system resource to memory transactions issued by the tasks being most impacted by memory transaction interference. In some examples, the given set of one or more task identifiers may comprise a single task identifier. Hence, the task-associated interference-specific performance value may be indicative of a performance impact associated with memory transaction interference caused by interfering memory transactions issued by a particular task, or a performance impact associated with interference of affected memory transactions issued by a particular task. However, in other examples, the given set may comprise two or more task identifiers. Hence, the task-associated interference-specific performance value may provide information corresponding to groups of tasks. Providing task-associated interference-specific performance values corresponding to groups of tasks may provide a more efficient way to track performance for a particular set of tasks, as this can for example reduce the total number of task-association interference-specific performance values which may be tracked. In addition, it may often be the case that several tasks, such as several related threads, may have similar memory transaction properties which means useful information may be derived from tracking those tasks together. In some examples, the task-associated interference-specific performance value could be determined for the task identifier of the affected memory transaction, or the task identifier of the interfering memory transaction. However, in some examples, the performance monitoring circuitry may be configured to determine at least one task-associated interference-specific performance value for which the affected memory transaction is associated with a first task identifier within a first group of one or more task identifiers and for which the interfering memory transaction is associated with a second task identifier within a second group of one or more task identifiers. Hence, at least one task-associated interference-specific performance value may correspond to a particular combination of affected and interfering tasks, and may reflect the performance impact caused by memory transactions issued by a first group of one or more tasks (responsible for issuing memory transactions specifying task identifiers in the first group of task identifiers) interfering with memory transactions issued by a second group of one or more tasks. In some examples, a plurality of task-associated interference-specific performance values could be determined corresponding to a plurality of pairs of groups of task identifiers. Determining a performance value representing the performance impact caused by memory transaction interference between particular groups of tasks can allow performance to be improved further, because this can allow the resource control agent to determine which groups of tasks are most likely to issue memory transactions which interfere with each other. This can allow performance to be improved by, for example, issuing memory transactions issued by particularly interfering groups of tasks to different memory system components, or updating scheduling to reduce overlap of particularly interfering tasks. In some examples, the performance monitoring circuitry may be configured to determine at least one task-associated interference-specific performance value for which the first task identifier is the same as the second task identifier. Such a self-interference performance value may provide a resource control agent with useful information indicating the extent to which memory transactions issued by a task are interfering with other memory transactions issued by the same task, which can be useful for quantifying memory system performance and identifying how performance can be improved. It will be appreciated that self-interference performance values may be determined instead of or as well as cross-interference performance values for which the first and second task identifiers are different. The performance monitoring circuitry may in some examples be configured to calculate a self-interference or cross-interference interference-specific performance value quantifying a degree of self-interference and / or cross-interference generally. Such a value may be calculated by comparing task identifiers issued with memory transactions to determine whether a particular instance of interference is self-interference or cross-interference, but the reported interferencespecific performance value may not be specific to a particular group of task identifiers. In some examples, what data is accessed in response to a given memory transaction is independent of the task identifier. Hence, the task identifier may not provide any of the address bits of a memory transaction used to identify target data, but may instead be an independent identifier indicating the task associated with that memory transaction. Hence, the task identifier may indicate which software execution environment is associated with a memory transaction independently of which location is to be accessed. In some examples, the apparatus may comprise processing circuitry configured to process instructions in one of a plurality of operating states, and the processing circuitry may be configured to issue a memory transaction specifying a task identifier selected in dependence on a task identifier selection value stored in a selected task identifier register selected in dependence on a current operating state of the processing circuitry. For example, the task identifier selection value may be the task identifier itself, or may be a value used to determine the task identifier. The processing circuitry may for example comprise a set of task identifier registers corresponding to different operating states, and the current operating state may be used to select which task identifier to issue. This can be useful to avoid software needing to rewrite a task identifier register each time there is a supervisor call or exception taken to a more privileged operating state or an exception return back to a less privileged operating state. As described earlier, the memory system resource used by the interfering memory transaction is not particularly limited. In some examples, the memory system resource may comprise a plurality of cache entries, for example when the given memory system component comprises a cache. As will be discussed in greater detail below, performance may be reduced for example when an interfering memory transaction causes eviction from the cache of data to be accessed by a future affected memory transaction, thereby causing data to be returned later in response to the affected memory transaction. The interference-specific performance value may therefore be indicative of a degree to which interfering memory transactions reduce performance by reducing availability of cache entries for handling affected memory transactions. In some examples, the memory system resource may comprise memory system bandwidth, for example when the given memory system component comprises an interconnect. The bandwidth available for handling an affected memory transaction may for example be reduced due to handling of an interfering memory transaction. In some examples, the memory system resource may comprise a plurality of entries within a queue or buffer, and performance may be impacted due to a reduction in available entries for handling an affected memory transaction caused by an interfering memory transaction occupying those entries. The interference-specific performance value may therefore be indicative of a degree to which interfering memory transactions reduce performance by reducing availability of bandwidth or buffer / queue entries for handling affected memory transactions. As mentioned above, in some examples where the memory the memory system resource comprises a plurality ofcache entries for storing information associated with memory transactions. In such examples the performance monitoring circuitry may be configured to determine the least one interference-specific performance value based on a number of interfering cache evictions from the plurality of cache entries caused by handling of the at least one interfering memory transaction. For example, the performance monitoring circuitry may provide a value indicative of a number (e.g., a proportion) of cache evictions from a cache which were caused by an interfering memory transaction, where the evicted cache entries would have otherwise been available for handling an affected memory transaction, and hence performance of the affected memory transaction has been impacted. The interference-specific performance value may be associated with a particular evicting or affected task identifier (or group of task identifiers), a particular pair of evicting and affected task identifiers (or pairs of groups of task identifiers), or may not be associated with any task identifiers and may instead generally indicate a degree of memory transaction interference in a particular cache. In some examples, the performance monitoring circuitry may determine a cache eviction to be an interfering cache eviction when the evicted cache entry is predicted to be accessed by a future affected memory transaction. Eviction of the victim cache entry predicted to be used in future may cause a performance impact because the predicted future affected memory transaction will miss in the cache and the target data for that memory transaction may need to be accessed in a further level of cache or memory. For example, the affected memory transaction may be predicted to be issued in a particular upcoming window of time. It may be difficult to predict when an affected memory transaction will be issued in the future, and therefore it may be difficult to assess whether a particular eviction is an interfering eviction or not. However the inventors have recognised that cache entries are often accessed with temporal locality, meaning that a particular cache entry may be accessed several times close together in time. Therefore, if a particular cache entry has been accessed recently then it may be more likely that the cache entry will be accessed in the near future. Therefore, in some examples the performance monitoring circuitry may be configured to determine that a given cache eviction is an interfering cache eviction in response to determining that the victim cache entry of the given cache eviction has been accessed within a preceding time window. The size of the preceding time window may be implementation specific, but in general the determination of whether a particular cache eviction is an interfering cache eviction may be implemented efficiently using the temporal locality principle to predict whether the victim cache entry was likely to have been accessed in the near future based on how recently it was previously accessed. In some examples, cache entries may be configured to provide an interfering eviction field indicating whether eviction of said entry constitutes an interfering cache eviction. For example, the interfering eviction field may provide one or more bits indicating whether eviction of that entry is an interfering cache eviction or not. The interfering eviction field may for example be set to indicate that eviction of a particular cache entry would be an interfering eviction in response to determining that the cache entry has been accessed, and the field may thereafter be reset to indicate that eviction of that cache entry would no longer be an interfering eviction after a particular time has passed (corresponding to the preceding time window discussed above). However, other prediction mechanisms may also be used to control setting of the interfering eviction field, for example based on a different prediction of whether the cache entry is likely to be accessed in the near future. In some examples, cache entries may also indicate a task identifier associated with a memory transaction responsible for allocating that cache entry, which may be predicted to be the same task identifier as a task identifier specified by a future memory transaction to be handled using that cache entry, and can hence be used to determine a task-associated interferencespecific performance value. For example, the task identifier specified by an evicting memory transaction may be compared to the task identifier stored in the evicted cache entry to determine the task-associated interference-specific performance value for the particular combination of interfering (evicting) and affected (evicted) memory transactions. In some examples, the performance monitoring circuitry may be configured to provide the interference-specific performance value to performance value aggregation circuitry configured to aggregate performance values associated with a plurality of memory system components. The inventors have realised that aggregating memory performance values from a plurality of memory system components may provide insights into memory system resource utilisations which may not be possible when considering an interference-specific performance value for a single memory system component. In some examples, the apparatus may comprise processing circuitry configured to issue memory transactions for accessing data, and the processing circuitry may comprise the performance monitoring circuitry. Providing performance monitoring circuitry at the processing circuitry may enable performance values to be determined for a wider range of memory system components encountered by memory transactions, and may also enable simpler aggregation of performance values from a range of memory system components. In some examples, the performance monitoring circuitry may be configured to determine an interference-specific performance value indicative of a performance impact on the given memory system component attributable to memory transaction interference for which the affected memory transaction is a read transaction. This can support provision of performance monitoring circuitry in the processing circuitry. As read transactions lie on a task’s critical path, they are likely to be more significant in terms of quantifying a performance impact. In addition, certain properties of a read transaction such as completion time may be indicative of memory interference and can hence enable certain interference-specific performance values to be determined at the processing circuitry itself, whereas the timing of a write transaction may be less relevant to performance. In some examples, performance monitoring circuitry may report the interference-specific performance value, as well as other performance values, to the processing circuitry in a response to a given memory transaction. To allow the processing circuitry (which may report the value to a resource control agent) to determine which memory system component is associated with a particular performance value, the response may comprise a memory system component identifier identifying the memory system component responsible for handling the given memory transaction. In some examples, the resource control agent may comprise software executed by processing circuitry. In such examples the performance monitoring circuitry may be configured to make the interference-specific performance value available in a software-accessible location, such as an architectural register. Hence, the resource control agent itself may not be a hardware feature of the apparatus described above. The resource control software may not yet be installed on the apparatus at the time the apparatus is manufactured or sold, so the actions taken by the resource control software are not an essential part of the apparatus. Hence, from a platform point of view, in some examples the performance monitoring circuitry may have an interface (e.g. a set of memory-mapped registers accessible by software) to expose the interference-specific performance value to the resource control agent, but the apparatus itself may not actually comprise the resource control agent (the software may not yet be installed). In other examples, the resource control agent may comprise memory system resource control circuitry implemented in hardware. For example, a dedicated memory system resource management processor may be provided in hardware to process the interference-specific performance value and make control decisions for setting resource control parameters. Particular examples will now be discussed with reference to the Figures. Figure 1 illustrates an example of a processing system 2 comprising one or more memory transaction initiators. In this example the memory transaction initiators include one or more processing elements capable of instruction execution. The processing elements include one or more central processing units (CPUs) 4. The memory access initiators may include other types of initiator 5, such as other types of processing element such as a graphics processing unit (GPU) and / or other non-processing-element memory access initiators such as an input / output (I / O) device. While Figure 1 for sake of example shows one CPU 4, it will be appreciated that the system 2 could include different numbers of initiators of a given type, could include additional types of memory access initiators not shown in Figure 1 (e.g. a hardware accelerator) and may not necessarily include all of the types of memory access initiators. The memory transaction initiators communicate with each other and with memory storage 20 via a system interconnect 14. Some of the access initiators (e.g. CPUs 4) may have private caches 12 for caching data or instructions obtained from memory 20. The system 2 may also comprise a system cache 16 which is shared between multiple initiators. It will be appreciated that Figure 1 is merely a simplified representation of some components of a possible processing system, and the system could include other elements not illustrated for conciseness. The system 2 has at least one instance of performance monitoring circuitry 30 for generating one or more performance values indicating a level of current performance associated with use of a memory system component. For example, the memory system component may comprise a cache 12, 16, the interconnect 14 or a given sub-component of the interconnect 14 (such as a queue or buffer structure or a network-on-chip router), or other resources such as memory controller queues in a memory controller 15 for controlling a specific memory storage unit 20. While Figure 1 shows an example where instances of the resource monitoring circuitry 30 are provided at a CPU 4, the private cache 12, the memory system interconnect 14 (e.g. for monitoring usage of bandwidth on the interconnect 14 or an interconnect sub-component), the system cache 16, and the memory controller 15, it is not essential for each of these instances of resource monitoring circuitry 30 to be provided, and some examples could only provide a subset of these instances. Also, it would be possible to provide resource monitoring circuitry 30 associated with other processing system resources (e.g. associated with a memory management unit (MMU) responsible for address translation, for monitoring metrics relating to the current utilisation of the memory management unit). Each instance of performance monitoring circuitry 30 may monitor performance associated with a particular memory system component, such as the memory system component with which it is provided, or may monitor performance associated with one or more memory system components other than the memory system component with which it is provided. For example, the performance monitoring circuitry 30 in a cache 12 may also monitor performance of downstream components such as the interconnect 14 and memory controller 15. The system 2 may comprise (at least when in use at runtime), at least one resource control agent 34 which is responsible for managing utilisation of system resources. For example, the resource control agent 34 could control resource control settings that set limits on the amount of memory system resource (e.g. cache capacity, bus / interconnect bandwidth, MMU bandwidth, etc.) that can be consumed by requests from a given hardware requester or a given task executing on the CPU 4. The (system) resource control agent 34 could, in some examples, be implemented as software running on a CPU 4, in which case the hardware system 2 does not include a permanent component corresponding to the system resource control agent 34 (at the time the system 2 is manufactured or sold, the software of the system resource control agent 34 may not have been installed yet). Alternatively, the system resource control agent 34 could be implemented as a hardware processing unit which controls resource utilisation configuration settings (e.g. as a processing unit coupled to the memory system interconnect 14). Either way, the performance values captured by the performance monitors 30 may be exposed to the system resource control agent 34 (e.g. by writing the metrics to locations in memory 20 or to registers accessible to the system resource control agent 34), and the system resource control agent 34 may use those metrics to decide how to set the resource utilisation configuration settings for particular memory system resources such as a cache 12, 16, interconnect 14, request buffer, memory controller, etc., or control scheduling of tasks to enable more efficient utilisation of the memory system resources. The performance value could, for example, comprise a measure of request latency experienced by requests being processed by a given part of the memory system, a measure of request throughput (the number of requests processed in a given time), and / or a measure of the amount of spare resource capacity (e.g. cache capacity, bus / interconnect bandwidth or buffer / queue slots). It will be appreciated that there can be a wide variety of metrics that could be gathered by the performance monitoring 30. The performance monitoring circuitry 30 shown in Figure 1 is configured to determine at least one interference-specific performance value. Examples of an interference-specific performance value will be discussed below. In general, the interference-specific performance value is indicative of an impact on performance associated with a given memory system component handling at least one interfering memory transaction, where handling of an interfering memory transaction comprises use of memory system resource which would have otherwise been available for handling at least one affected memory transaction. Hence, the interference-specific performance value may provide a metric directly indicating a degree to which performance is affected by memory transactions interfering with each other at the given memory system component. If handling of a memory transaction comprises use of memory system resource, but use of that memory system resource does not affect handling of any other memory transaction (e.g., there is available remaining memory system resource for handling other transactions, or the memory system resource is not predicted to be used by any future memory transactions) then the interference-specific performance value may indicate no performance impact associated with handling such a memory transaction. This differs from a utilisation-based performance value, which may represent all changes in utilisation of memory system resource. The performance monitoring circuitry 30 is configured to make the interference-specific performance value available to a resource control agent 34, allowing the resource control agent 34 to control usage of memory system resource to improve performance taking into account a degree to which performance is being impacted by memory transactions interfering with each other at memory system components. The inventors have realised that providing interference-specific performance values can resolve issues with previous performance monitors which may have only monitored overall memory system resource utilisation. In particular, a greater insight into memory performance can be achieved when considering how interference between memory transactions affects performance. This can allow memory system resources to be allocated to improve performance, and can allow the effects on memory transaction interference of reallocating memory system resource to be evaluated. To provide one illustrative example, consider a periodic partition “P” alternating, in every period, between fetching a new working set from main memory and performing computations on it, with the working sets of previous periods to be never used again - a behaviour typical of sensor processing applications. In this setting and assuming the caches to be managed use an exclusive inclusion policy, should the partition’s code and working sets fit entirely in its core’s level one (L1) cache, the number of P’s cache lines (CLs) in the level two (L2) cache would belong to P’s past working sets which have been evicted from L1, and whose evictions from L2 to memory are associated with no performance impact as they will no longer be used. A performance monitor monitoring cache utilisation by the partition P might observe an increased utilisation of the L2 cache by P as lines are evicted from L1, and a subsequent decrease in utilisation of L2 as lines are evicted from L2. This utilisation value may provide misleading information, as whilst decreasing utilisation of L2 may suggest that the partition P is being affected by a lack of L2 memory system resource, under these conditions the changes in utilisation are in fact not associated with performance. In contrast to the utilisation value, an interference-specific performance value could indicate that, since eviction of the lines from L2 does not affect handling of any future memory transactions (as those lines will not be accessed in future), there is in fact no memory transaction interference associated with eviction of those lines from L2 and hence no performance impact attributable to memory system interference. Hence, whilst looking at the utilisation value alone may have suggested that additional L2 resource should be allocated to the partition P (to reduce the number of P’s L2 cache lines being evicted), the interference-specific performance value allows a resource control agent to determine that in fact no additional L2 cache resource needs to be allocated to partition P and hence allows the L2 cache resource to be utilised more effectively. Hence, providing a performance monitor configured to determine an interference-specific performance value can enable memory system resource to be allocated more effectively. Figures 1 and 2 illustrate an example for determining task-associated interference-specific performance values. Figure 2 illustrates an example of the processing unit (CPU) 4 in more detail. The processor includes a processing pipeline including a number of pipeline stages, including a fetch stage 40 for fetching instructions from the instruction cache 10, a decode stage 42 for decoding the fetched instructions, an issue stage 44 comprising an issue queue 46 for queueing instructions while waiting for their operands to become available and issuing the instructions for execution when the operands are available, an execute stage 48 comprising a number of execute units 50 for executing different classes of instructions to perform corresponding processing operations, and a write back stage 52 for writing results of the processing operations to data registers 54. Source operands for the data processing operations may be read from the registers 54 by the execution stage 48. In this example, the execute stage 48 includes an ALU (arithmetic / logic unit) for performing arithmetic or logical operations, a floating point (FP) unit for performing operations using floating-point values and a load / store unit for performing load operations to load data from the memory system into registers 54 or store operations to store data from registers 54 to the memory system. It will be appreciated that these are just some examples of possible execution units and other types could be provided. Similarly, other examples may have different configurations of pipeline stages. The processor has a memory management unit (MMU) 70 for controlling access to the memory system in response to memory transactions. For example, when encountering a load or store instruction, the load / store unit issues a corresponding memory transaction specifying a virtual address. The virtual address is provided to the memory management unit (MMU) 70 which translates the virtual address into a physical address using address mapping data stored in a translation lookaside buffer (TLB) 72. Each TLB entry may identify not only the mapping data identifying how to translate the address, but also associated access permission data which defines whether the processor is allowed to read or write to addresses in the corresponding page of the address space. In addition to the TLB 72, the MMU may also comprise other types of cache, such as a page walk cache 74 for caching data used for identifying mapping data to be loaded into the TLB during a page table walk. The processing element 4 supports an instruction set architecture which provides software with the ability to define, for a given software execution environment, one or more task identifiers (also referred to as partition identifiers and / or performance monitoring group identifiers) which distinguish one software task from another. Such identifiers can be specified in memory system requests sent out to the memory system, and propagate through the memory system along with those requests, so that memory system components can identify which software execution environment (task) a given request relates to. In particular, the registers 54 of the processing element 4 include a set of task identifier registers 68 used to set one or more task identifiers which are specified by a memory system request sent to the cache 12 of the processing element 4 or other parts of the memory system. The processing element 4 includes task identifier selection circuitry 70 (e.g. the load / store unit 50) which selects which items of task identifying information are specified by the memory system request, based on the information stored in the one or more task identifier registers 68. The task identifying information specified by the memory system request may include one or more identifiers which act as a label to distinguish memory system requests issued on behalf of different execution environments (e.g. different software execution environments executed by the processing element 4). The task identifying information does not influence which addresses in memory are allowed to be accessed by a particular execution environment, but is used for resource allocation control for regulating the level of performance seen for memory accesses issued by a particular execution environment and / or for control of performance monitoring so that separate performance metrics can be gathered for different tasks (execution environments). As shown in Figure 1, a given memory system component (e.g. system cache 16, interconnect 14, memory controller of memory 20, or private cache 12) could include a resource control agent 34 which uses the task identifying information for selecting resource allocation control settings, e.g. which limit the amount of memory system bandwidth which a particular execution environment is allowed to use, or limit a maximum fraction of cache capacity that a given execution environment is allowed to allocate for its own information. The memory system component resource control agent 34 may have access to a number of sets of resource allocation setting information (each set corresponding to a given value of the task identifying information), which specifies how to control resource allocations for requests specifying that value of the task identifying information. For example, the resource allocation settings may be defined in a memorybased table structure stored in memory 20 which is accessed at an address determined based on a base address defined in a register of the memory system component 16, 14, 20, 12. The base address register may be memory-mapped so that the base address can be set by software executing on a CPU 4 by executing a store instruction specifying an address mapped to the base address register. Alternatively, other configuration interface mechanisms may be provided as a configuration interface of the memory system component 16, 14, 20, 12 to allow software to control the resource allocation settings for handling requests with a given value of the task identifying information. In general, such resource allocation controls can be useful to prevent a “noisy” execution environment (which generates frequent cache requests) monopolizing a significant fraction of the available memory system resource (which may otherwise harm performance for other execution environments with less frequent requests which might not be able to gain sufficient usage of memory system resource if the amount of resource used by the “noisy” execution environment was not limited). Also, at least one instance of the performance monitoring circuitry 30 can have a similar configuration interface for configuring the performance monitoring circuitry 30 to maintain performance metrics for one or more distinct tasks as identified by respective values of the task identifying information. When a memory system request specifying a given value of the task identifying information is detected, the performance monitoring circuitry 30 may check whether that value of the task identifying information corresponds to a task for which a task-associated performance value is to be maintained, and if so updates a corresponding one of the task-associated performance values. The gathered performance values are made available for access by the system resource control agent 34 (e.g. software on a CPU or a hardware processor responsible for resource allocation control). Hence, the task identifier control registers 68 provide a mechanism by which software executing on the CPU 4 may control labelling of memory transactions (memory access requests, e.g., load requests and store requests) to assign task identifiers to memory access requests, which can be used to control resource allocation and / or gathering of task-associated performance values at a memory system component within the memory system. In some examples, the task identifying information assigned to a given memory access request may include more than one identifier, e.g.: - a partition identifier (PARTI D) which is used to control which set of resource allocation settings are applied by resource allocation control circuitry 34 of a memory system component 16, 14, 20, 12; and a performance monitoring group identifier (PMG) which is used by resource monitoring circuitry 30 to select which of several task-specific performance metrics are to be updated based on the memory access request. In some examples, the PARTID and PMG may be considered independent identifiers, with the resource control circuitry 34 selecting between resource allocation control settings based on the PARTID (independent of PMG) and the performance monitoring circuitry 30 selecting which performance metric to update based on the PMG (independent of PARTID). Alternatively, one of the PARTID and PMG may be regarded as a sub-identifier which distinguishes between different sub-classes of tasks corresponding to a given value of the other identifier. For example, while resource allocation control may be based on PARTID only (independent of PMG), performance monitor selection may be based on the combination of PARTID and PMG (so that tasks having the same PARTID but different PMGs might have different performance values maintained specific to each of those tasks even though the tasks share the same resource allocation settings controlled based on PARTID). The opposite approach is also possible, with performance monitor selection based on PMG only and resource allocation control based on the combination of PARTID and PMG. Regardless of the particular approach taken, providing multiple identifiers can give more flexibility in providing different granularity of control over resource allocation control compared to performance monitoring. However, it will be appreciated that providing multiple identifiers is not essential, and other approaches may provide a single identifier used to control both selection of resource allocation settings applied by resource allocation control circuitry 34 (e.g. caps on maximum cache allocation or maximum bandwidth consumption) and for selection of which performance monitoring value to update. Another item of task identifying information that can be specified for a given memory access request may be a task identifier space indicator. For example the processing element 4 may support execution of software in one of a number of security states, and software in different security states might specify the same value of PARTID or PMG in the task identifier registers 68, but to maintain secure isolation between the software operating in different security states, it may be desirable to prevent software in one security state influencing resource allocation control or performance monitoring associated with memory system requests issued from another security state. Therefore, as well as the PARTID / PMG identifiers, the task identifying information specified by a memory access request could also include a task identifier space indicator (e.g. an identifier of the security state from which the memory access request is issued), which distinguishes between multiple task identifier spaces which are assigned separate resource allocation settings or separate task-associated performance values even for requests specifying the same values of PARTID / PMG. The processing element 4 could select the task identifier space indicator for a given memory access request based on a current security state of the processing element 4 at the time the memory access request is issued. Hence, it will be appreciated that there can be a wide variety of ways in which task identifying information can be specified. In general, any information can be assigned to memory access requests which enables the software execution environment that caused that request to be issued to be distinguished by a memory system component or instance of performance monitoring circuitry 30. The architectural mechanism for setting the task identifiers used for particular software execution environments may be based on providing a set of one or more task identifier control registers 68. In some examples, there may be a single task identifier control register 68 to which one or more task identifiers (e.g. PARTID and / or PMG) can be written by software. In such implementations, memory access requests issued by the processing circuitry 4 specify the task identifier(s) currently specified in the register 68. When switching between different portions of software requiring their memory access requests to be distinguished from each other for performance resource control or resource utilisation monitoring purposes (e.g. on a context switch), software updates the task identifier control register 68 to specify the task identifier(s) for the new software to be executed after the switch, and then subsequent memory access requests will specify the new task identifier. Other examples could implement multiple task identifier control registers 68 specifying task identifiers associated with different operating states (e.g. privilege levels or exception levels associated with the processing element 4), and the current operating state of the processing element 4 at the time a memory system request is issued may be used to select which task identifier control register is selected, and hence which task identifier(s) is / are specified in the memory transaction. For example, this can be useful to avoid software needing to rewrite task identifier control registers 68 each time there is a supervisor call or exception taken to a more privileged operating state or an exception return back to a less privileged operating state, which may be relatively frequent events. Some implementations may provide an architectural mechanism for enabling different task identifiers to be specified for different classes of memory access request issued in the same software execution environment (with the same setting for the task identifier control registers 68). For example, there may be fields within the task identifier control registers 68 for specifying different task identifiers for data access requests issued in response to load / store instructions executed by the processing pipeline, instruction fetch requests issued to fetch instructions for processing by the pipeline, and / or page table walk requests issued by the processing circuitry 4 to request access to page table information used to translate addresses of memory access requests. Also, in some cases the task identifier(s) specified in the memory access request sent to the memory system may not be exactly the same as the task identifier value(s) stored in the relevant task identifier control register 28. Some implementations may support a task identifier virtualisation scheme where a virtual task identifier written by software to the task identifier control registers 28 is remapped to a physical task identifier appended to the memory system request, based on task identifier remapping information which can be defined by software. This can allow a number of different pieces of less privileged software (e.g. operating systems) to coexist on the same hardware platform while independently setting the task identifiers to be used for different software execution environments managed by the less privileged software, with more privileged software (e.g. a hypervisor) defining the task identifier remapping information so that conflicting task identifiers set by different operating systems can be mapped to different task identifiers as seen by the memory system. Hence, it will be appreciated that there are a wide variety of ways in which the task identifier of the memory access request could be determined, but in general the processing element 4 includes selection circuitry to select task identifying information to be specified for a given memory system request, with the particular value of the task identifying information being selected depending on at least one task identifier specified in a software-writable architectural register 28 of the processing element 4 (and also possibly depending on a current operating state, e.g. exception level and / or security state, of the processing element 4). While the interference-specific performance value described in this application is useful even in a system which does not support the tagging of memory transactions with task identifiers, the interference-specific performance value can be particularly effective in a system offering the ability for tracking of memory transactions, and allocation of memory system resource to memory transactions, at the granularity of a task. In particular, the performance monitoring circuitry 30 may determine one or more task-associated interference-specific performance values for which at least one of the affected memory transaction and the interfering memory transaction is associated with a task identifier within a given set of one or more task identifiers. Hence, a performance value may be determined by performance monitoring circuitry 30 indicating an extent to which a performance impact is caused by interference from memory transactions issued by a particular task (or group of tasks), an extent to which a performance impact caused by memory transaction interference affects memory transactions issued by a particular task (or group of tasks), and / or an extent to which a performance impact is attributable to interference between memory transactions issued by a particular combination of tasks (or groups of tasks). This can be useful information for improving performance, because it allows cases of memory interference associated with a particular task or combination of tasks to be identified. A resource control agent 34 may use this information to improve performance by, for example, increasing an allocation of resource to memory transactions from a task which is particularly affected by memory transaction interference or which is a particular source of memory transaction interference. The resource control agent 34 (particularly when provided at the CPU 4) may also use a task-associated interference-specific performance value to update a scheduling of tasks on the CPU 4. For example, if a particular task is found to be a particular source of memory transaction interference on the basis of the task-associated interference-specific performance value then the resource control agent 34 may choose to cancel or delay that task to reduce interference and improve overall system performance. Similarly, if a particular combination of tasks are found to issue memory transactions which interfere regularly, then such tasks may be rescheduled to be performed at different times. It will be appreciated that it is providing the task-associated interference-specific performance value which allows the aforementioned performance improvements to be achieved, and not the specific details of how the resource control agent decides to act on that information. Figure 3 provides an example of determining an interference-specific performance value in an example where the memory system component is a cache 12, 16 and the memory system resource comprises a plurality of cache entries 300. The interference-specific performance value in this example indicates a number of evictions for which eviction of a cache line to make a cache entry available for handling a first memory transaction interferes with handling of any further memory transactions. For example, the performance monitoring circuitry may determine whether the evicted cache line is predicted to be used to handle any upcoming memory transactions. If the evicted cache line is not predicted to be used in the near future, then eviction of that cache line does not affect handling of any memory transactions. Hence, the first memory transaction in this example would not be an interfering memory transaction because its handling does not require use of memory system resource (the cache entry storing the evicted cache line) which would otherwise have been used to handle an affected memory transaction. In comparison, if the evicted cache line is predicted to be used in the near future, then eviction of that cache line does affect handling of the future memory transaction which would have accessed the evicted cache line, because that affected memory transaction would need to obtain the evicted data from a slower to access further level of cache or memory. In this case, the first memory transaction would be an interfering memory transaction because its handling requires use of memory system resource (the cache entry storing the evicted cache line) which would otherwise have been used to handle an affected memory transaction. In particular, the performance monitoring circuitry for determining the interference-specific performance value for the example of Figure 3 may determine whether, for a particular eviction, the evicted cache line has an interfering eviction field 302 indicating that the eviction of that line is an interfering eviction or not. If the interfering eviction field 302 indicates that the eviction interferes with handling of an affected memory transaction (e.g., the field indicates “1”) then the performance monitoring circuitry 30 may update a value to indicate an increased performance impact associated with the particular cache due to memory transaction interference. If the interfering eviction field 302 indicates that the eviction does not interfere with handling of an affected memory transaction (e.g., the field indicates “0”) then the performance monitoring circuitry 30 may not update a value to indicate an decreased performance impact associated with the particular cache due to memory transaction interference, or may not update the interferencespecific performance value at all. The interfering eviction field 302 may be set for cache lines to indicate that eviction of that cache line is indicative of memory transaction interference when that cache line is predicted to be used in the near future. Due to temporal locality of cache accesses, a cache line being accessed recently may indicate that the cache line is more likely to be accessed in the near future and hence the interfering eviction field 302 may be set (e.g., to “1”) for a cache line 300 in response to that cache line being accessed. The interfering eviction field 302 may then be reset (e.g., to “0”) after some predetermined time has passed, after which it is determined that the likelihood of the cache line being accessed in the near future is low enough that eviction of that cache line is no longer likely to interfere with handling of a future memory transaction, and is hence not an interfering cache eviction. In some examples, the interference-specific performance value determined by a performance monitor 30 for the cache 12, 16 may also take into account the task identifier of either the evicting line, the evicted line, or both. For example, task-associated interference-specific performance values may be maintained for a group of tasks to monitor how much memory transactions issued by those tasks interfere with other memory transactions. Hence, in response to determining that a particular cache eviction is an interfering cache eviction (as described above) the task identifier (e.g., PARTID and / or PMG) associated with the memory transaction causing the eviction may be used to select which task-associated interference-specific performance value to update to record the instance of memory transaction interference. Similarly, task-associated interference-specific performance values may be maintained for a group of tasks to monitor how much memory transactions issued by those tasks are interfered with by other memory transactions. Hence, in response to determining that a particular cache eviction is an interfering cache eviction (as described above) the task identifier (e.g., PARTID and / or PMG) associated with the evicted cache line (indicative of the task identifier of the future affected memory transaction) may be used to select which task-associated interference-specific performance value to update to record the instance of memory transaction interference. In some examples, the task identifiers of both the evicted cache line and the evicting memory transaction may be used to select which task-associated interference-specific performance value should be updated to record whether handling the evicting memory transaction comprised memory transaction interference (as shown in Figures 5 and 6). For example, a set of task-associated interference-specific performance values may be maintained for various combinations of tasks associated with the interfering and affected memory transactions, which can enable tracking of interference between particular software execution environments. In some examples, the performance monitoring circuitry 30 may comprise a task identifier comparator 306. The comparator 306 may compare the task identifier of the evicted cache line with the task identifier associated with the memory transaction causing the cache line to be evicted. This can then be used to determine whether the evicting cache line and the evicted cache line are associated with the same or different task identifiers, to allow the performance monitoring circuitry 30 to determine a degree of self-interference for which the evicted and evicting cache lines are associated with the same task identifier and a degree of cross-interference for which the evicting cache line is associated with a different task identifier to the evicted cache line. For example, the performance monitoring circuitry 30 may maintain a self-interference interferencespecific performance value and a cross-interference interference-specific performance value, and the selection of which value to update to record whether the eviction was an interfering cache eviction may be made based on the outcome of the comparator 306. Hence, some examples may be configured to determine an interference-specific performance value on the basis of the task identifiers of the interfering and affected memory transactions, which may not be associated with a particular group of task identifiers but more generally quantifies a degree of self-interference or cross-interference experienced by the given memory system component (although it will be appreciated that cross- and self-interference may also be determined for specific task identifiers). Figure 4 provides a further example of determining an interference-specific performance value. In the example of Figure 4 the memory system component is an input queue, for example a queue in an interconnect 14 or within a memory controller 15, and the memory system resource comprises a plurality of entries 400 in the queue. The interference-specific performance value in this example indicates a degree to which occupation of entries in the queue by memory transactions interferes with handling of a memory transaction. For example, if the queue is full due to previously issued memory transactions, then this may interfere with handling of a new memory transaction and hence there may be memory transaction interference between the memory transactions in the queue and the memory transaction for which there is not space in the queue. The interference-specific performance value may not indicate any performance impact associated with the level of utilisation of the queue until the point that a given memory transaction is unable to be handled at a normal timing due to the reduction in availability of entries in the queue, at which point handling of the given memory transaction is determined to have been affected due to handling of other memory transactions. In the context of a queue, the memory system resource may not correspond to a particular queue entry (which may be assigned arbitrarily to memory transactions), but for example the overall capability of the queue to accommodate a new memory transaction. As with the example of Figure 3, a task-associated interference-specific performance value may be determined for the queue. The task-associated interference-specific performance value may for example indicate for a particular task whether memory transactions associated with that task have occupied space in a queue to the detriment of other memory transactions, or indicate for a particular task whether memory transactions associated with that task have been affected due to limited capacity in a queue, or for example may indicate for a particular combination of tasks whether memory transactions associated with those tasks have competed for space in a queue in a memory system component. Such task-associated performance values may for example be generated based on a comparison between task identifiers associated with entries in the queue and a task identifier associated with a memory transaction to be allocated to the queue, such a comparison for example being performed by comparator 402. It will be appreciated that whilst Figures 3 and 4 provide specific examples of memory transaction interference, the present techniques are not so limited. Interference-specific performance values may be determined for a range of memory system components having memory system resource, for example also including the interconnect for which memory transactions may compete for bandwidth. Figure 5 illustrates an example set of task-associated interference-specific performance values which may be determined by performance monitoring circuitry for monitoring memory transaction interference in a cache. In the example of Figure 5, a partition is a particular example of a software execution environment which may be associated with a particular task identifier. Hence, memory transactions issued by different partitions (Pt to PN) are propagated through the memory system in association with different task identifiers, allowing memory transactions to be associated with different partitions by the performance monitoring circuitry 30. Figure 5 illustrates an example in which the performance monitoring circuitry 30 determines a task-associated interference-specific performance value for each combination of evictor and victim partitions. For example, in response to determining that a particular eviction constitutes an interfering cache eviction the performance monitoring circuitry 30 may select which task-associated interference-specific performance value to update based on the combination of which task identifier was specified by the memory transaction causing the eviction (the evictor partition) and which task identifier was associated with the victim cache line (the victim partition). It will be appreciated that the precise values indicated in Figure 5 are merely an example, and may for example count a number of interfering cache evictions within a particular monitoring period for each combination of victim and evictor partitions. It will be noted that Figure 5 illustrates task-associated interference-specific performance values for combinations where the evictor partition is the same as the victim partition (i.e., self-interference) and for combinations where the evictor partition differs from the victim partition (i.e., cross-interference). Figure 6 illustrates an alternative example set of task-associated interference-specific performance values which may be determined by performance monitoring circuitry for monitoring memory transaction interference in a cache. Compared to Figure 5, the example of Figure 6 indicates performance values associated with groups of one or more partitions. Groups may comprise more than one partition, meaning that a smaller overall number of performance values may need to be maintained by the performance monitoring circuitry. For example, in response to determining that a particular eviction constitutes an interfering cache eviction the performance monitoring circuitry 30 may select which task-associated interference-specific performance value to update based on the combination of which group of task identifiers the task identifier specified by the memory transaction causing the eviction (the evictor partition) belonged to, and which group of task identifiers the task identifier associated with the victim cache line (the victim partition) belonged to. The present techniques may be used in a system comprising at least one performance monitor configured to measure an average memory performance (AMP) of tasks in the system as defined in a mathematical memory performance model. An AMP may be computed as a sum of a task’s average memory performance contributions (APCs) on each memory system component in the system which can act as a completer for a memory transaction (e.g., a cache or DRAM channel controller). An APC for a particular task on a particular completer memory system component (CMSC) may be computed as a product between the task’s average memory request performance (ARP) on the CMSC, and the fraction of task’s completed memory requests which were completed by the particular CMSC. The ARP for a task on a CMSC may be calculated in various ways, to generally represent the average performance of the memory requests for the task on that CMSC. For example, a throughput metric may be calculated based on a total number of memory requests completed by the CMSC divided by the total duration of its execution (in which case higher ARPs are better). Alternatively, a latency metric may be calculated to indicate a duration for completion of memory transactions from the task on the CMSC. The calculation of an average memory performance for a task can enable characterisation of the effect on performance of a task due to memory transaction interference, as a variation from the maximum interference-free memory performance baseline which may be achieved in an ideal interference-free counterpart of a real system. The interference-free memory performance baseline may for example be estimated via simulation to allow the interference variation to be determined from the measured AMP. The overall interference variation can be divided into a cross-interference (Cl) contribution representing memory transaction interference effects on a task caused by other tasks, and a self-interference (SI) contribution representing memory transaction interference effects on a task caused by its own memory transactions. In addition to its interference-free memory performance baseline, a task may also be characterized by its AMP when executed in isolation in a real system (dedicated execution), a quantity which, along with its associated minimum self-interference variation in such system, can be exploited as a baseline for evaluating the overall memory interference variation induced by other partitions and the current allocation of memory system resource. The defined memory performance model can be used to characterize a particular choice of memory system resource allocation settings in a system. A monitor for calculating an AMP value for a task may be provided in a CPU 4, or elsewhere in the system. The AMP monitor may for example calculate APCs based on read requests, as these lie on a task’s critical path and are therefore significant (e.g., more significant than write requests) for estimating an AMP, and because completion time and therefore latency of a read request may be determined in the CPU 4 as the contents is returned (while the time at which a write operation completes may not be propagated outside its CMSC), reducing the requirement for monitoring circuitry to be provided in CMSCs. In some examples, a determination of latency may be used when calculating a task’s ARP. Figure 7 provides an example of monitoring circuitry which may be used to calculate a latency for a memory transaction issued by a task. A set of reserving stations may be provided, each storing: • the physical memory address of a memory transaction, used as an identifier for identifying which reserving station to update when a response is received from a CMSC; • A temporal reference to the time at which a memory transaction associated with a particular reserving station was issued; and • A latency for the memory transaction, expressed in clock cycles, which may be calculated when the response is observed, by calculating an offset of a current time from the start time for the reserving station corresponding to the returned address. The latency may be reported for a particular memory transaction, and the AMP monitor may need to determine which CMSC completed the request to determine how to update the AMP for that task. In some examples, the completer memory system component may be assumed based on the latency. For example, if the latency is below a particular threshold then it may be assumed that the completer was a fast to access CMSC such as a L1 cache, if the latency is in a higher range then it may be assumed that the CMSC was an L2 cache, or if the latency was above a particular threshold it may be assumed that the CMSC was a memory 20 (for example). In other examples, however, CMSCs may tag memory transaction responses with a CMSC identifier to enable the CMSC for a particular memory transaction to be determined at the AMP monitor. In some examples, the interference-specific performance value determined by performance monitoring circuitry 30 may contribute to determination of an AMP, and hence may be returned to an AMP monitor (in some examples in association with an identifier of a particular CMSC). This can for example allow a determination to be made regarding how much of the AMP variation is due to self-interference or due to cross-interference. Figure 8 is a flow diagram illustrating a method of using performance monitoring circuitry to calculate an interference-specific performance value. At step 800, a memory transaction is 25 received to be handled by a given memory system component using at least one type of memory system resource. At step 802, it is determined whether handling the received memory transaction uses memory system resource for which it is predicted or known that the memory system resource would otherwise be used for handling another memory transaction. The performance monitoring circuitry hence determines whether handling of the received memory transaction uses memory system resource in a way which decreases performance of handling at least one affected memory transaction. It is also determined whether handling of the received memory transaction is affected due to a reduction in available memory system resource caused by handling of another memory transaction, and hence whether the received memory transaction is an affected memory transaction affected by memory transaction interference. Hence, the performance monitoring circuitry may determine whether or not the received memory transaction is an affected memory transaction, an interfering memory transaction, or both. At step 804, one or more interference-specific performance values are determined. The performance monitoring circuitry may for example maintain a set of interference-specific performance values and choose which performance value to update to record the occurrence / absence of memory transaction interference. For example, the selection may be made based on a task identifier associated with at least one of the affected or interfering memory transactions. At step 806, the interference-specific performance value calculated at step 804 is made available to a resource control agent. Concepts described herein may be embodied in computer-readable code for fabrication of an apparatus that embodies the described concepts. For example, the computer-readable code can be used at one or more stages of a semiconductor design and fabrication process, including an electronic design automation (EDA) stage, to fabricate an integrated circuit comprising the apparatus embodying the concepts. The above computer-readable code may additionally or alternatively enable the definition, modelling, simulation, verification and / or testing of an apparatus embodying the concepts described herein. For example, the computer-readable code for fabrication of an apparatus embodying the concepts described herein can be embodied in code defining a hardware description language (HDL) representation of the concepts. For example, the code may define a register-transfer-level (RTL) abstraction of one or more logic circuits for defining an apparatus embodying the concepts. The code may define a HDL representation of the one or more logic circuits embodying the apparatus in Verilog, System Verilog, Chisel, or VHDL (Very High-Speed Integrated Circuit Hardware Description Language) as well as intermediate representations such as FIRRTL. Computer-readable code may provide definitions embodying the concept using system-level modelling languages such as SystemC and SystemVerilog or other behavioural representations of the concepts that can be interpreted by a computer to enable simulation, functional and / or formal verification, and testing of the concepts. Additionally or alternatively, the computer-readable code may define a low-level description of integrated circuit components that embody concepts described herein, such as one or more netlists or integrated circuit layout definitions, including representations such as GDSII. The one or more netlists or other computer-readable representation of integrated circuit components may be generated by applying one or more logic synthesis processes to an RTL representation to generate definitions for use in fabrication of an apparatus embodying the invention. Alternatively or additionally, the one or more logic synthesis processes can generate from the computer-readable code a bitstream to be loaded into a field programmable gate array (FPGA) to configure the FPGA to embody the described concepts. The FPGA may be deployed for the purposes of verification and test of the concepts prior to fabrication in an integrated circuit or the FPGA may be deployed in a product directly. The computer-readable code may comprise a mix of code representations for fabrication of an apparatus, for example including a mix of one or more of an RTL representation, a netlist representation, or another computer-readable definition to be used in a semiconductor design and fabrication process to fabricate an apparatus embodying the invention. Alternatively or additionally, the concept may be defined in a combination of a computer-readable definition to be used in a semiconductor design and fabrication process to fabricate an apparatus and computer-readable code defining instructions which are to be executed by the defined apparatus once fabricated. Such computer-readable code can be disposed in any known transitory computer-readable medium (such as wired or wireless transmission of code over a network) or non-transitory computer-readable medium such as semiconductor, magnetic disk, or optical disc. An integrated circuit fabricated using the computer-readable code may comprise components such as one or more of a central processing unit, graphics processing unit, neural processing unit, digital signal processor or other components that individually or collectively embody the concept. Figure 9 illustrates a simulator implementation that may be used. Whilst the earlier described embodiments implement the present invention in terms of apparatus and methods for operating specific processing hardware supporting the techniques concerned, it is also possible to provide an instruction execution environment in accordance with the embodiments described herein which is implemented through the use of a computer program. Such computer programs are often referred to as simulators, insofar as they provide a software based implementation of a hardware architecture. Varieties of simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on a host processor 730, optionally running a host operating system 720, supporting the simulator program 710. In some arrangements, there may be multiple layers of simulation between the hardware and the provided instruction execution environment, and / or multiple distinct instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations which execute at a reasonable speed, but such an approach may be justified in certain circumstances, such as when there is a desire to run code native to another processor for compatibility or re-use reasons. For example, the simulator implementation may provide an instruction execution environment with additional functionality which is not supported by the host processor hardware, or provide an instruction execution environment typically associated with a different hardware architecture. An overview of simulation is given in “Some Efficient Architecture Simulation Techniques”, Robert Bedichek, Winter 1990 IISENIX Conference, Pages 53 - 63. To the extent that embodiments have previously been described with reference to particular hardware constructs or features, in a simulated embodiment, equivalent functionality may be provided by suitable software constructs or features. For example, particular circuitry may be implemented in a simulated embodiment as computer program logic. Similarly, memory hardware, such as a register or cache, may be implemented in a simulated embodiment as a software data structure. In arrangements where one or more of the hardware elements referenced in the previously described embodiments are present on the host hardware (for example, host processor 730), some simulated embodiments may make use of the host hardware, where suitable. The simulator program 710 may be stored on a computer-readable storage medium (which may be a non-transitory medium), and provides a program interface (instruction execution environment) to the target code 700 (which may include applications, operating systems and a hypervisor) which is the same as the interface of the hardware architecture being modelled by the simulator program 710. Thus, the program instructions of the target code 700 described above, may be executed from within the instruction execution environment using the simulator program 710, so that a host computer 730 which does not actually have the hardware features of the apparatus 2 discussed above can emulate these features. The simulator code 710 may for example comprise memory system component logic 712 to provide the functionality of at least one memory system component, and performance monitoring program logic 714 to provide the functionality of the performance monitoring circuitry 30. In the present application, the words “configured to...” are used to mean that an element of an apparatus has a configuration able to carry out the defined operation. In this context, a “configuration” means an arrangement or manner of interconnection of hardware or software. For example, the apparatus may have dedicated hardware which provides the defined operation, or a processor or other processing device may be programmed to perform the function. “Configured to” does not imply that the apparatus element needs to be changed in any way in order to provide the defined operation. In the present application, lists of features preceded with the phrase “at least one of mean that any one or more of those features can be provided either individually or in combination. For example, “at least one of: [A], [B] and [C]” encompasses any of the following options: A alone (without B or C), B alone (without A or C), C alone (without A or B), A and B in combination (without C), A and C in combination (without B), B and C in combination (without A), or A, B and C in combination. 5 Although illustrative embodiments of the invention have been described in detail herein with reference to the accompanying drawings, it is to be understood that the invention is not limited to those precise embodiments, and that various changes and modifications can be effected therein by one skilled in the art without departing from the scope of the invention as defined by the appended claims. 10

Claims

1. An apparatus, comprising:at least one memory system component configured to handle memory transactions for accessing data; andperformance monitoring circuitry configured to determine an interference-specific performance value indicative of a performance impact associated with a given memory system component handling at least one interfering memory transaction; wherein:handling of an interfering memory transaction comprises use of memory system resource which would have otherwise been available for handling at least one affected memory transaction, andthe performance monitoring circuitry is configured to make the interference-specific performance value available to a resource control agent.

2. The apparatus according to claim 1, wherein the at least one memory system component is configured to handle memory transactions specifying a task identifier identifying a software execution environment associated with said memory transaction; andthe performance monitoring circuitry is configured to determine at least one task-associated interference-specific performance value for which at least one of the affected memory transaction and the interfering memory transaction is associated with a task identifier within a given set of one or more task identifiers.

3. The apparatus according to claim 2, wherein the given set comprises two or more task identifiers.

4. The apparatus according to any of claims 2 and 3, wherein the performance monitoring circuitry is configured to determine at least one task-associated interference-specific performance value for which the affected memory transaction is associated with a first task identifier within a first group of one or more task identifiers and for which the interfering memory transaction is associated with a second task identifier within a second group of one or more task identifiers.

5. The apparatus according to claim 4, wherein the performance monitoring circuitry is configured to determine at least one task-associated interference-specific performance value for which the first task identifier is the same as the second task identifier.

6. The apparatus according to any of claims 2 to 5, wherein what data is accessed in response to a given memory transaction is independent of the task identifier.

7. The apparatus according to any of claims 2 to 6, comprising processing circuitry configured to process instructions in one of a plurality of operating states,wherein the processing circuitry is configured to issue a memory transaction specifying a task identifier selected in dependence on a task identifier selection value stored in a selected task identifier register selected in dependence on a current operating state of the processing circuitry.

8. The apparatus according to any preceding claim, wherein the memory system resource comprises at least one of:a plurality of cache entries, memory system bandwidth, and a plurality of entries within a memory system component queue.

9. The apparatus according to any preceding claim, wherein the memory system resource comprises a plurality of cache entries for storing information associated with memory transactions; andthe performance monitoring circuitry is configured to determine the least one interferencespecific performance value based on a number of interfering cache evictions from the plurality of cache entries caused by handling of the at least one interfering memory transaction.

10. The apparatus according to claim 9, wherein an interfering cache eviction is a cache eviction for which a victim cache entry is predicted to be accessed by an affected memory transaction.

11. The apparatus according to claim 10, wherein the performance monitoring circuitry is configured to determine that a given cache eviction is an interfering cache eviction in response to determining that the victim cache entry of the given cache eviction has been accessed within a preceding time window.

12. The apparatus according to any of claims 9 to 10, wherein cache entries in the plurality of cache entries are configured to provide an interfering eviction field indicating whether eviction of said entry constitutes an interfering cache eviction.

13. The apparatus according to any preceding claim, wherein the performance monitoring circuitry is configured to provide the interference-specific performance value to performance value aggregation circuitry configured to aggregate performance values associated with a plurality of memory system components.

14. The apparatus according to any preceding claim,comprising processing circuitry configured to issue the memory transactions for accessing data,wherein the processing circuitry comprises the performance monitoring circuitry.

15. The apparatus according to claim 14, wherein the performance monitoring circuitry is configured to determine an interference-specific performance value indicative of a performance impact on the given memory system component attributable to memory transaction interference for which the affected memory transaction is a read transaction.

16. The apparatus according to any of claims 14 and 15, wherein the given memory system component is configured to provide to the processing circuitry a response to a given memory transaction, wherein the response comprises a memory system component identifier identifying the memory system component responsible for handling the given memory transaction.

17. The apparatus according to any preceding claim, wherein the performance monitoring circuitry is configured to make the interference-specific performance value available in a software-accessible location.

18. Computer-readable code for fabrication of the apparatus of any preceding claim.

19. A method, comprising:handling memory transactions for accessing data;determining an interference-specific performance value indicative of a performance impact associated with a given memory system component handling at least one interfering memory transaction; andmaking the interference-specific performance value available to a resource control agent;wherein handling of an interfering memory transaction comprises use of memory system resource which would have otherwise been available for handling at least one affected memory transaction.

20. A computer program for controlling a host data processing apparatus to provide an instruction execution environment, the computer program comprising:memory system component program logic configured to handle memory transactions for accessing data; andperformance monitoring program logic configured to determine an interference-specific performance value indicative of a performance impact associated with given memory system component program logic handling at least one interfering memory transaction; wherein:handling of an interfering memory transaction comprises use of memory system resource which would have otherwise been available for handling at least one affected memory transaction, andthe performance monitoring program logic is configured to make the interference-specific 5 performance value available to a resource control agent.33

Citation Information

Patent Citations

  • Monitoring memory consumption

    US20100153475A1