Collecting diagnostic information
The data processing apparatus addresses the challenge of multiple monitoring contexts by using a base interval and context sampling intervals to efficiently profile software operations, ensuring accurate and efficient diagnostic information collection for multiple software agents.
Patent Information
- Application Number
- GB2023018062
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-06-11
AI Technical Summary
Existing data processing systems face challenges in efficiently profiling software operations due to the high cost and complexity of implementing multiple monitoring contexts, which can lead to conflicts and contamination of diagnostic information when different software agents require different sampling intervals.
A data processing apparatus is equipped with profiling circuitry that supports multiple monitoring contexts by using a base interval and context sampling intervals that are multiples of this base interval, allowing independent configuration of sampling rates while preventing conflicts, and filtering collected sample records based on these intervals.
This approach enables simultaneous profiling for multiple software agents with reduced circuitry overhead, ensuring accurate and efficient collection of diagnostic information without contamination, supporting various monitoring contexts and privilege levels.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND The present technique relates to the field of data processing. More particularly, the present technique relates to collecting diagnostic information. A data processing apparatus may have circuitry to collect diagnostic information about software executing on the apparatus. This diagnostic information can be used to analyse software performance in order to identify part of a software program that may be causing poor performance and possible reasons for any performance issues. The diagnostic information can be used by software engineers to optimise their software to reduce execution time and allow better utilisation of available resources in the data processing apparatus. SUMMARY In one example arrangement, there is provided circuitry for profiling operations within a data processing apparatus, the circuitry comprising: sampling circuitry to select a subset of operations within a data processing apparatus as sampled operations to be profiled; profiling circuitry to collect a sample record for each operation selected as a sampled operation by the sampling circuitry, the sample record comprising diagnostic information obtained from the data processing circuitry relating to the sampled operation; and base interval storage circuitry to store an indication of a base interval with which operations are to be sampled; wherein the profiling circuitry supports collection of sample records for each of a plurality of monitoring contexts, each monitoring context specifying a context sampling interval as a multiple of the base interval; wherein the profiling circuitry is configured to filter, for each monitoring context, the collected sample records based on the respective context sampling interval to produce a series of filtered sample records for the respective monitoring context. In another example arrangement, there is provided a method comprising: selecting a subset of operations within a data processing apparatus as sampled operations to be profiled; collecting a sample record for each operation selected as a sampled operation, the sample record comprising diagnostic information obtained from the data processing circuitry relating to the sampled operation; storing an indication of a base interval with which operations are to be sampled; wherein collecting the sample record for each operation selected as a sampled operations comprises, collecting sample records for each of a plurality of monitoring contexts, each monitoring context specifying a context sampling interval as a multiple of the base interval; and filtering, for each monitoring context, the collected sample records based on the respective context sampling interval to produce a series of filtered sample records for the respective monitoring context. In a yet further example arrangement, there is provided a computer program for controlling a host data processing apparatus to provide an instruction execution environment, the computer program comprising: sampling program logic to select a subset of operations carried out by data processing program logic as sampled operations to be profiled; profiling program logic to collect a sample record for each operation selected as a sampled operation by the sampling circuitry, the sample record comprising diagnostic information obtained from the data processing program logic relating to the sampled operation; and a base interval data structure to store an indication of a base interval with which operations are to be sampled; wherein the profiling program logic supports collection of sample records for each of a plurality of monitoring contexts, each monitoring context specifying a context sampling interval as a multiple of the base interval; wherein the profiling program logic is configured to filter, for each configured monitoring context, the collected sample records based on the respective context sampling interval to produce a series of filtered sample records for the respective monitoring context. BRIEF DESCRIPTION OF THE DRAWINGS Further aspects, features, and advantages of the present technique will be apparent from the following description of examples, which is to be read in conjunction with the accompanying drawings. Figure 1 schematically illustrates an example of a data processing apparatus having profiling circuitry; Figure 2 is a flowchart illustrating the operation of sampling circuitry and profiling circuitry according to an example; Figure 3 is a flowchart illustrating the operation of sampling circuitry and profiling circuitry according to another example; Figure 4 illustrates the use of multiple context sampling intervals according to an example; Figure 5 schematically illustrates another example of a data processing apparatus having profiling circuitry; Figure 6 schematically illustrates another example of a data processing apparatus having profiling circuitry; and Figure 7 schematically illustrates a simulation example. DESCRIPTION OF EXAMPLES One type of diagnostic information collection that may be performed is statistical profiling. With statistical profiling, a subset of instructions of micro-operations handled by the data processing apparatus are selected to be profiled, for example by selecting operations based on a regular interval. During the handling of the selected operation by the data processing apparatus, information about the behaviour of the selected operations is then collected to form a sample record which can be output for analysis. The information included in the sample record may directly indicate events that happened during the processing of the sample or a state of the data processing apparatus at the time. For example, the sample record may indicate whether events such as a cache miss in a given level of cache or a branch misprediction for a given branch occurred. The sample record may also or instead indicate cycle counts measuring the latency of certain events during the processing of the sampled operation or the latencies at each cache stage for memory access requests in a memory system. Where such memory access requests get handled by main memory, the diagnostic information may contain an indication of a dynamic random access memory (DRAM) bank that serviced the request. With statistical profiling, only a subset of all the operations handled by the data processing apparatus are selected for profiling and so it is possible to collect more detailed information about the handling of the sampled operation than would otherwise be possible with diagnostic techniques that aim to collect information relating all of the instructions handled by the data processing apparatus or count certain events in the data processing apparatus as they occur. To support statistical profiling within a data processing apparatus, the apparatus may be provided with profiling circuitry that is able to collect diagnostic information relating to a particular sampled operation from potentially numerous components in the data processing apparatus and aggregate this information into a sample record. To avoid the diagnostic information collected for one sampled operation being wrongly associated with another sampled operation, the profiling circuitry may support the collection of diagnostic information for only a single in-flight sampled operation at a time. The complexities involved in collecting this diagnostic information from around the data processing apparatus means that this profiling circuitry is relatively costly to implement; however, the insights provided into the execution of software on data processing apparatus may nonetheless be valuable enough to justify the inclusion of this circuitry. One approach to implementing statistical profiling would be to provide a single instance of profiling circuitry and corresponding sampling circuitry that selects operations to sample and marks them for sampling. Since the sampling circuitry / profiling circuitry are configured by software however, this approach may provide only a single monitoring context that can be configured to carry out this statistical profiling. That is to say that this approach may enable only one software agent to configure statistical profiling with one set of configuration information at a time. However, in some situations, it may be desirable to support multiple monitoring contexts to allow multiple software agents to make use of statistical profiling at the same time and / or for a particular software agent to make use of multiple monitoring contexts with different configurations. For example, where the data processing apparatus executes processes associated with a hypervisor that manages one or more guest operating systems, each of those guest operating systems executing applications, it may be desirable to support statistical profiling at the hypervisor level (to monitor the operation of the hypervisor) as well as at the operating system and / or application levels (to monitor the operating system / applications). If a single monitoring context is supported and it used by the hypervisor, the benefits of statistical profiling are not available to the operating system or applications. While the data processing apparatus may be provided with multiple instances of profiling circuitry to support such multiple monitoring contexts, with each instance able to collect diagnostic information for a different set of sample records, such an approach quickly becomes costly in terms of the circuitry required to implement this profiling and the area on-chip required to accommodate such circuitry. As described herein, techniques are provided which allow for the configuration of a plurality of monitoring contexts (and consequently the collection of diagnostic information for a plurality of different software agents) at the same time. The circuitry to implement this form of statistical profiling restricts the allowed sampling intervals to multiples of a defined base interval. That is, rather than allowing the monitoring contexts a free choice of sampling interval for their particular context, a base interval is defined and the monitoring contexts specify a context sampling interval as a multiple of the base interval. If instead the monitoring contexts were able to freely select the context sampling interval for the particular monitoring context, a conflict could arise between monitoring contexts where multiple monitoring contexts selected different sampled operations to be sampled that were in-flight at the same time. This could lead to contamination of the diagnostic information collected for the sampled operations where the collected information for one operation was wrongly associated with a different operation. Thus, by allowing the monitoring contexts to configure a sampling interval as a multiple of a base interval, each monitoring context is provided with the ability to independently specify the rate at which it decides to sample operations while mitigating against conflict between competing monitoring contexts making use of the same profiling circuitry. Once the sample records have been collected according to the base interval, the profiling circuitry can filter the collected sample records based on the context sampling interval for each monitoring context to produce a series of filtered sample records for each monitoring context. In accordance with the techniques described herein, there is therefore provided circuitry for profiling operations within a data processing apparatus. The circuitry comprises sampling circuitry to select a subset of operations within a data processing apparatus as sampled operations to be profiled and profiling circuitry to collect sample records for the sampled operations. The sampling circuitry may for example select the operations to be sampled based on expiration of a base interval. In this way, the sampling circuitry may be configured to select a next sampled operation in response to elapse of a sampling interval counted by an interval counter (which may count up or down as operations are observed). The profiling circuitry collects diagnostic information from various components of the data processing apparatus which may report to the profiling circuitry information relating to the sampled operations. Base interval storage circuitry (which may take the form of a register) is provided to store an indication of the base interval with which operations are to be sampled. In some examples, the base interval may be fixed; however, in other examples, the base interval may be modifiable by software to support different sampling intervals (those sampling intervals being multiples of the base interval). The profiling circuitry can then support a plurality of monitoring contexts, each monitoring context independently configurable with its own context sampling interval, where the context sampling interval is specified as a multiple of the base interval. The profiling circuitry then collects sample records for all of the configured monitoring contexts which can then be filtered based at least on the specified context sampling interval for each monitoring context to result in the sample records for the various monitoring contexts. The sample records may then be written to a buffer in memory and / or to ports to be read off-chip. In addition to counting the operations and marking operations to be sampled, the sampling circuitry in some examples, also marks the sampled operations to indicate which, if any, monitoring contexts are associated with the sampled operations. Thus, as the operations are handled (e.g., as instructions progress down a pipeline of the data processing circuitry), the operations may carry with them an indication of the monitoring context(s) for which the sampled operation may be included in the set of sample records. In such examples, the sampling circuitry may maintain separate counters for each of the monitoring contexts to identify when to tag the sampled operations as relating to the respective monitoring contexts. As an illustrative example, if the base interval were set to sample every 100 instructions, a first monitoring context samples with a multiple of 2 (i.e., every 200 instructions) and a second monitoring context samples with a multiple of 5 (i.e., every 500 instructions), every 200th instruction (subject to the application of perturbation as discussed below) will be tagged as relevant to the first monitoring context and every 500th instruction will be tagged as relevant to the second monitoring context. Consequently, every 1000th instruction will be tagged as relevant to both the first and second monitoring contexts. With the operations tagged in this way, the profiling circuitry can filter the collected sample records based on the tagging in order to determine which sample records should be reported for each monitoring context. While the sampling circuitry may indicate operations for sampling based on elapse of a sampling interval corresponding to the base interval, with this approach, for some configurations of the monitoring contexts, there will be operations tagged as sampled operations for which no monitoring context is actually monitoring. Continuing the illustrative example above, although the 100th, 300th, 700th and 900th instructions may be selected based on elapse of the base interval, neither of the monitoring contexts are configured to collect diagnostic information for such instructions. Accordingly, in some examples, the sampling circuitry is configured to suppress the selection of operations that are not associated with any monitoring context. That is, for operations that would otherwise be selected for sampling based on the base interval alone, but for which there is not a monitoring context to sample record would be reported, the sampling circuitry may suppress the selection as a sampled operation. This approach can prevent the profiling circuitry unnecessarily collecting diagnostic information relating to the operation. To support the tagging of operations as relevant to one or more monitoring contexts, the sampling circuitry may maintain a plurality of context counters, where each context counter is associated with a respective monitoring context. Then, when the context sampling interval for a given monitoring context elapses (as determined using the context counter), the sampling circuitry may mark a next sampled operation (as may be determined using a base interval counter) as being associated with the given monitoring context. While the sampling circuitry may tag the operations based on the monitoring context to which they relate as discussed above, in some examples, the sampling circuitry is configured to tag operations as sampled operations according to the base interval without indicating the monitoring context to which they relate. This may reduce the amount of information that needs to be passed along with the operation as it is handled by the data processing apparatus (e.g., as an instruction / micro-operation progresses along a pipeline). To filter the collected sample records, the profiling circuitry may then count the observed sample records and select only a subset according to the respective multiple. With this approach, context sampling counters may be implemented in the profiling circuitry rather than the sampling circuitry as described above. The context sampling intervals for each of the plurality of monitoring contexts may be independently programmable, for example, by allowing software to update registers storing the respective multiples by which the context sampling intervals are defined, thereby enabling different software processes to define sampling rates appropriate to the profiling that they are looking to carry out. To extract the collected diagnostic information, the circuitry may comprise a plurality of profiling buffers and the profiling circuitry may be arranged to write the sample records for each monitoring context to a respective profiling buffer. In some examples, the buffers may then be read from a trace output port which allows external hardware to read the collected diagnostic information. In some examples however, the collected diagnostic information for each monitoring context is written to a respective buffer in memory for access by the associated software process. The sampling circuitry may support random perturbation of the sampling interval so that if random perturbation of the sampling interval is enabled, the profiling circuitry may adjust the sampling interval by a random or pseudorandom value, to set the sampling interval counted by the interval counter in a given period. For example, the sampling interval may be configurable by a user (by setting control data in a register) to specify a nominal number of instructions or micro-operations to count between two successive sampled operations, but this nominal value may be adjusted by the random / pseudorandom value to vary the exact interval in one period of counting compared to the next. This reduces the risk that the sampling circuitry may repeatedly select the same operation within a program loop on multiple iterations through the loop, which could risk skewing profiling results. Since the monitoring contexts may be configured independently, in some cases, there may be no enabled / configured monitoring contexts at a given point. To prevent the collection of diagnostic information by the profiling circuitry in such cases, the sampling circuitry may be configured to suppress the collection of sample records entirely when all monitoring contexts are disabled. This may be done for example by stopping one or more counters at the sampling circuitry or otherwise by preventing the sampling circuitry from marking any operations for sampling. Statistical profiling may be implemented within a number of possible forms of data processing apparatus and / or to monitor a range of different operations and components within such data processing apparatuses. For example, the data processing apparatus may comprise processing circuitry such as the processing circuitry of a central processing unit (CPU) or graphics processing unit (GPU). In such instances, the sampled operations may be instructions executed by the processing circuitry and / or micro-operations decoded from such instructions. In some examples, the data processing apparatus includes a memory system and the circuitry may be arranged to monitor bus transactions (e.g., memory access requests) occurring within the memory system. The data processing apparatus could also comprise devices connected via a transport network such as an interconnect with the monitored operations being bus transactions. Where the data processing circuitry is able to operate at different privilege levels (which may also be referred to as exception levels), the circuitry may be provided with mechanisms to prevent processes operating at different privilege levels from accessing data associated with other privilege levels that they are not permitted to access. Traditionally, processes executing at a lower privilege level should be prevented from accessing data a higher privilege level while processes at the higher privilege level are allowed to access data associated with lower privilege levels. Accordingly, the profiling circuitry may be arranged to prevent, for a monitoring context associated with a given privilege level, the collection of sample records associated with processes executing at a higher privilege level. Therefore, if the monitoring context for the given privilege level is enabled and the data processing apparatus begins operating in a higher privilege level, the profiling circuitry may prevent the collection of sample records for the given monitoring context until execution returns to the given privilege level (or a lower privilege level). This may be done to maintain security by preventing processes at a particular privilege level obtaining information about processes at a higher privilege level. In some examples, however, the data processing apparatus may also restrict processes at a higher privilege level from accessing data associated with processes at a lower privilege level (for example a hypervisor may be prevented from accessing data associated with a guest operating system). As such, the profiling circuitry may be also operable to prevent higher privilege levels from collecting diagnostic information about processes operating at lower privilege levels. The circuitry may also support the use of selection criteria to further filter the sample records reported for a particular monitoring context. As part of the configuration of a monitoring context, one or more selection criteria may be specified. For example, the selection criteria may be used to cause the profiling circuitry to only report diagnostic information relating to certain types of operation (e.g., branch instructions or load / store operations). Thus, the amount of data collected can be reduced by filtering out information relating to unwanted operations. The use of selection criteria may be used in combination with the provision of multiple monitoring contexts (even within a single process) to enable the collection of diagnostic information for different types of operation with different frequencies. That is, by using different monitoring contexts defining different context sampling intervals and different selection criteria, a process may make use of several monitoring contexts to collect the desired amounts of information. In some examples, statistical profiling may be implemented in combination with one or more other diagnostic techniques. One such diagnostic technique is performance monitoring for which a plurality of event counters are provided, each to maintain a respective event count value based on monitoring of events during processing of software by a data processing apparatus. Each counter can be configured separately to count certain types of events such as cache misses, branch mispredictions, pipeline stalls or a number of clock cycles elapsed. In combination with the techniques described herein for statistical profiling, a data processing apparatus may be provided with performance monitoring circuitry and the performance monitoring circuitry may support the counting of a population event to count elapse of the base interval and / or the context sampling interval associated with any of the monitoring contexts. In this way, performance monitoring may be used in combination with statistical profiling with these population events used to associate the data collected according to each technique. Particular examples will now be described with reference to the figures. Figure 1 schematically illustrates an example of a data processing apparatus 2 comprising processing circuitry 4 for performing data processing. The processing circuitry 4 comprises a processing pipeline having a number of pipeline stages, including a fetch stage 6, a decode stage 8, an issue stage 10, an execute stage 12 and a writeback stage 14 in this example. It will be appreciated that this is just one example of possible pipelined configuration and other examples may have a different number of stages and may include additional types of pipeline stage. The fetch stage 6 fetches instructions from a level 1 instruction cache 16 for execution by the pipeline. A branch predictor 18 provides predictions of outcomes of branch instructions which can be used by the fetch stage 6 to decide which instructions to fetch beyond a branch. Instructions fetched by the fetch stage 6 are decoded by the decode stage 8 to generate control signals for controlling later pipeline stages to perform the operations represented by the instructions. The decode stage 8 may map the instructions to microoperations which represent operations to be performed at a granularity at which the execute stage 12 can execute them. In some pipeline implementations, there is always a one-to-one mapping between the instructions fetched from memory and micro-operations as seen at later stages so the micro-operations can simply be seen as equivalent to, or a representation of, the originally fetched instructions themselves (although it is still possible for the micro-operations to be represented in a different form the corresponding instructions - e.g. a micro-operation could be tagged with additional information, such as the “sampled operation” tag described below). Such a pipeline implementation may either be viewed as executing instructions directly (without decoding instructions to micro-operations, as there is no change of mapping) or could be viewed as executing micro-operations, which are decoded from instructions with a one-to-one mapping. Both views can be seen as equivalent descriptions of the same pipeline. In other examples, for at least some instructions the pipeline may support one-to-many or many-to-one mappings of instructions to micro-operations. Some instructions may still correspond to a single micro-operation. Other instructions may be split into multiple micro-operations by the decode stage 8. The decode stage 8 could also support fusion of two or more program instructions fetched from the cache 16 to form a single combined micro-operation supported by the execute stage 12. Hence, the micro-operations seen by the execute stage 12 may differ from the architectural definition of the instructions defined in an instruction set architecture supported by the data processing apparatus 2. In the description below, the term “micro-operation” refers to the form of the instruction as seen at the execute stage. This could simply be the original instructions themselves for some pipelines, or could be a modified form of the instruction or a microoperation obtained by splitting one instruction into multiple micro-operations or fusing multiple instructions into a combined micro-operation. The issue stage 10 queues micro-operations generated by the decoder 8 while awaiting for operands to become available and issues a micro-operation for execution when its operands are available (or when it is known that an operand will become available by the time the micro-operation reaches the relevant cycle of the execute stage at which the operand is needed). The execute stage 12 includes a number of execution units 22-28 for executing different types of micro-operation. Processing operations are performed by the execution units 22-28 based on operands read from registers 20. For example, the execution units of the execute stage 12 may include an arithmetic / logic unit (ALU) 22 for performing arithmetic or logical operations, a floating-point unit 24 for performing operations involving numbers represented as floating-point values, a branch unit 26 for determining whether branch instructions should be taken or not taken and for adjusting program flow to perform a non-sequential change of program flow when a branch is taken, and a load / store unit 28 for processing load operations to load data from a memory system to the registers 20 and store operations to store data from the registers 20 to the memory system. It will be appreciated that the particular set of execution units 22-28 shown in the execute stage 12 in the example of Figure 1 is just one possible arrangement, and other examples may have different types of execute unit or could have multiple execute units of the same type (e.g. multiple ALUs 22). The write back stage 14 writes results of executed micro-operations to the registers 20. The processing pipeline 4 could be an in-order processing pipeline, which is restricted to executing the micro-operations in an order corresponding to the program order in which the instructions are defined by the programmer or compiler and in which the instructions are fetched by the fetch stage 6. Alternatively, the processing pipeline 4 could support out-of-order processing, where the issue stage 10 is allowed to issue microoperations for execution in an order which may differ from the program order. For example, while one micro-operation is stalled awaiting its operands to become available, a later microoperation associated with a later instruction in program order may be executed ahead of the stalled instruction. If the pipeline supports out-of-order processing then additional pipeline stages could be provided, such as a rename stage to remap architectural register specifiers specified by decoded program instructions to physical register specifiers identifying particular hardware registers in the register bank 20. In this example, the memory system includes the level 1 instruction cache 16, a level 1 data cache 30, a shared level 2 cache 32 which can be used for both data and instructions, and main memory 34, which may include a number of memory storage units, peripheral devices and other devices accessible via load / store instructions executed by the pipeline 4. It will be appreciated that the particular cache hierarchy shown in Figure 1 is just one example and other examples may have a different number of cache levels or could provide a different arrangement of the instruction cache is relative to the data caches (e.g. the level 2 cache 32 could be split into separate level 2 instruction and data caches). A memory management unit 36 is provided for controlling address translation between virtual addresses (generated based on operands of instructions processed by the processing pipeline 4) to physical addresses identifying locations to be accessed in the memory system. The MMU 36 may include at least one translation lookaside buffer (TLB) 38 for caching information dependent on page table data obtained from the memory system 34 which defines address translation mappings and may also provide access permissions which control whether certain regions of an address space are accessible to particular processes executing on the pipeline 4. While for conciseness the MMU 36 is shown in Figure 1 as connected to the load / store unit 28 for use in data accesses, the MMU 36 may also be used for translating addresses and checking access permissions of instruction fetch addresses issued by the fetch unit 6 when fetching instructions for execution. Figure 1 is just one example of a possible processor architecture and other examples may have other components not explicitly indicated in Figure 1. To assist with software development, the processor 2 is provided with hardware resources which allow gathering of profiling information about the behaviour of instructions processed by the processing pipeline 4, which a software developer can use to perform code optimisation with the aim of modifying their code to run more efficiently. Sampling circuitry 50 is provided to select certain instructions or micro-operations as sampled operations to be profiled by profiling circuitry 52. The sampling circuitry 50 can select the sampled operations at different stages of the pipeline. For example, the sampling circuitry 50 may select certain fetched instructions as sampled operations and tag those fetched instructions at the fetch stage 6 to label those instructions to indicate that, as the instruction progresses down the pipeline 4, the profiling circuitry 52 should gather information on the behaviour of the sampled operation. Alternatively, the tagging of instructions of sampled operations by the sampling circuitry could take place at the decode stage 8 or at a later stage. Also it is possible that sampled operations are selected at the granularity of individual microoperations rather than at the granularity of the architectural program instructions fetched from memory. An advantage of selecting only a subset of operations as a sampled operation is that this greatly reduces the overhead in tracking information for profiling. For example, the sampling interval could be set to be long enough that in practice only a single operation in flight within the pipeline 4 is selected as a sampled operation at a time, so that the profiling circuitry 52 need only be provided with sufficient hardware resources to track behaviour of a single sampled operation at a time. This avoids the overhead of having to index storage structures which can store information for multiple operations, based on an operation identifier associated with a particular sampled operation, to select which entry of the storage structure to update based on information for the particular sampled operation. The profiling circuitry 52 comprises monitoring circuitry 70 for gathering information about the behaviour of sampled operations selected by the sampling circuitry 50. Although the monitoring circuitry 70 is shown as a single block within the profiling circuitry 52, in practice the monitoring circuitry may include a number of elements distributed about the processor to gather information from different components of the processor. For example the monitoring circuitry 70 may include event detection circuitry to detect occurrence of various types of events for a sampled operation. The types of events detected may depend on the type of sampled operation. For example, for a branch operation selected as a sampled operation, the events could track whether a branch misprediction occurred or whether the branch predictor 18 correctly predicted the branch. For load / store operations the events could include, for example, whether the load / store operation missed in a certain level of cache 30, 32, whether the address translation lookup for the load / store instruction missed in the TLB 38 or in a particular level of TLB, or whether an address fault occurred for the load / store instruction. Other types of events which may be monitored may be instruction fetches missing in the instruction cache 16, faults such as an undefined instruction exception, or whether certain instructions were delayed due to contention for resources. The monitoring circuitry 70 may also capture information about particular instructions, such as the instruction address of the sampled operation, a target address of a load / store operation or a branch target address of a branch operation, and items of architectural state from the registers 20 that are captured at the point when the sampled operation reaches a certain stage of processing (for example, context identifiers identifying the processing context in which the sampled operation was processed). The monitoring circuitry 70 may also have cycle counters which count the number of processing cycles taken for certain operations to complete, such as measuring the latency of an address translation or cache lookup, or the number of cycles for an operation to progress between a first point of processing and a second point of processing, for example. Hence, it will be appreciated that a variety of information can be gathered by the monitoring circuitry 70. The captured monitoring information can be recorded in a sample record stored in sample record storage circuitry 72 (e.g. registers or a buffer) of the profiling circuitry 52. Within a sample record captured for a given sampled operation, the record may specify the type of operation associated with the sampled operation (e.g. whether it is a branch, a load / store operation or an ALU operation, etc.) and also provides various information directly attributed to the sampled operation. The capture of the sample record in the sample record storage 72 is performed in hardware in the background of processing being performed on the pipeline 4, so does not require any specific software instructions to be executed to gather the information within the sample record. The profiling circuitry 52 and sampling circuitry 50 are arranged to support the collection of profiling data for multiple software agents concurrently. In the example shown in Figure 1, the apparatus supports two monitoring contexts, enabling the collection of profiling data for two software agents at the same time; however, it will be appreciated that more monitoring contexts may be supported by providing configuration registers 78, filtering circuitry 86, criteria application circuitry 74 (if used), and sample record writing circuitry 79 for each additional monitoring context. The present approach supports the use of a plurality of monitoring contexts without providing additional monitoring circuitry to collect diagnostic information for each monitoring context and without providing circuitry to distinguish between diagnostic information collected for multiple in-flight sampled operations at the same. To do this, base interval storage 80 is provided which stores a base interval, at integer multiples of which operations can be specified for sampling. The base interval itself may be fixed for the hardware implementation or may be user-modifiable by software. The profiling circuitry also provides configuration registers 78 for each of the monitoring contexts (of which two are shown in Figure 1) to allow software to specify certain configuration details for the monitoring context. The configuration registers 78 for each monitoring context include an enable register 84 to indicate whether the associated monitoring context is enabled or not. There is also provided a context multiple register 82 in which a context sampling interval for the monitoring context can be specified as a multiple of the base interval. For example, if the base interval (as indicated by the base interval storage circuitry 80) was 500, the context multiple register 82 for one of the monitoring contexts may specify a multiple of 4 to indicate that sampling is to be carried out based on every 2000th operation. It will be appreciated that the configuration registers 78 may store other items of configuration information such as configuration information to govern the application of filtering criteria carried out by the criteria application circuitry 74 as discussed below. The context sampling intervals may also be specified in other manners (e.g., by specifying the interval directly rather than as a multiple); however, in general the context sampling intervals are restricted to being integer multiples of the base interval. The sampling circuitry 50 has an interval counter 54 for counting instructions or micro-operations to determine when the next sampled operation should be selected. In some examples, the interval counter 54 triggers elapse of the counter based on the base interval as maintained in the base interval storage circuitry 80 and tags operations for sampling according to that frequency. The sampling circuitry 50 may also be provided with respective counters for the monitoring contexts such as the A counter 92 and B counter 93 shown in Figure 1 and which may be used to count when the context sampling intervals have elapsed. For example, the counters may increment / decrement their counts in response to elapse of the base interval as counted by the counter 54 and trigger elapse of their counts when the specified multiple for the respective monitoring context has been reached. Alternatively, the context counters 92, 93 may count each operation and trigger elapse of the counter when the context sampling interval (as defined based on the relevant multiple and the base interval) has been reached. The sampling circuitry 50 may make use of the counts for the different monitoring contexts to suppress the selection of sample records that are not being monitored by any of the monitoring contexts. Additionally or alternatively, the sampling circuitry may tag the sampled operations with an indication of the monitoring contexts to which they are relevant to allow the filtering circuitry 86 of the profiling circuitry 52 to more easily select the sample records relevant to each monitoring context. The sampling circuitry 50 may also support a random perturbation function where the sampling interval is perturbed by a random or pseudorandom value generated by a random number generator or pseudorandom number generator 56. While Figure 1 shows the pseudo random number generator 56 being within the sampling circuitry 50, in other examples the sampling circuitry could reuse (pseudo)random numbers generated by a (pseudo)random number generator provided elsewhere in the processing system which is also used for other purposes other than instruction or micro-operation sampling. For example the (pseudo)random number generator 56 could generate a random value in a certain range which is added to (or subtracted from) the base sampling interval specified, to generate the sampling interval to be counted by the interval counter 54 for the next period of operation. Enabling random perturbation can be useful, because it reduces the risk that when processing a loop of instructions, the same instructions or micro-operations keep being selected as the sampled operation on multiple iterations of the loop. In some implementations it may be optional whether random perturbation is enabled or disabled, with a user-configurable control parameter selecting whether random perturbation is enabled. Regardless of how the elapse of the sampled operations is determined by the sampling circuitry 50, the operations selected for sampling are tagged as sample operations. For example, a tag bit associated with an instruction or micro-operation may be set, and this tag bit may accompany the instruction or micro-operation selected as the sampled operation as it progresses down the pipeline 4. Similarly, a monitoring context bit for each supported monitoring context may be set to indicate whether the sampled operation is relevant to each of the monitoring contexts. Returning to the profiling circuitry 52, there is provided filtering circuitry 86 to filter the collected sample records to extract only the sample records that are relevant to the monitoring context. For example, where the multiple for a particular monitoring context is set to be 10, the filtering circuitry 86 may filter out all but the 10th sample records collected by the profiling circuitry 52. To support this filtering, the filtering circuitry 86 may be provided with a counter 87, 88 to count the collected sample records and to discard the sample records for which the counter has not elapsed. Where the sampling circuitry 50 tags the operations based on the monitoring context(s) to which they relate however, the filtering circuitry 86 need not be provided with such counters 87, 88 and the profiling circuitry 52 may instead use the tags to perform the filtering. Regardless of the how the filtering is done, the filtering circuitry 86 enables the use of different sampling intervals for different monitoring contexts by removing sample records that are not of interest to a particular monitoring context. Criteria application circuitry 74 is provided to allow further criteria to be applied by the profiling circuitry 52 when selecting whether a sample record captured for a particular sampled operation is made accessible for diagnostic analysis. In the embodiment shown in Figure 1, a sample record can be made accessible for diagnostic analysis by writing it to a respective profiling buffer structure 90, 91 stored in the memory system 34. The address range allocated for use as the profiling buffers 90, 91 may be determined based on buffer address identifying information stored in configuration registers 78 of the profiling circuitry 52, which can be set by a user under software control. The configuration registers 78 may also include configuration information which specifies what types of information should be included within the sample record for particular types of sampled operation, and for setting the criteria used by the criteria application circuitry 74. The writing of the sample record to the profiling buffers 90, 91 in memory may be performed without needing to interrupt the processing on the pipeline 4, so that no specific software instructions are needed to cause the sample record to be stored to the memory system. Hence, a certain number of sample records can be output to the profiling buffers 90, 91 without any interrupts occurring, until sufficient number of sample records have been generated and written out that the profiling buffer risks overflowing, and at that point a performance monitoring interrupt could be triggered to cause the processing to be interrupted and so that the exception handler could then take action to ensure that the sample records previously stored to the profiling buffers 90, 91 can remain accessible for diagnostic analysis (e.g. by updating the address parameters in the configuration registers 78 to allow subsequent sample records to be stored to a profiling buffer in a different region of the address space, or by reading out the sample records from the profiling buffer and storing them elsewhere or outputting them for external analysis). Although Figure 1 shows an example where sample records are made accessible for diagnostic analysis by writing them to the memory, an alternative would be to output them to trace buffers 94, 95 which is a dedicated hardware structure separate from the memory system 34 for storing diagnostic information on-chip, and / or to output the captured sample record over a trace output port 75 (either directly or via the trace buffers 94, 95), where the trace output port is a set of integrated circuit pins via which the sample record can be output to an external off-chip trace analyser or storage device. Some systems may support this trace output functionality instead of the ability to write sample records to memory, while other may support both approaches and may use configuration information in the registers 78 to select which approach to use. Figure 2 is a flowchart illustrating the operation of sampling circuitry and profiling circuitry according to an example. The sampling circuitry 50 monitors the operations handled by the data processing apparatus and in response to elapse of a counter corresponding to the base interval (subject to any additional perturbation) detected at step 202, the sampling circuitry marks the operation for sampling at step 204. This marking may for example be done by setting a tag bit within the operation to flag that the operation is to be sampled. With the operation marked in this way, the profiling circuitry 52 then collects diagnostic information about the handling of the operation from various components in the data processing apparatus at step 206 as the operation is handled. For each enabled monitoring context, the profiling circuitry 52 maintains a counter to allow the profiling circuitry 52 to determine whether the context sampling interval for that monitoring context has elapsed at steps 208, 214. Two such steps are shown in Figure 2 but it will be appreciated that this determination could be repeated for each monitoring context where more than two monitoring contexts are supported. If the context sampling interval for a particular monitoring context has not elapsed, the sampled record is not recorded as shown at step 220. Since, the sampled records may be collected at a frequency governed by the base interval which is more frequent than the context sampling interval, in general not all of the sample records collected by the profiling circuitry 52 will be applicable to each monitoring context. In some examples, an additional filtering step 210, 216 may be performed for each monitoring context to determine whether certain additional selection criteria are satisfied. These selection criteria may for example be used to restrict the sample records to only certain types of operation. If the criteria are not satisfied, the sample record is not written out for the monitoring context, as shown with step 220 whereas if any criteria imposed are satisfied, the sample record is written out at steps 212, 218. Figure 2 illustrates an example in which the sampling circuitry marks operations for sampling based on elapse of the base interval but does not track to which monitoring context(s) the operation will be relevant. Figure 3 is flowchart illustrating the operation of sampling circuitry and profiling circuitry according to another example in which the sampling circuitry 50 tracks the context sampling intervals and tags the operations based on the monitoring contexts for which they are relevant. In this example, the sampling circuitry 50 again tracks the base interval. However, in response to elapse of the base interval, as determined at step 302, the sampling circuitry 50 determines whether the context sampling interval for each monitoring context has also elapsed at step 304 and 306. If the context sampling interval for a particular monitoring context has not elapsed, no action needs to be taken and the sampling circuitry can continue its operation. However, if the context sampling interval for a given monitoring context has elapsed, the sampling circuitry 50 marks the operation as being for sampling by the relevant monitoring context at steps 308, 310. Thus, in this example, if the operation is not relevant to any of the monitoring contexts, the operation is not marked for sampling at all. The profiling circuitry collects diagnostic information about the marked operations to produce a sample record for the operation. In this example, the filtering performed by the profiling circuitry 52 can be done on the basis of the indications provided by the sampling circuitry 50. At steps 314, 320, it is determined whether the sample record relates to each monitoring context and the sample record is not written out for that monitoring context (as shown in step 326) if the sample record is not relevant. Similar monitoring and record writing steps as discussed with respect to Figure 2 can then be carried out for each monitoring context at steps 316, 318, 322, 324. Figure 4 illustrates conceptually the use of a plurality of context sampling intervals defined as a multiple of a base interval. Figure 4 shows a timeline in which each dash represents an operation handled by a data processing apparatus. Above some of the dashes is shown a letter to represent a monitoring context that will profile that operation. As shown in Figure 4, the base sampling interval, n, is 10 and the context sampling intervals for each of two monitoring contexts are expressed as a multiple of that base sampling interval. Specifically, the context sampling interval A will sample every 20 operations (expressed as 2n) and context sampling interval B will sample every 30 operations (expressed as 3n). Counting the initial operation as the 0th operation, the 0th operation will be profiled by both monitoring context A and monitoring context B. Then the 20th operation will be sampled by monitoring context A only, the 30th operation by monitoring context B only, the 40th operation by monitoring context A only and then the 60th operation by both monitoring context A and monitoring context B. In this way, the different monitoring contexts can be configured with different sampling intervals while allowing a plurality of monitoring contexts to make use of the same profiling circuitry. Figure 5 schematically illustrates another example of a data processing apparatus 200 having profiling circuitry 52 and sampling circuitry 50. Here, the data processing apparatus has two central processing units (CPUs) 104 and an input / output unit 106 for controlling input or output of data from / to a peripheral device. It will be appreciated that many other types of devices could also be provided, such as graphics processing units (GPUs), a display controller for controlling display of data on a monitor, direct memory access controllers for controlling access to memory, etc. At least some of the devices may have internal data or instruction caches 108 for caching instructions or data local to the device. Other devices such as the input / output interface 106 may be uncached devices. Coherency between data in the respective caches and accessed by the respective devices may be managed by a coherent interconnect 100 which tracks requests for accesses to data from a given address and controls snooping of data in other devices’ caches when required for maintaining coherency. It will be appreciated that in other embodiments such coherency operations could be managed in software, but a benefit of providing a hardware interconnect 100 for tracking such coherency is that the programmers of the software executed by the system do not need to consider coherency. As shown in Figure 5, some devices may include a memory management unit (MMU) 112 which may include at least one address translation cache for caching address translation data used for translating addresses specified by the software into physical addresses referring to specific locations in memory 114. It is also possible to provide a system memory management unit (SMMU) 116 which is not provided within a given device, but is provided as an additional component between a particular device 106 and the coherent interconnect 100, for allowing simpler devices which are not designed with a built-in MMU to use address translation functionality. In other examples the SMMU 116 could be considered part of the interconnect 100. The devices 104, 106 and memory 114 may communicate with each via the interconnect 100 by exchanging bus transactions. A user may wish to monitor these bus transactions to extract diagnostic information from a subset of the transactions. Within the data processing apparatus 200, there is provided profiling circuitry 52 and sampling circuitry 50 to implement the profiling techniques discussed above. That is to say that the sampling circuitry 50 may tag selected bus transactions occurring within the apparatus 200 for profiling by the profiling circuitry 52. As discussed above, the sampling circuitry 50 and profiling circuitry 52 may support a plurality of monitoring contexts to allow multiple software agents to configure profiling and / or to allow software agents to configure multiple instances of profiling. Figure 6 schematically illustrates another example of a data processing apparatus 2 having profiling circuitry 52. In this case, many of the components shown in Figure 6 are similar to those shown in Figure 1 and a full discussion of their features and functionality will not be repeated. In this example however, the profiling circuitry is arranged to monitor bus transactions between the processing circuitry 4 and the memory system 16, 30, 32, 34. That is, the sampling circuitry 50 is arranged to monitor the transactions (e.g., memory access requests) from the execute stage 12 of the pipeline 4 and / or transactions (such as instruction fetches) from the fetch stage 6 of the pipeline, and to tag sampled bus transactions for profiling by the profiling circuitry 52 in accordance with the techniques described herein. Figure 7 illustrates a simulator implementation that may be used. Whilst the earlier described embodiments implement the present invention in terms of apparatus and methods for operating specific processing hardware supporting the techniques concerned, it is also possible to provide an instruction execution environment in accordance with the embodiments described herein which is implemented through the use of a computer program. Such computer programs are often referred to as simulators, insofar as they provide a software based implementation of a hardware architecture. Varieties of simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on host hardware 730, optionally running a host operating system 720, supporting the simulator program implemented by simulator code 710. In some arrangements, there may be multiple layers of simulation between the hardware and the provided instruction execution environment, and / or multiple distinct instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations which execute at a reasonable speed, but such an approach may be justified in certain circumstances, such as when there is a desire to run code native to another processor for compatibility or re-use reasons. For example, the simulator implementation may provide an instruction execution environment with additional functionality which is not supported by the host processor hardware, or provide an instruction execution environment typically associated with a different hardware architecture. An overview of simulation is given in “Some Efficient Architecture Simulation Techniques”, Robert Bedichek, Winter 1990 USENIX Conference, Pages 53 - 63. To the extent that embodiments have previously been described with reference to particular hardware constructs or features, in a simulated embodiment, equivalent functionality may be provided by suitable software constructs or features. For example, particular circuitry may be implemented in a simulated embodiment as computer program logic. Similarly, memory hardware, such as a register or cache, may be implemented in a simulated embodiment as a software data structure. In arrangements where one or more of the hardware elements referenced in the previously described embodiments are present on the host hardware (for example, host processor 730), some simulated embodiments may make use of the host hardware, where suitable. The simulator program 710 may be stored on a computer-readable storage medium (which may be a non-transitory medium), and provides a program interface (instruction execution environment) to the target code 700 (which may include applications, operating systems and a hypervisor) which is the same as the interface of the hardware architecture being modelled by the simulator program 710. Thus, the program instructions of the target code 700 may be executed from within the instruction execution environment using the simulator program 710, so that host hardware 730 which does not actually have the hardware features of the circuitry discussed above can emulate these features. Thus, the features of the sampling circuitry 50 and the profiling circuitry 52 may be emulated by corresponding program logic 712, 714. The simulator code 710 may also have code to manage a base internal data structure 716 to emulate the base interval storage circuitry 80. Providing the simulation of the circuitry described herein can allow software for interacting 5 with the sampling circuitry and profiling circuitry to be developed before the hardware is actually available. In the present application, the words “configured to...” are used to mean that an element of an apparatus has a configuration able to carry out the defined operation. In this context, a “configuration” means an arrangement or manner of interconnection of hardware 10 or software. For example, the apparatus may have dedicated hardware which provides the defined operation, or a processor or other processing device may be programmed to perform the function. “Configured to” does not imply that the apparatus element needs to be changed in any way in order to provide the defined operation. Although illustrative embodiments of the invention have been described in detail 15 herein with reference to the accompanying drawings, it is to be understood that the invention is not limited to those precise embodiments, and that various changes and modifications can be effected therein by one skilled in the art without departing from the scope and spirit of the invention as defined by the appended claims.
Claims
1. Circuitry for profiling operations within a data processing apparatus, the circuitry comprising:sampling circuitry to select a subset of operations within a data processing apparatus as sampled operations to be profiled;profiling circuitry to collect a sample record for each operation selected as a sampled operation by the sampling circuitry, the sample record comprising diagnostic information obtained from the data processing circuitry relating to the sampled operation; andbase interval storage circuitry to store an indication of a base interval with which operations are to be sampled;wherein the profiling circuitry supports collection of sample records for each of a plurality of monitoring contexts, each monitoring context specifying a context sampling interval as a multiple of the base interval;wherein the profiling circuitry is configured to filter, for each monitoring context, the collected sample records based on the respective context sampling interval to produce a series of filtered sample records for the respective monitoring context.
2. The circuitry according to any preceding claim, wherein the sampling circuitry is configured to select a next sampled operation in response to elapse of a sampling interval counted by an interval counter.
3. The circuitry according to claim 1 or claim 2, wherein:the sampling circuitry is configured to mark the sampled operations to indicate any monitoring contexts with which the sampled operations are associated; andthe profiling circuitry is configured to filter the collected sample records for each configured monitoring context based on the marking by the sampling circuitry.
4. The circuitry according to claim 3, wherein the sampling circuitry is configured to suppress the selection of operations not associated with any monitoring context.
5. The circuitry according to claim 3 or claim 4, wherein:the sampling circuitry maintains a plurality of context counters, each context counter associated with a respective monitoring context; andthe sampling circuitry is responsive to elapse of a context sampling interval for a given monitoring context counted by an associated context counter to mark a next sampled operation as being associated with the given monitoring context.
6. The circuitry according to claim 2, wherein:the sampling interval corresponds to the base interval; andthe profiling circuitry is configured to filter, for each configured monitoring context, the collected sample records based on the specified multiple for the monitoring context.
7. The circuitry according to any preceding claim, wherein the context sampling intervals for the plurality of monitoring contexts are independently programmable.
8. The circuitry to any preceding claim, the circuitry comprising a plurality of profiling buffers, wherein the profiling circuitry is configured to write the sample records for each configured monitoring context to a respective profiling buffer.
9. The circuitry according to any of claims 2-8, wherein when random perturbation of the sampling interval is enabled, the sampling circuitry is configured to apply a random or pseudorandom perturbation to the sampling interval counted by the interval counter.
10. The circuitry according to any preceding claim, wherein the sampling circuitry is configured to suppress the collection of sample records when all monitoring contexts are disabled.
11. The circuitry according to any preceding claim, wherein the data processing apparatus comprises processing circuitry to execute instructions and the operations comprise instructions and / or micro-operations decoded from the instructions.
12. The circuitry according to any of claims 1-10, wherein the data processing apparatus comprises a memory system and the operations comprise bus transactions.
13. The circuitry according to any of claims 1-10, wherein the data processing apparatus comprises two or more devices connected via an interconnect and the operations comprise bus transactions.
14. The circuitry according to claim 11, wherein:the data processing apparatus is operable in a plurality of privilege levels, whereby by default processes executing at a given privilege level have access to data associated with processes executing a lower privilege level and are prevented from accessing data at a higher privilege level; andfor a monitoring context associated with a particular privilege level, the profiling circuitry is configured to prevent the collection of sample records associated with processes executing at a privilege level higher than the particular privilege level.
15. The circuitry according to claim 14, wherein:the data processing apparatus is operable to restrict data associated with processes at a given privilege level from being accessed by processes at a higher privilege level; andthe profiling circuitry is responsive to one or more higher privilege levels being prevented from accessing processes at the given privilege level to prevent, for monitoring contexts at the one or more higher privilege levels, the collection of sample records associated with processes at the given privilege level.
16. The circuitry according to any preceding claim, wherein the profiling circuitry is responsive to a configured monitoring context specifying one or more selection criteria to suppress the inclusion of collected sample records for the monitoring context that do not meet the one or more selection criteria.
17. The circuitry according to any preceding claim, further comprising:performance monitoring circuitry comprising a plurality of event counters, each to maintain a respective event count value based on monitoring of specified events during the processing of operations by the data processing apparatus;wherein the performance monitoring circuitry is responsive to performance monitoring configuration information for a given event counter specifying a population event for a particular monitoring context, to update the given event counter in response to elapse of the context sampling interval for the particular monitoring context.
18. An apparatus comprising the circuitry of any preceding claim and the data processing apparatus.
19. A method comprising:selecting a subset of operations within a data processing apparatus as sampled operations to be profiled;collecting a sample record for each operation selected as a sampled operation, the sample record comprising diagnostic information obtained from the data processing circuitry relating to the sampled operation;storing an indication of a base interval with which operations are to be sampled;wherein collecting the sample record for each operation selected as a sampled operations comprises, collecting sample records for each of a plurality of monitoring contexts, each monitoring context specifying a context sampling interval as a multiple of the base interval; andfiltering, for each monitoring context, the collected sample records based on the respective context sampling interval to produce a series of filtered sample records for the respective monitoring context.
20. A computer program for controlling a host data processing apparatus to provide an instruction execution environment, the computer program comprising:sampling program logic to select a subset of operations carried out by data processing program logic as sampled operations to be profiled;profiling program logic to collect a sample record for each operation selected as a sampled operation by the sampling circuitry, the sample record comprising diagnostic information obtained from the data processing program logic relating to the sampled operation; anda base interval data structure to store an indication of a base interval with which operations are to be sampled;wherein the profiling program logic supports collection of sample records for each of a plurality of monitoring contexts, each monitoring context specifying a context sampling interval as a multiple of the base interval;wherein the profiling program logic is configured to filter, for each configured monitoring context, the collected sample records based on the respective context sampling interval to produce a series of filtered sample records for the respective monitoring context.27
Citation Information
Patent Citations
Diagnostic data capture
US20190340097A1