Apparatus, method and computer program for collecting diagnostic information

The diagnostic information collection circuit filters, collects or suppresses diagnostic information based on SIMD processing configuration information, thereby solving the information confusion problem caused by array size in the prior art and achieving accurate analysis and optimization of processing circuit performance.

CN120787340APending Publication Date: 2025-10-14ARM LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480012963.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-24
Filing Date
2024-01-22
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

When collecting diagnostic information, existing technologies cannot effectively distinguish SIMD processing of different array sizes, resulting in information confusion and difficulty in gaining an in-depth understanding of the performance and operating status of the processing circuit.

Method used

Diagnostic information is collected or suppressed through diagnostic information collection circuitry based on SIMD processing configuration information filtering, including mode filtering configuration state and array size filtering configuration state, to ensure that information collection matches array size and operation mode.

Benefits of technology

It enables accurate collection of diagnostic information of the processing circuit under different operating modes and array sizes, providing deeper performance analysis and optimization opportunities and avoiding information confusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120787340A_ABST
    Figure CN120787340A_ABST
Patent Text Reader

Abstract

An apparatus has processing circuitry to perform single instruction multiple data (SIMD) processing on an array having a plurality of data items, and the processing circuitry supports the SIMD processing for multiple array sizes. The processing circuitry is capable of selecting an array size for performing the SIMD processing based on SIMD processing configuration information. Diagnostic information collection circuitry is provided to collect diagnostic information about software executing on the processing circuitry, and the diagnostic information collection circuitry screens the collection of the diagnostic information based on the SIMD processing configuration information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present technology relates to the field of data processing. More specifically, the present technology relates to collecting diagnostic information.

[0002] A data processing system can have diagnostic information collection circuitry for collecting diagnostic information about software executing on processing circuitry. The diagnostic information can be used, for example, to analyze software performance to help identify software program portions that can be causing poor performance and possible causes of any performance problems. Software engineers can utilize the diagnostic information to optimize the software to reduce its execution time and to more fully utilize available resources in the data processing system.

[0003] At least some examples provide an apparatus comprising: processing circuitry operable to perform single instruction multiple data (SIMD) processing on arrays comprising a plurality of data items, the processing circuitry supporting the SIMD processing for a plurality of array sizes; wherein the processing circuitry is configured to select an array size for performing the SIMD processing based at least in part on SIMD processing configuration information; diagnostic information collection circuitry for collecting diagnostic information about software executing on the processing circuitry; wherein the diagnostic information collection circuitry is configured to filter collection of the diagnostic information based on the SIMD processing configuration information.

[0004] At least some examples provide a computer readable medium for storing computer readable code for manufacturing the apparatus mentioned above.

[0005] At least some examples provide a method for collecting diagnostic information, the method comprising: collecting diagnostic information about software executing on processing circuitry; wherein the processing circuitry is operable to perform single instruction multiple data (SIMD) processing on arrays comprising a plurality of data items, the processing circuitry supporting the SIMD processing for a plurality of array sizes, wherein the processing circuitry is configured to select an array size for performing the SIMD processing based at least in part on SIMD processing configuration information; and wherein collecting the diagnostic information comprises filtering collection of the diagnostic information based on the SIMD processing configuration information.

[0006] At least some examples provide a computer program comprising instructions which, when executed by a host data processing apparatus, control the host data processing apparatus to provide an instruction execution environment for executing target program code, the computer program comprising: processing program logic operable to perform single instruction multiple data (SIMD) processing on an array comprising a plurality of data items, the processing program logic supporting the SIMD processing for a plurality of array sizes; wherein the processing program logic is configured to select an array size for performing the SIMD processing based at least in part on SIMD processing configuration information; diagnostic information collection program logic for collecting diagnostic information about software executing on the processing program logic; wherein the diagnostic information collection program logic is configured to filter the collection of the diagnostic information based on the SIMD processing configuration information.

[0007] At least some examples provide a computer readable medium storing the computer program mentioned above. The computer readable medium can be a non-transitory computer readable storage medium.

[0008] Further aspects, features, and advantages of the technology will be apparent from the example descriptions that follow, read in conjunction with the accompanying drawings, in which:

[0009] Figure 1 Examples of an apparatus having diagnostic information collection circuitry are illustrated;

[0010] Figure 2 Examples of a vector register and a matrix register are illustrated;

[0011] Figures 3A to 3B Examples of SIMD processing configuration information are illustrated;

[0012] Figure 4 is a flowchart illustrating a process of selecting an array size for performing SIMD processing;

[0013] Figure 5 is a flowchart illustrating a process of collecting or suppressing diagnostic information based on an operating mode of processing circuitry;

[0014] Figure 6 is a table illustrating a dependency of collecting or suppressing diagnostic information on an operating mode of processing circuitry;

[0015] Figure 7 is a flowchart illustrating a process of collecting or suppressing diagnostic information based on an array size specified by SIMD processing configuration information;

[0016] Figure 8 Examples of performance monitoring circuitry are illustrated;

[0017] Figure 9Examples of devices with profiling circuitry are illustrated;

[0018] Figure 10 Examples of multiple processors sharing matrix processing circuitry are illustrated; and

[0019] Figure 11 Analog examples are illustrated.

[0020] The following example description is provided before examples are discussed with reference to the accompanying drawings.

[0021] According to the techniques described herein, processing circuitry is provided that is operable to perform single instruction multiple data (SIMD) processing. With SIMD processing, a single instruction can cause operations to be applied to data in an array comprising multiple data items. The operations performed on different data items can be performed in parallel using multiple processing elements, although it will be appreciated that in some cases the SIMD processing can be implemented with repetition by a single processing element (or more generally using a number of processing elements that is smaller than the size of the array). In this sense, SIMD processing represents an architectural feature (i.e. an option that can be specified by program code) and microarchitectural implementations can or can not use multiple processing elements to implement the SIMD processing. Examples of SIMD processing include vector processing, in which operations are performed on vectors consisting of multiple elements logically arranged in one dimension, or matrix processing, in which operations are performed on matrices consisting of multiple elements logically arranged in two dimensions (or in some cases, a vector).

[0022] According to the techniques described herein, the processing circuitry is configured to select an array size for performing the SIMD processing based at least in part on SIMD processing configuration information. As the SIMD processing configuration information is stored in a memory-mapped register or a system register that can be modified using software, the SIMD processing configuration information is accessible. This allows software (and hence a user) to control the array size used for performing the SIMD processing.

[0023] However, if diagnostic information is collected for the processing circuit irrespective of the content of the SIMD processing configuration information, and hence irrespective of the array size, it can not be possible to gain an in-depth understanding of the performance of the processing circuit. For example, if the diagnostic information is used to count the number of cycles taken to perform a particular algorithm using SIMD processing operations, depending on whether a small array size is used (in which case a larger number of SIMD operations can be required) or a large array size is used (in which case a smaller number of SIMD operations can be required), the count can differ significantly. Thus, being able to filter the collection of diagnostic information based on the array size used can assist in gaining an accurate understanding of the health of the device, in particular the health of the SIMD processing performed by the processing circuit. Additionally, without relying on the array size used, it can be difficult or even impossible to categorise diagnostic information collected for SIMD processing performed using a range of different array sizes. Diagnostic information associated with one array size can become confused with one another due to the presence of diagnostic information collected during use of another array size. For example, even if the number of SIMD processing operations performed and the array size used during software execution are known, without knowing how many operations were performed in each mode, it can not be possible to determine the total number of data items on which operations were performed. This can hinder a software engineer intending to examine the execution of the software from gaining a meaningful insight into the performance of the software.

[0024] In some cases, a software engineer can only be interested in the execution of the software associated with a particular value of the SIMD processing configuration information, and thus by filtering the diagnostic information collected based on the SIMD processing configuration information, he / she can select only the information of interest.

[0025] Thus, in accordance with the techniques described herein, the device is provided with diagnostic information collection circuitry configured to filter the collection of diagnostic information based on the SIMD processing configuration information.

[0026] In some examples, filtering the collection of diagnostic information comprises collecting the diagnostic information depending on the SIMD processing configuration information, or inhibiting the collection of the diagnostic information. However, in some examples, filtering the collection of diagnostic information can take other forms. For example, basic diagnostic information can be collected irrespective of the SIMD processing configuration information, and the filtering can determine whether to collect an additional level of diagnostic information based on the SIMD processing configuration information.

[0027] The diagnostic information can take a number of possible forms, but in some examples the diagnostic information is indicative of one or more of:

[0028] • a cycle count when measuring certain event latencies during processing sample operations;

[0029] • elapsed clock cycles;

[0030] • execution of an instruction (any general instruction or a specific type of instruction);

[0031] • a memory access request (any general memory access, or a specific type of memory access such as a load or store) being issued;

[0032] • a cache access, cache linefill, or cache miss occurring (in some cases, this can be specific to a particular level or type of cache);

[0033] • a TLB access, TLB linefill, or TLB miss occurring (again, this can be a general event tracked for any TLB, or can be specific to a particular TLB instance (e.g., a data-side TLB or instruction-side TLB) or a particular TLB level (e.g., level 1 or level 2));

[0034] • a branch misprediction occurring;

[0035] • a queue or buffer becoming full (variants can be provided for particular buffers such as instruction issue queues, load buffers, store buffers, etc.); or

[0036] • a pipeline stall occurring due to a particular reason (e.g., a cache miss, TLB miss, or load or store buffer becoming full);

[0037] • a number of loops of a page table walk being performed to fill a TLB;

[0038] • a number of cycles taken to service a linefill request to bring data into a cache following a cache miss; and

[0039] • an indication of current occupancy of a particular queue or buffer.

[0040] It will be appreciated that this list is not exhaustive, and there are many more aspects of the operation of a processing circuit that can be monitored.

[0041] In some examples, the processing circuitry supports at least two operating modes associated with SIMD processing with different array sizes. More specifically, the processing circuitry is capable of operating in a first operating mode and a second operating mode, wherein the first operating mode performs SIMD processing with an array size (i.e., the number of elements) determined based on a first array size definition state, and the second operating mode performs SIMD processing with an array size based on a second array size definition state. In this case, the SIMD processing configuration information may indicate whether the SIMD processing circuitry is operating in the first operating mode or the second operating mode. Therefore, the diagnostic information collection circuitry may filter the collection of the diagnostic information (e.g., determine whether to collect the diagnostic information or suppress the collection of the diagnostic information) based on whether the processing circuitry is operating in the first operating mode or the second operating mode. Therefore, although the SIMD processing configuration information does not directly define the array size itself, the SIMD processing configuration information may be indirectly used to identify the array size to be used (e.g., by referencing the first array size definition state / the second array size definition state).

[0042] In addition to or in lieu of determining whether to collect the diagnostic information based on the operating mode of the processing circuitry, the diagnostic information collection circuitry may include in the collected diagnostic information an indication of the operating mode that the processing circuitry was operating in during the period associated with the diagnostic information. The indication of the operating mode included in the diagnostic information can be used to gain insight into how software operates in different operating modes or to implement calculations of values ​​that depend on array size for SIMD processing.

[0043] The diagnostic information collection circuit may also be configurable so that an operating mode in which diagnostic information is collected may be selected. To achieve this, the device is provided with a mode filtering configuration circuit to store a mode filtering configuration state. Based on the value of the mode filtering configuration state, the diagnostic information collection circuit may determine diagnostic information for each operating mode in the operating mode. For example, the diagnostic information collection circuit may collect the diagnostic information when the processing circuit is in the first operating mode in response to the mode filtering configuration state having a first value, and suppress the collection of the diagnostic information when the processing circuit is in the second operating mode. Therefore, the first value of the mode filtering configuration state can be used to cause the diagnostic information collection circuit to collect diagnostic information only when the processing circuit is operating in the first operating mode.

[0044] In some examples, the diagnostic information collection circuitry can collect the diagnostic information when the processing circuitry is in the second operating mode in response to the mode filtering configuration state having the second value, and inhibit collection of the diagnostic information when the processing circuitry is in the first operating mode. Thus, the second value of the mode filtering configuration state can be used to selectively collect diagnostic information only when operating in the second operating mode.

[0045] Thus, with these techniques, it can be selected which operating mode diagnostic information should be collected for and collected for only. Thus, it can be possible to gain insight into the behaviour of the software when executing in a particular operating mode. As these operating modes are associated with different array size definition states, this can be used to ensure that diagnostic information is not collected for SIMD processing performed with array sizes defined by different array size definition state entries, which might otherwise confuse the information collected.

[0046] In some cases, it can be necessary to collect diagnostic information regardless of the operating mode. To support this, the diagnostic information collection circuitry can collect the diagnostic information when the processing circuitry is in the first operating mode or the second operating mode in response to the mode filtering configuration state having a third value. That is, collection of the diagnostic information is thus independent of the operating mode.

[0047] It will be appreciated that the techniques are not limited to examples in which only two operating modes are used. Indeed, one or more additional operating modes can be provided in which the array size for SIMD processing is defined based on one or more corresponding additional entries of the array size definition state.

[0048] In some examples, instead of or in addition to specifying an operating mode in the SIMD processing configuration information, the SIMD processing configuration information specifies an array size to be used. This can take the form of an array size being included in the SIMD processing configuration information or can be encoded (e.g. by a value corresponding to a particular array size selected from a list of possible array sizes). The diagnostic information collection circuitry can then determine how to filter the diagnostic information based on the array size used to perform the SIMD processing, as can be established using the SIMD processing configuration information.

[0049] To support filtering of the diagnostic information based on a specified array size, the apparatus can comprise an array size filter configuration storage circuit to store an array size filter configuration state indicating one or more array sizes for which diagnostic information is to be collected. The diagnostic information collection circuit can then collect diagnostic information regarding execution of SIMD processing using an array size of the one or more array sizes specified by the array size filter configuration state. Conversely, for SIMD processing executed using an array size not specified by the array size filter configuration state, the diagnostic information collection circuit can be arranged to suppress collection of the diagnostic information.

[0050] One example of a diagnostic information collection circuit in which the present technology can be implemented is a performance monitoring circuit. A performance monitoring circuit is provided for monitoring performance of software executing on a processing circuit. The performance monitoring circuit comprises event counters each maintaining a respective event count value based on monitoring events during processing of the software by the processing circuit. A control circuit is also provided for configuring the event counters based on counter configuration information. The counter configuration information can comprise event type assignment information indicating which types of events are assigned to be monitored by the event counters. Such a performance monitoring circuit can be used to investigate possible causes of poor performance when software is executing on a processing circuit, as the event count values can expose information about internal events occurring within the processing circuit when the software is executing (such as cache misses, branch mispredictions, instruction stalls, buffer becoming full, etc.).

[0051] The configuration information can also comprise filter configuration information which can be used to select how the performance monitoring circuit should filter updates to the event counters in dependence on the SIMD processing configuration information. Thus, one or more of the event counters can be specific to events occurring, for example, in a particular operational mode or in relation to one or more particular array sizes. The performance monitoring circuit can update a given event counter in response to filter configuration information for the given event counter having a first value for events occurring when the processing circuit is in the first operational mode, and suppress updating the given event counter when the processing circuit is in the second operational mode, or otherwise filter the updates to the event counter based on the SIMD processing configuration information. Similarly, the performance monitoring circuit can update a given event counter in response to filter configuration information for the given event counter specifying one or more particular array sizes for events occurring when the processing circuit executes SIMD processing using one of those particular array sizes, and suppress updating the given event counter when using a different array size. Thus, event counts can be maintained in relation to periods of interest, and events collected outside of those periods ignored without counting.

[0052] In some examples, the performance monitoring circuitry can be operable to scale the updates to the event counter by a scaling factor before updating the event counter. Scaling can be used to reduce the amount by which the count in the event counter needs to be updated when the number of events expected to be counted is particularly large, thereby reducing the probability of the event counter overflowing (the count becoming too high to be recorded by the event counter) and reducing the demands placed on the circuitry provided to handle the event counter (e.g. by reducing the amount of circuitry required to add two counters together). Thus, the performance monitoring circuitry can be responsive to the counter configuration information, which specifies a scaling factor to be applied to a particular event counter, to scale the updates to the particular event counter by the scaling factor.

[0053] In some examples, the scaling factor can depend on the array size used by the processing circuitry to perform the SIMD processing. For example, the counter configuration information can configure a particular event counter to count the number of data items operated on. In the case of the processing circuitry performing SIMD processing, this number can become very large as a single instruction can result in operations on multiple data items. Thus, scaling can be used to scale the count maintained by the particular event counter by a scaling factor to a more manageable size. The scaling factor can be the array size used to perform the SIMD processing. With this scaling functionality, the event counter can be incremented by performing some SIMD operations. The scaling can also be undone later to restore the count of the number of elements operated on.

[0054] However, in the case where the processing circuitry is capable of operating with different array sizes, it can not be possible to properly restore the number of elements operated on as it is not known which array size was used when the event was recorded. Thus, by providing performance monitoring circuitry capable of selectively filtering events based on SIMD processing configuration information associated with array sizes, the present technology can allow event counters to be maintained which relate to operations performed only for a specified array size / operation mode. This can enable restoration of un-scaled counts (e.g. the number of data items operated on) from scaled counters.

[0055] Another case in which the present technology can be applied is for statistical profiling of software. To perform software profiling, the apparatus can be provided with sampling circuitry to select a subset of instructions or micro-operations as sample operations to be profiled. Profiling circuitry can also be provided to capture a sample record in response to processing of an instruction or micro-operation selected as a sample operation by the sampling circuitry, the sample record comprising diagnostic information about the behaviour of the sample operation, which can be directly attributable to the sample operation.

[0056] The information included in the sample record may directly indicate events that occurred during processing of the sampled operation, such as whether a cache miss occurred in a given level of cache or whether a branch misprediction occurred for a given branch, or may indicate a cycle count that measures the latency of certain events during processing of the sampled operation.

[0057] With this approach, because the sample record captures information directly attributable to the sampling operation, this avoids the "skid" problem that may arise with interrupt-based profiling mechanisms, which use counters to count specific events that may cause poor performance (such as cache misses or branch mispredictions) and, when a given number of such events have been detected, generate an interrupt to cause an exception handling routine to read architectural state or other diagnostic information from a register. This architectural state or other diagnostic information can then be used for diagnostic analysis. With such interrupt-based profiling mechanisms, there may be a "skid" delay between the time an interrupt signal is issued to cause the exception handling routine to read architectural state or other performance monitoring information from a register and the time the exception handler begins collecting the captured architectural state or diagnostic information.

[0058] Because the sampling circuitry selects only a subset of instructions or micro-operations as sampled operations, this significantly reduces the hardware and power overhead associated with collecting information about the sampled operations. The sample records captured for the sampled operations can provide a statistical view of overall program performance, rather than attempting to capture the behavior of every operation. Furthermore, providing sampling circuitry to select specific instructions or micro-operations as sampled operations to be profiled allows for tracking events occurring at different stages of the pipeline as the sampled operations pass through the pipeline, which may not be feasible in examples where profiling is based on capturing architectural state or counters indicating the occurrence of specific events, as this information may not be directly attributable to a specific operation but may instead be based on a variety of different operations.

[0059] The sampling circuitry and the profiling circuitry can be used to provide detailed information about the results and performance of specific operations processed by the processing circuitry, which can be useful for identifying possible causes of poor performance when executing a given program. However, another aspect of profiling can be to identify which parts of program code are executed more frequently than other parts. Software developers may have only a limited amount of time to devote to code optimization and may wish to focus their time on improving the performance of more frequently executed parts of the code rather than the performance of less frequently executed sections of the code.

[0060] To gain insight into the operation of the processing circuit (depending on the array size used for SIMD processing), the profiling circuitry may be configured such that the collection of the sample records depends on the SIMD processing configuration information. Thus, the profiling circuitry may be configured to selectively filter the capture of sample records based on the SIMD processing configuration information (e.g., whether the processing circuitry is in the first operating mode or the second operating mode, or the array size used).

[0061] The profiling circuitry may also include in the sample record an indication of the SIMD processing configuration information, such as the operating mode of the processing circuitry when processing the sample operation, or the array size used. The method may classify the collected information based on the operating mode / array size when analyzing the sample record.

[0062] When the SIMD processing configuration information indicates an operating mode to be used for SIMD processing, the first operating mode and the second operating mode may correspond to the processing circuit using different elements of the circuit to perform the SIMD processing. Thus, the processing circuit may be configured to: use a first SIMD processing circuit to perform the SIMD processing when in the first operating mode; and use a second SIMD processing circuit to perform the SIMD processing when in the second operating mode. This may correspond, for example, to using a vector processing circuit and a matrix processing circuit, both of which are capable of performing SIMD processing. For example, one of the first SIMD processing circuit and the second SIMD processing circuit may be a vector processing circuit, and the other of the first SIMD processing circuit and the second SIMD processing circuit may be a matrix processing circuit. In some examples, the matrix processing circuit may be operable to perform vector processing, such that for vector processing operations, the processing circuit can use either the vector processing circuit or the matrix processing circuit. However, the vector lengths (i.e., array sizes) used by the vector processing circuit and the matrix processing circuit when performing vector processing are based on different vector length definition states and may therefore differ. Therefore, by controlling the collection of diagnostic information depending on whether the vector processing circuit or the matrix processing circuit is used, relevant diagnostic information applicable to the SIMD processing circuit used can be obtained. However, in general, the first SIMD processing circuit and the second SIMD processing circuit do not need to correspond to the matrix processing circuit and the vector processing circuit. Both the first SIMD processing circuit and the second SIMD processing circuit may correspond to the vector processing circuit or the matrix processing circuit, or to another form of SIMD processing circuit.

[0063] In some microarchitectures, at least one of the first SIMD processing circuit and the second SIMD processing circuit may be shared among multiple processors (e.g., central processing units (CPUs), graphics processing units (GPUs), neural processing units (NPUs)). For example, a processor may be provided with vector processing dedicated to that processor and access to matrix processing circuitry shared with other processors.

[0064] The first array size definition state and the second array size definition state, as well as the dependency of the array size on the states, can take various forms. For example, the first array size definition state and the second array size definition state can be stored in corresponding registers, or in different areas of the same register.

[0065] In some examples, the first array size definition state and the second array size definition state directly specify the array size to be used, causing the processing circuit to perform the SIMD processing with the specified array size. However, in some examples, the array size definition state specifies a requested array size to be used. However, the actual array size used may additionally depend on the array sizes supported by the hardware and / or software configurable constraints that determine which array sizes can be selected.

[0066] Thus, for a particular operating mode, the processing circuitry may be arranged to determine the array size for SIMD processing based on the array size definition state for that mode by prioritizing the use of the requested array size. However, if the requested array size is unsuitable because it exceeds a software-configured or hardware-imposed maximum / minimum supported array size, the processing circuitry may use a different array size supported for that operating mode.

[0067] Thus, a device has been described that can support SIMD processing with a range of array sizes, wherein SIMD processing configuration information is used to control the array size to be used. Thus, the array sizes used to perform SIMD processing can vary. To avoid aliasing diagnostic information collected by diagnostic information collection circuitry due to the use of different array sizes, the collection of diagnostic information can be dependent on the array size used, or more generally, on the SIMD processing configuration that controls the array size used.

[0068] A device may include the processing circuit and the diagnostic information collecting circuit mentioned above, and any one or both of the first SIMD processing circuit and the second SIMD processing circuit.

[0069] A computer readable medium can store computer readable code for making the above-mentioned apparatus. As further described below, this can provide an electronic representation of a circuit design which can be propagated to another party to enable that party (or another party further downstream in the manufacturing chain) to make the apparatus.

[0070] The techniques discussed above can be implemented using hardware circuitry provided to implement the processing circuitry and diagnostic information collection circuitry discussed above.

[0071] However, the same techniques can also be implemented within a computer program which is executed on a host data processing apparatus to provide an instruction execution environment for executing target program code. Such a computer program can control the host data processing apparatus to emulate an architectural environment provided on a hardware apparatus which actually supports target code according to a certain instruction set architecture, even though the host data processing apparatus itself does not support that architecture. The computer program can have processing program logic which emulates the functionality of the processing circuitry mentioned above, and diagnostic information collection logic which emulates the functionality of the diagnostic information collection circuitry mentioned above. Such emulation can allow software development of target program code to begin before the hardware with the processing circuitry and / or diagnostic information collection circuitry is actually ready, for example, to debug software intended for use with an apparatus having the processing circuitry and diagnostic information collection circuitry discussed above. By executing target program code on an emulated execution environment, this can enable testing of target code in parallel with developing hardware devices which support the new features of the performance monitoring circuitry. The emulating program can be stored on a storage medium, which can be a non-transitory storage medium.

[0072] Particular examples will now be described with reference to the accompanying drawings.

[0073] Figure 1An example of an apparatus 2 having diagnostic information collection circuitry 60 is illustrated. The data processing apparatus has a processing pipeline 4 comprising several pipeline stages, each pipeline stage being implemented by corresponding circuitry. In this example, the pipeline stages include a fetch stage 6 for fetching instructions from an instruction cache 8. A branch predictor 7 is also provided to predict the outcome of a branch instruction, which can be used by the fetch stage 6 to determine which instructions to fetch following the branch. The pipeline stages also include a decode stage 10 for decoding the fetched program instructions to generate micro-operations (decoded instructions) to be processed by the remaining stages of the pipeline; an issue stage 12 for checking whether operands required for a micro-operation are available in a register file 14 and issuing the micro-operation once the required operands for a given micro-operation are available; and an execute stage 16 for executing the data processing operation corresponding to the micro-operation by processing the operands read from the register file 14 to generate a result value. This result value can then be written back to the register file 14. It should be understood that this is only one example of a possible pipeline arrangement, and other systems may have additional stages or different stage configurations. For example, in an out-of-order processor, a register renaming stage may be included that is used to map architectural registers specified by program instructions or micro-operations to physical register descriptors that identify physical registers in register file 14. In some examples, there may be a one-to-one relationship between program instructions decoded by decode stage 10 and corresponding micro-operations processed by execute stage 10. There may also be a one-to-many or many-to-one relationship between program instructions and micro-operations, such that, for example, a single program instruction may be split into two or more micro-operations, or two or more program instructions may be fused to be processed as a single micro-operation.

[0074] The execution stage 16 includes multiple processing units for performing different categories of processing operations. For example, the execution units can include scalar processing units 20 (e.g., including scalar arithmetic / logic units (ALUs) 20 for performing arithmetic or logical operations on scalar operands read from the registers 14), vector processing units 22 for performing vector operations on vectors including multiple data elements, matrix processing units 24 for performing matrix operations on vectors and matrices, and load / store units 28 for performing load / store operations to access data in the memory systems 8, 30, 32, 34. Here, the vector processing units 22 and the matrix processing units 24 each represent an example of SIMD processing circuitry. The matrix processing units 24 are capable of not only performing matrix processing on matrix inputs, but are also operable to perform vector processing. In some examples, the matrix processing units 24 provide all of the functionality of the vector processing units 22; however, the matrix processing units 24 can only implement a subset of the vector processing operations supported by the vector processing units 22. Other examples of processing units that can be provided at the execution stage include floating point units for performing operations involving values represented in floating point format, or branch units for handling branch instructions.

[0075] The apparatus 2 also includes a diagnostic information collection circuit 60 for collecting diagnostic information about the operation of the apparatus 2. This diagnostic information can for example include information about cache misses, branch prediction errors, instruction stalls, buffer becoming full, counts of instructions executed / operations performed, and so on.

[0076] The registers 14 include scalar registers 25 for storing scalar values, vector registers 26 for storing vector values, and matrix registers 27 for storing matrix values. The register file 14 also contains a SIMD processing configuration information register 66 for storing SIMD processing configuration information.

[0077] In some cases, the SIMD processing configuration information may indicate an operating mode such that when operating in a first operating mode, the device 2 is configured to perform vector processing 22 using the vector processing unit 22, and when operating in a second operating mode (as indicated by the SIMD processing configuration information 66), the device is configured to perform vector processing using the matrix processing unit 24. In such cases, the register file 14 also contains a first array size definition state register 62 containing a first array size definition state specifying the array size used by the vector processing unit 22. The register file 14 also contains a second array size definition state register 64 containing a second array size definition state specifying the array size used by the matrix processing unit 24.

[0078] In some cases, SIMD processing configuration information 66 may directly specify array sizes to be used when performing SIMD processing using vector processing unit 22 or matrix processing unit 24 .

[0079] like Figure 1 As shown, the memory system includes a level one data cache 30, a level one instruction cache 8, a shared level two cache 32, and a main system memory 34. It should be understood that this is only one example of a possible memory hierarchy, and other arrangements of caches may be provided. The specific type of processing units 20 through 28 shown in the execution level 16 is only one example, and other implementations may have different groups of processing units or may include multiple instances of the same type of processing unit so that multiple micro-operations of the same type can be processed in parallel. It should be understood Figure 1 This is just a simplified representation of some components of a possible processor pipeline arrangement, and the processor may include many other elements that are not illustrated for the sake of brevity.

[0080] Figure 2Illustrated are examples of vector registers 200 and matrix registers 210. Vector registers 200 include multiple elements 202-a, 202-b, ... 202-n (collectively referred to as element 202). Each element 202 can be used to store separate data items. Then, the vector and / or matrix processing unit can operate on vector 200 as a whole, and each element 202 of vector 200 is applied to the data item separately. Although in some cases all elements 202 of vector 200 can be applied, in some cases, predicates can also be adopted to allow the value of the predicate stored in the predicate register to selectively apply operations to some elements 202 in the vector 200. Matrix registers 210 also include multiple elements 204-a, 204-b, ... 204-mn (collectively referred to as element 204), wherein each element 204 can be used to store separate data items. Matrix processing unit 24 can operate on vector and / or matrix data, again performing operations on the data items in corresponding element 204.

[0081] Figure 3A and Figure 3B The following schematically illustrates possible forms of SIMD processing configuration information 66. Figure 3A As shown, SIMD processing configuration information 66 includes SIMD mode information indicating in which of the first and second operating modes the processing circuit 4 operates. When the processing circuit 4 operates in the first operating mode, the array size used by the processing circuit 4 is specified by the first array size definition state 62, and when the processing circuit 4 operates in the second operating mode, the array size used by the processing circuit 4 is specified by the second array size definition state 64.

[0082] exist Figure 3B , an example of SIMD processing configuration information 66 specifying an array size is shown. The array size may be stored in SIMD processing configuration information registers 66 or may be indicated by SIMD processing configuration information registers 66 in other ways.

[0083] Figure 4 3 is a flow chart illustrating a process of selecting an array size for performing SIMD processing. The array size may correspond to the length of a vector used in vector processing or the dimension of a matrix used in matrix processing. After starting at step 302, the processing circuit 4 determines whether the requested array size (e.g., by setting a value in a requested array size register or by indicating the requested size in a field of the instruction) is less than the minimum supported array size, which is requested by the software. The minimum supported array size may be a minimum size set by the software, or may be a minimum array size supported by the hardware performing the SIMD processing.

[0084] If the requested array size is smaller than (or equal to) the supported minimum array size, then at step 306 , the processing circuit 4 uses the supported minimum array size as the array size for the operation.

[0085] If the requested array size is greater than the minimum array size, the process proceeds to step 308, where it is determined whether the requested array size is greater than the maximum supported array size. If the requested array size is greater than the maximum supported array size, the process proceeds to step 310, where the maximum supported array size is used. Otherwise, the requested array size is used at step 312.

[0086] Figure 5 is a flow chart illustrating a process for collecting or suppressing diagnostic information based on the operating mode of the processing circuitry 4. After starting the process at step 402, the mode filter configuration state is checked at step 404. The mode filter configuration state allows selection of the operating mode in which diagnostic information should be collected.

[0087] If the mode filter configuration state has a first value (indicating that diagnostic information should be collected only for the first operating mode), the process proceeds to step 406, where it is determined whether processing circuit 4 is operating in the first operating mode. If processing circuit 4 is operating in the first operating mode, the process proceeds to step 410, where diagnostic information is collected. On the other hand, if processing circuit 406 is operating in an operating mode other than the first operating mode, the process proceeds to step 412, which corresponds to suppressing the collection of diagnostic information.

[0088] If the mode filter configuration state checked at step 404 instead has the second value, the flow proceeds to step 408, where it is determined whether the processing circuit 4 is operating in the second operating mode. If the processing circuit 4 is operating in the second operating mode, the flow proceeds to step 410, where diagnostic information is collected. On the other hand, if the processing circuit 406 is operating in an operating mode other than the second operating mode, the flow proceeds to step 412, which corresponds to suppressing the collection of diagnostic information.

[0089] When the mode filter configuration state has the third value (the third value is used to indicate that diagnostic information should be collected regardless of the operating mode), the process proceeds to step 410 where diagnostic information is collected without checking in which mode the processing circuit 4 operates.

[0090] Figure 6is a table illustrating the dependency of collecting or suppressing diagnostic information on the operating mode of processing circuit 4. As illustrated in the table, a mode filtering configuration state can be implemented using two bits, thereby providing four possible values ​​of the mode filtering configuration state. When the mode filtering configuration state has a value of 0b00, the diagnostic information collection circuit is arranged to collect the diagnostic information regardless of the mode in which processing circuit 4 operates. When the mode filtering configuration state has a value of 0b01, the diagnostic information collection circuit is configured to collect diagnostic information only when processing circuit 4 operates in the first mode, and therefore suppress the collection of diagnostic information when processing circuit 4 operates in the second mode. When the mode filtering configuration state has a value of 0b10, the diagnostic information collection circuit collects diagnostic information only when processing circuit 4 operates in the second mode (rather than when processing circuit 4 operates in the first mode). Therefore, the mode filtering configuration state can be used to control the dependency of diagnostic information collection on the operating mode of processing circuit 4. When the mode filtering configuration state has the final possible value of 0b11, in this case, the behavior of the diagnostic information collection circuit is undefined.

[0091] Figure 7 7 is a flow chart illustrating a process for collecting or suppressing diagnostic information based on array sizes specified by SIMD processing configuration information. In this example, after starting at step 702, at step 704, the array size used by processing circuitry 4 to perform SIMD processing is compared to one or more array sizes specified by the array size screening configuration state (which indicates which array sizes diagnostic information is to be collected).

[0092] The array size filtering configuration state can indicate in various ways which array sizes diagnostic information is to be collected. For example, the array size filtering configuration state can be a bitmap where each bit indicates whether the corresponding array size is enabled / disabled for diagnostic information capture, or can indicate a size threshold that specifies whether diagnostic capture is enabled / disabled based on a comparison of the current array size to the threshold.

[0093] If the processing circuitry 4 uses the array size specified by the array size filter configuration state (which may be established by comparing the SIMD processing configuration information with the array size filter configuration state), diagnostic information is collected for processing at step 708. Otherwise, collection of diagnostic information is suppressed at step 706.

[0094] It should be understood that, as reference Figures 5 to 7 The collection / suppression of diagnostic information described above represents an example of how the collection of diagnostic information can be filtered. In other examples, some basic diagnostic information may be generated regardless of the SIMD processing configuration information (e.g., operating mode), and additional diagnostic information may be selectively generated based on the SIMD processing configuration information, such as to provide more detailed information.

[0095] Figure 8 The performance monitoring circuit is illustrated, which represents an example of the diagnostic information collection circuit 60 to which the present technology can be applied. Figure 8 As shown, performance monitoring circuitry 70 includes a plurality of event counters 42, each of which maintains a corresponding event count value 43. The performance monitoring circuitry also includes control circuitry 44 that configures how the event counters behave based on counter configuration information 46 set by a user. For example, counter configuration information 46 may be state information stored in registers 14 of the processor (e.g., system registers), may be stored in memory-mapped registers implemented as distinct hardware separate from memory systems 30, 32, 34, or may be stored in memory systems 30, 32, 34 themselves (where memory-mapped registers or data structures in memory themselves are used to provide counter configuration information, control circuitry 44 may access those registers / structures based on a base address programmable by the user). Thus, generally speaking, a programming interface is provided to allow a user (e.g., a software developer performing debugging) to program counter configuration information 46 so that event counters 42 can be configured to collect various types of performance monitoring information of interest when debugging a particular program executing on processing circuitry 4. For example, debugging software may be executed to set the counter configuration information. The target program being debugged may then be executed. During the execution of the target program, the performance monitoring circuit 70 operates according to the previously set counter configuration information.

[0096] The performance monitoring circuit 70 includes an event selection circuit 48 that receives a plurality of event signals 45 from the processing circuit 4 or other components of the data processing system 2, the plurality of event signals indicating the status of events of corresponding types. Figure 8 4. The event selection circuit is shown as a single logic block in FIG. 4, but the event selection circuit may include a separate event selector for each event counter that independently selects the event signal 45 to be monitored by the corresponding event counter.

[0097] For example, event signals may be generated to indicate a wide variety of types of information about various components of the data processing apparatus 2 .

[0098] Some event signals may indicate the occurrence of a particular action (or a count of how many times that action has occurred). For example, such actions may include any of the following:

[0099] After a clock cycle;

[0100] Execute instructions (either general instructions or specific types of instructions);

[0101] Making a memory access request (any general memory access, or a specific type of memory access, such as a load or store);

[0102] A cache access, cache line fill, or cache miss occurs (in some cases this may be specific to a particular level or type of cache);

[0103] A TLB access, TLB line backfill, or TLB miss occurs (again, this can be an event tracked generally for any TLB, or can be specific to a particular TLB instance (e.g., data-side TLB or instruction-side TLB) or a particular TLB level (e.g., level 1 or level 2));

[0104] A branch misprediction occurs;

[0105] A queue or buffer becomes full (variants of which may be provided for specific buffers such as the instruction issue queue, load buffer, store buffer, etc.); or

[0106] • A pipeline stall occurs due to a specific reason (e.g., a cache miss, a TLB miss, or a load or store buffer becoming full).

[0107] Other event signals may specify quantitative information that provides a quantitative status value indicating an attribute of the event that has occurred, such as:

[0108] The number of page table walk cycles performed to fill the TLB;

[0109] The number of cycles it takes to serve a line backfill request to bring the data into the cache following a cache miss; or

[0110] An indication of the current occupancy of a particular queue or buffer.

[0111] It should be understood that the above list of event types is not exhaustive and that a wide variety of event types may be monitored.

[0112] The counter configuration information 46 includes event type assignment information that specifies the event type to be monitored by each event counter 42. For example, each event counter 42 may have a corresponding event type field in the counter configuration information that has an encoding that selects which event signal 45 is to be used for the particular event counter 42. For each event counter, the event selection circuit 48 selects, from the event signals 45, one event signal to be passed to the corresponding event counter 42 as an event status indication 47 based on the event type assignment information for that counter. The event status indication represents the status of the event assigned to the event counter 42 by the counter configuration information 46.

[0113] The counter configuration information 46 also includes filtering configuration information 52 to allow a user to filter updates to the event counter 42 based on the SIMD processing configuration information that defines the array size to be used. The filtering configuration information 52 may be used, for example, Figure 6 The encoding shown indicates in which operating mode or modes the performance monitoring circuitry 70 is to count events for a particular event counter 42 associated with an item of the mode filtering configuration information 52. When the processing circuitry 4 is operating in an operating mode in which the mode filtering configuration information 52 indicates that counting is to be suppressed (or otherwise filtered), the performance monitoring circuitry 70 will not increment the event count value 43 of one or more of the associated event counters 42. The filtering configuration information 52 may also or alternatively specify certain array sizes for which event counters 43 are to be updated, such that updates to event counters 43 associated with other array sizes are suppressed.

[0114] Counter configuration information 46 also includes scaling configuration information 54, which a user can use to indicate a scaling factor to be applied to the count. Based on scaling configuration information 54, performance monitoring circuitry 70 can adjust the amount of updates to event counter 42. Thus, when it is expected that the event count will require a large number of updates or that the event count is likely to be particularly high, the count can be scaled down, for example, to prevent event counter 42 from overflowing and reduce the demand on downstream circuitry that operates on the event count value. In some cases, the scaling factor may depend on the size of the array used by the processing circuitry. For example, if a count of the number of data items operated on is maintained while performing SIMD processing, to avoid excessively large updates to the event count value, the count can be scaled down by the array size used by the processing circuitry. This can be a convenient way to scale the count, as performance monitoring circuitry 70 only needs to count the number of SIMD operations performed by the processing circuitry. However, if the processing circuitry uses different array sizes while performing SIMD processing, it can be difficult or impossible to recover the number of data items operated on from the scaled count. Thus, applying a scaling factor can be used in conjunction with filtering so that events are counted only in specific operating modes or with specific array sizes, making it easier to recover the information of interest from the scaled counts.

[0115] For each event counter 42, a set of hardware circuit logic is provided, including: a storage circuit for storing a corresponding event count value 43; and a counter control logic circuit (implemented in hardware) for updating the event count value based on an event status indication 47 provided to the counter 42 by an event selection circuit 48. For example, an increment value may be selected based on the event status indication 47, and a new value of the event counter value 43 may be calculated by adding the increment value to the previous value of the event counter value 43. Control signals 49 may be provided to each event counter 42 by a control circuit 44 based on counter configuration information 46. These control signals 49 may configure how a given counter selects a function to be applied to the event status indication 47, and how an increment value is selected based on the result of applying the function to the event status indication 47.

[0116] The performance monitoring circuit 70 provides an event counter read interface 50 that allows software to read the event count value of each counter 42. For example, the read interface 50 can be provided by exposing each event count value 43 to the software as a system register that can be read by a system register read instruction executed by the processing circuit 4. Alternatively, the event count value 43 of each event counter 42 can be exposed through a memory mapping interface so that the event count value can be read by software executing a load instruction that specifies a memory address mapped to the storage location where the corresponding event count value 43 is stored. In summary, debugging software can read the current value of each event count value to determine information about events that occurred when the target software was processed by the processing circuit. In use, for example, when the target software has reached a desired point that needs to be investigated (e.g., a desired instruction address reached in the program flow, or a desired data address accessed by a memory access instruction), the debugging software can use a breakpoint or a watchpoint to trigger an exception. Then, when the exception is triggered, the exception handler provided by the debugging software can read the event count value 43 and analyze the information provided by each event count value 43 to determine what occurred. This can be used to diagnose potential performance inefficiencies in program code to help identify possible improvements that can be made to the executing program code to allow the program code to run more efficiently.

[0117] Figure 9 An example of an apparatus 2 having profiling circuitry that may be used to implement the present technique is illustrated. Figure 9 Many components in Figure 1 A complete discussion of these components will not be repeated here.

[0118] To assist in software development, apparatus 2 includes hardware resources that allow for the collection of profiling information about the behavior of instructions processed by processing pipeline 4. Software developers can use this profiling information to perform code optimizations, thereby modifying their code to run more efficiently. Sampling circuitry 82 is provided to select certain instructions, or micro-operations, as sampled operations, which are profiled by profiling circuitry 84. Sampling circuitry 82 can select sampled operations at different stages of the pipeline. For example, sampling circuitry 82 can select certain fetched instructions as sampled operations and mark those fetched instructions at fetch stage 6 to indicate that profiling circuitry 84 should collect information about the behavior of the sampled operations as they progress down pipeline 4. Alternatively, marking of instructions for sampled operations by the sampling circuitry can be performed at decode stage 10 or at a later stage. Furthermore, sampled operations can be selected based on the granularity of individual micro-operations (rather than the granularity of architectural program instructions fetched from memory).

[0119] The sampling circuitry may use an interval counter to count instructions or micro-operations to determine when the next sampling operation should be selected. The sampling interval may be defined by a user-configurable parameter in a control register, or may be fixed to a specific interval. The interval counter counts the number of operations processed by the processing circuitry 4. The operations counted may be fetched instructions, decoded instructions, or decoded micro-operations. If the sample interval has passed, the next operation (e.g., a fetched instruction, a decoded instruction, or a decoded micro-operation) is marked as a sampling operation. For example, a flag bit associated with the instruction may be set, and the flag bit may be advanced down the pipeline 4 along with the instruction or micro-operation selected as the sampling operation. The sampling circuitry 82 then resets the counter once again based on the new sampling interval.

[0120] An advantage of selecting only a subset of operations as sampled operations is that this significantly reduces the overhead of tracking information for profiling. For example, the sampling interval can be set long enough so that, in practice, only a single operation is selected at a time from the operations running within pipeline 4 as the sampled operation, so that profiling circuitry 84 only requires hardware resources sufficient to track the behavior of a single sampled operation at a time. This avoids the overhead of having to index a storage structure that stores information for multiple operations based on an operation identifier associated with a particular sampled operation, in order to select which entry in the storage structure to update based on the information for that particular sampled operation.

[0121] However, other implementations may choose to incur significant hardware costs and may choose to support the selection of multiple sampling operations simultaneously. Compared to implementations that would attempt to trace every instruction, in those embodiments, sampling a subset of the sampled operations still has the advantage of significantly reducing the amount of profiling information generated, making it feasible to capture a wider range of data for each sampled operation, thereby allowing for more meaningful analysis.

[0122] The profiling circuitry 84 is configured to gather information about the behavior of the sampled operations selected by the sampling circuitry 82. The monitoring circuitry may include event detection circuitry for detecting various types of events occurring during the sampled operations. The type of event detected may depend on the type of sampled operation. For example, for a branch operation selected as the sampled operation, the event may track whether a branch prediction error occurred, or whether the branch predictor 7 correctly predicted the branch. For load / store operations, the event may include, for example, whether the load / store operation missed in a particular level of the cache 30, 32, whether the address translation lookup for the load / store instruction missed in the TLB or in a particular level of the TLB, or whether an address error occurred for a load / store instruction. Other types of events that may be monitored may include instruction fetch misses in the instruction cache 8, errors such as undefined instruction exceptions, or delays in certain instructions due to resource contention. The monitoring circuitry may also capture information about specific instructions, such as the instruction address of a sampled operation, the target address of a load / store operation, or the branch target address of a branch operation, as well as an item of architectural state captured from register 14 at the point in time when the sampled operation reaches a certain processing stage (e.g., a context identifier identifying the processing context in which the sampled operation is processed). The monitoring circuitry may also have a cycle counter that counts the number of processing cycles required to complete certain operations, such as measuring the latency of an address translation or cache lookup, or the number of cycles for an operation performed between a first processing point and a second processing point, for example. It will be appreciated that the monitoring circuitry may collect a variety of information.

[0123] The captured monitoring information may be recorded in sample records stored in sample record storage circuitry (e.g., registers or buffers) of the profiling circuitry 84. Within a sample record captured for a given sample operation, the record may specify the type of operation associated with the sample operation (e.g., whether it is a branch, load / store operation, vector processing or matrix processing operation, etc.), and may also provide various information directly attributable to the sample operation. The capture of sample records in the sample record storage device is performed in hardware in the context of processing performed on the pipeline 4, and therefore, no specific software instructions need to be executed to collect the information within the sample record.

[0124] The profiling circuit 84 can filter the collection of sample records depending on the SIMD processing configuration information 66. For example, as described herein, the profiling circuit 84 can generate sample records depending on whether the processing circuit 4 is operating in a first mode of operation (in which the array size is determined by the first array size definition state 62) or a second mode of operation (in which the array size is determined using the second array size definition state 64). In some cases, the profiling circuit 84, in response to the SIMD processing configuration information 66 specifying the array size to be used by the SIMD processing performed by the processing circuit 4, filters the collection of sample records based on that array size. Here, filtering the collection of sample records can correspond, for example, to selectively inhibiting or collecting the sample records based on the SIMD processing configuration information (e.g., array size / mode of operation), or the profiling circuit 84 can determine the amount of information included in the sample records depending on the SIMD processing configuration information.

[0125] In some cases, the profiling circuit 84 also includes information in the sample records relating to the array size used to perform the SIMD operation. The profiling circuit 84 can include an indication of the mode of operation or the array size used in the sample records.

[0126] The sample records can be made available for diagnostic analysis by writing the sample records to a profiling buffer structure stored in the memory system 30, 32, 34. When writing the sample records to the profiling buffer in memory, it can not be necessary to interrupt processing on the pipeline 4, so that no special software instructions are required to cause the sample records to be stored to the memory system. Thus, a number of sample records can be output to the profiling buffer without any interruption occurring, until a sufficient number of sample records have been generated and written out that the profiling buffer is at risk of overflowing, at which point a performance monitoring interrupt can be triggered to interrupt processing so that the exception handler can take action to ensure that the sample records previously stored to the profiling buffer remain accessible for diagnostic analysis. Alternatively, the sample records can be output to a trace buffer, which is a dedicated hardware structure separate from the memory system 30, 32, 34 for storing diagnostic information on-chip, and / or outputting the captured sample records through a trace output port (directly or via the trace buffer), where the trace output port is a set of integrated circuit pins through which the sample records can be output to an external off-chip trace analyzer or storage device.

[0127] Figure 10 An example is illustrated in which multiple processors 945, 940 share the matrix processing circuit 950. As shown, there are two processors CPU 0 940 and CPU 1 945. Although only the execution circuit 16, 966 is depicted in Figure 10 Figure 10 each processor can contain​ Figure 1 and / or Figure 9 The execution circuitry 16, 966 contains a scalar processing unit 20, 970, a vector processing unit 22, 972, and a load / store unit 28, 978, as described above with respect to Figure 1 972 . However, matrix processing functionality is provided by a matrix processing circuit 950 external to the individual processors 940, 945. The matrix processing circuit 950 is arranged to perform matrix processing and vector processing and has vector registers 952 and matrix registers 956 to store vector and matrix data, respectively, to support the processing. In some cases, the matrix processing circuit 950 will be operable to perform all vector processing operations supported by the vector processing units 22, 972; however, in some cases, the matrix processing circuit 950 may only support a more limited range of vector processing operations. Therefore, for those vector processing operations, the respective processors 940, 945 may use their respective vector processing units 22, 972, or may use the matrix processing unit 950 via the matrix processing interface 924, 974. Using the vector processing unit 22, 972 may correspond to one mode of operation as described herein, wherein using the matrix processing unit 950 to perform vector processing corresponding to another mode of operation. One or both of the vector processing units 22, 972 and the matrix processing unit 950 may support a range of vector lengths (i.e., array sizes) for performing vector processing, with the vector lengths used being specified by the first array size definition state and the second array size definition state. It will be appreciated that the processing units shared between processors need not be matrix processing units. For example, in some examples the processors may each have a dedicated vector processing unit and may access a shared vector processing unit. In practice, various arrangements of SIMD processing units may be made in which one or more SIMD processing units are shared between multiple processors.

[0128] The concepts described herein may be embodied in computer-readable code for fabricating devices embodying the described concepts. For example, the computer-readable code may be used in one or more stages of a semiconductor design and fabrication process, including an electronic design automation (EDA) stage, to fabricate integrated circuits including devices embodying the concepts. The computer-readable code above may additionally or alternatively enable the definition, modeling, simulation, verification, and / or testing of devices embodying the concepts described herein.

[0129] For example, computer readable code for making an apparatus embodying the concepts described herein can be embodied in code that defines a hardware description language (HDL) representative of the concepts. For example, the code can define a register transfer level (RTL) abstraction that is used to define one or more logic circuits that embody the concepts. The code can define an HDL representative of the one or more logic circuits that embody the apparatus in Verilog, SystemVerilog, Chisel, or VHDL (very high speed integrated circuit hardware description language), as well as intermediate representations such as FIRRTL. Computer readable code can provide a definition embodying the concepts using a system level modeling language such as SystemC and SystemVerilog, or other behavioral representations of the concepts that can be interpreted by a computer to implement simulation, functional and / or formal verification, and testing of the concepts.

[0130] Additionally or alternatively, the computer readable code can define a detailed description of an integrated circuit component embodying the concepts described herein, such as one or more netlists or integrated circuit layout definitions, including representations such as GDSII. One or more netlists or other computer readable representations of the integrated circuit component can be generated by applying one or more logic synthesis processes to an RTL representation, thereby generating a definition for making an apparatus embodying the invention. Alternatively or additionally, one or more logic synthesis processes can generate a bitstream from the computer readable code that is loaded into a field programmable gate array (FPGA) to configure the FPGA to embody the described concepts. The FPGA can be deployed for the purpose of verifying and testing the concepts in an integrated circuit prior to manufacturing, or the FPGA can be deployed directly in a product.

[0131] The computer readable code can include a mix of code representations for making an apparatus, such as a mix including one or more of an RTL representation, a netlist representation, or another computer readable definition for a semiconductor design and manufacturing process to make an apparatus embodying the invention. Alternatively or additionally, the concepts can be defined in a combination of a computer readable definition for making an apparatus in a semiconductor design and manufacturing process and computer readable code defining instructions to be executed by the defined apparatus after being made.

[0132] Such computer readable code can be provided in any known transient computer readable medium, such as sending the code over a network wired or wirelessly, or in a non-transient computer readable medium, such as a semiconductor, magnetic or optical disk. An integrated circuit made using the computer readable code can include components such as one or more of a central processing unit, a graphics processing unit, a neural processing unit, a digital signal processor, or other components that individually or collectively embody the concepts.

[0133] Furthermore, Figure 11Embodiments of simulators that can be used are illustrated. While the embodiments described earlier implement the application in terms of apparatus and methods for operating specific hardware that support the technology of interest, it is also possible to provide an instruction execution environment in accordance with the embodiments described herein, implemented by use of a computer program. Such computer programs provide a software-based implementation of a hardware architecture, and are therefore often referred to as simulators. Classes of simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. In general, a simulator implementation can run on a host processor 930, optionally running a host operating system 920, that supports the simulator program 910. In some arrangements, there can be multiple layers of simulation between the hardware and the provided instruction execution environment and / or multiple distinct instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator embodiments that execute at a reasonable speed, but such approaches can be reasonably used in certain situations, such as when it is desirable to run code native to another processor for compatibility or re-use reasons. For example, a simulator implementation can provide an instruction execution environment with additional functionality not supported by the host processor hardware, or provide an instruction execution environment generally associated with a different hardware architecture. For a review of simulation, see "Some Efficient Architecture Simulation Techniques", Robert Bedichek, Winter 1990 USENIX Conference, pages 53-63.

[0134] Where embodiments have been previously described with reference to specific hardware architectures or features, equivalent functionality can be provided by suitable software architectures or features in a simulated embodiment. For example, specific circuitry can be implemented as computer program logic in a simulated embodiment. Similarly, memory hardware such as registers or caches can be implemented as software data structures in a simulated embodiment. Where one or more of the hardware elements mentioned in previously described embodiments are present on a host hardware (e.g., host processor 930), some simulated embodiments can make use of the host hardware as appropriate.

[0135] The emulator program 910 can be stored on a computer readable storage medium (which can be a non-transitory medium) and provide a program interface (instruction execution environment) to the target program code 900 (which can include an application, operating system, and hypervisor) that is identical to the interface of the hardware architecture being modeled by the emulator program 910. Thus, program instructions of the target code 900, including instructions to set the operating mode, can be executed within the instruction execution environment using the emulator program 910 so that a host computer 930 that does not actually have the hardware features of the device 2 discussed above can emulate such features. The functions of the diagnostic information collection circuit 60 can be emulated by corresponding program logic 916. By providing emulation of the device shown in software, it is possible to develop debug software that interacts with the diagnostic information collection circuit 60 before the hardware is actually available. Figure 1

[0136] Thus, the emulator program 910 can have processing program logic 912 that emulates the state of the processing circuit 4 described above. For example, the processing program logic 912 can control transitions of the execution state (e.g., exception level, operating mode) in response to events that occur during the emulated execution of the target code 900. Instruction decode program logic 914 decodes instructions of the target code 900 and maps these instructions to a corresponding instruction set within the native instruction set of the host device 930. Register emulation logic 913 maps register accesses requested by the target code to accesses of a corresponding register emulation data structure 933 maintained by the host hardware of the host device 930, such as by accessing data in a register or memory 932 of the host device 930. Memory management program logic 915 implements address translation, page table walks, and access control checks in a manner corresponding to the MMU in a hardware implemented embodiment, but also has the additional function of mapping emulated physical addresses obtained by the emulated MMU 915 to host virtual addresses for accessing the host memory 932. These host virtual addresses can themselves be translated to host physical addresses using standard address translation mechanisms supported by the host (conversion of host virtual addresses to host physical addresses is outside the scope of the emulator program 910 control). Thus, the emulated physical address space accessed by the target code 900 can be mapped to a region 934 of the host memory 932 that represents the emulated target memory 8, 30, 32, 34 of the target processing device 2 being emulated by the emulator program 910. This emulated target address space 934 can be used to store diagnostic information 935 that is read out to the memory system 8, 30, 32, 34 from diagnostic information stored on-chip and represented by the emulated diagnostic information memory 936 in the corresponding hardware device.

[0137] The emulator program 910 has diagnostic information collection program logic 916 that emulates the behavior of the diagnostic information collection circuit 60. ​

[0138] In this application, the phrase "configured to" is used to mean that an element of a device has a configuration able to perform the defined operation. In this context, a "configuration" means an arrangement or manner of interconnection of hardware or software. For example, the device can have dedicated hardware which provides the defined operation, or a processor or other processing device can be programmed to perform the function. "Configured to" does not imply that the device element needs to be changed in any way in order to provide the defined operation.

[0139] While exemplary examples of the application have been described herein in detail with reference to the attached drawings, it is to be understood that the application is not limited to those precise examples, and that various changes and modifications can be effected therein by one of ordinary skill in the art without departing from the scope and spirit of the application as defined by the appended claims.

Claims

1. A device, comprising: processing circuitry operable to perform single instruction multiple data (SIMD) processing on an array comprising a plurality of data items, the processing circuitry supporting the SIMD processing for a plurality of array sizes; wherein the processing circuitry is configured to select an array size for performing the SIMD processing based at least in part on SIMD processing configuration information; diagnostic information collection circuitry configured to collect diagnostic information about software executing on the processing circuitry; The diagnostic information collection circuit is configured to filter the collection of the diagnostic information based on the SIMD processing configuration information.

2. The device according to claim 1, wherein: Filtering the collection of the diagnostic information based on the SIMD processing configuration information includes collecting the diagnostic information or suppressing the collection of the diagnostic information depending on the SIMD processing configuration information.

3. The device according to claim 1 or claim 2, wherein: the processing circuitry being operable in a first operating mode and a second operating mode, wherein in the first operating mode the processing circuitry is arranged to perform the SIMD processing with an array size determined based on a first array size definition state, and in the second operating mode the processing circuitry is arranged to perform the SIMD processing with an array size determined based on a second array size definition state; The SIMD processing configuration information indicates whether the SIMD processing circuitry is operating in the first operating mode or the second operating mode; and The diagnostic information collection circuitry is configured to filter the collection of the diagnostic information based on whether the processing circuitry is operating in the first operating mode or the second operating mode. 4 . The apparatus of claim 3 , wherein the diagnostic information collection circuitry is configured to include, in the diagnostic information collected for the SIMD processing, an indication of whether the processing circuitry is operating in the first operating mode or the second operating mode.

5. The apparatus according to any preceding claim, further comprising: A mode screening configuration storage circuit, wherein the mode screening configuration storage circuit is used to store a mode screening configuration state; The diagnostic information collection circuit collects the diagnostic information when the processing circuit is in the first operating mode in response to the mode filter configuration state having a first value, and refrains from collecting the diagnostic information when the processing circuit is in the second operating mode.

6. The apparatus of claim 5 , wherein the diagnostic information collection circuit collects the diagnostic information when the processing circuit is in the second operating mode in response to the mode filter configuration state having a second value, and suppresses collection of the diagnostic information when the processing circuit is in the first operating mode.

7. The apparatus of claim 5 or claim 6, wherein the diagnostic information collection circuit collects the diagnostic information when the processing circuit is in the first operating mode and when the processing circuit is in the second operating mode in response to the mode filter configuration state having a third value.

8. An apparatus according to any preceding claim, wherein: The processing circuitry is arranged to perform the SIMD processing with an array size specified by the SIMD processing configuration information; and The diagnostic information collection circuitry is configured to filter collection of the diagnostic information based on the array size specified by the SIMD processing configuration information.

9. An apparatus according to any preceding claim, wherein the diagnostic information collection circuitry is configured to include, in the diagnostic information collected for the SIMD processing, the array size used to perform the SIMD processing.

10. The apparatus according to any preceding claim, further comprising: An array size screening configuration storage circuit is used to store an array size screening configuration state, wherein: the diagnostic information collection circuitry collecting the diagnostic information in response to the array size screening configuration state indicating that diagnostic information is to be collected for a particular array size used by the processing circuitry to perform the SIMD processing; and The diagnostic information collection circuitry suppresses the collection of the diagnostic information in response to the array size screening configuration state specifying that diagnostic information is to be collected for one or more array sizes other than the particular array size used by the processing circuitry to execute the SIMD processing circuitry.

11. An apparatus according to any preceding claim, wherein: The diagnostic information collection circuit includes a performance monitoring circuit, the performance monitoring circuit being configured to monitor the performance of the software executed on the processing circuit; The performance monitoring circuit includes a plurality of event counters, each event counter maintaining a corresponding event count value based on monitoring events during execution of the software on the processing circuit; The diagnostic information collection circuit includes a control circuit configured to configure the event counter based on counter configuration information, the counter configuration information including filter configuration information; and The performance monitoring circuitry is configured to filter updates to a given event counter for an event based on the filtering configuration information and the SIMD processing configuration information.

12. The apparatus according to any preceding claim, further comprising: a sampling circuit configured to select a subset of instructions or micro-operations processed by the processing circuit as sampled operations to be profiled; and Profiling circuitry is configured to capture a sample record in response to processing by the sampling circuitry of an instruction or micro-operation selected as a sample operation, the sample record including diagnostic information indicative of behavior of the sample operation.

13. The apparatus of claim 12, wherein the profiling circuitry is configured to include an indication of at least a portion of the SIMD processing configuration information in the sample record.

14. An apparatus according to claim 12 or claim 13, wherein the profiling circuitry is configured to selectively suppress the capturing of sample records in dependence on the SIMD processing configuration information.

15. An apparatus according to claim 3 or any claim dependent thereon, wherein the processing circuit is configured to: use a first SIMD processing circuit to perform the SIMD processing in the first operating mode; and use a second SIMD processing circuit to perform the SIMD processing in the second operating mode.

16. The apparatus of claim 15, wherein the first array size definition state defines an array size for the first SIMD processing circuitry, and the second array size definition state defines an array size for the second SIMD processing circuitry.

17. Apparatus according to claim 15 or claim 16, wherein at least one of the first array size definition state and the second array size definition state defines a software-requested array size for the corresponding SIMD processing circuitry; and The respective SIMD processing circuitry is configured to perform SIMD processing with an array size determined based on the requested array size and at least one of: Software-configurable maximum array size; Software-configurable minimum array size; The maximum array size supported by the hardware; and The minimum array size supported by the hardware.

18. The apparatus of any one of claims 15 to 17, wherein at least one of the first SIMD processing circuit and the second SIMD processing circuit is one of a vector processing circuit and a matrix processing circuit.

19. The device according to any one of claims 15 to 18, wherein: At least one of the first SIMD processing circuitry and the second SIMD processing circuitry is shared among a plurality of processors.

20. A device comprising: The device according to any one of claims 15 to 19; a first SIMD processing circuit; and A second SIMD processing circuit.

21. A computer readable medium for storing computer readable code for producing an apparatus according to any preceding claim.

22. A method for collecting diagnostic information, the method comprising: collecting said diagnostic information about software executing on the processing circuitry; wherein the processing circuitry is operable to perform single instruction multiple data (SIMD) processing on an array comprising a plurality of data items, the processing circuitry supporting the SIMD processing for a plurality of array sizes, wherein the processing circuitry is configured to select an array size for performing the SIMD processing based at least in part on SIMD processing configuration information; and The collecting of the diagnostic information includes filtering the collection of the diagnostic information based on the SIMD processing configuration information.

23. A computer program comprising instructions which, when executed by a host data processing device, control the host data processing device to provide an instruction execution environment for executing target program code, the computer program comprising: processing program logic operable to perform single instruction multiple data (SIMD) processing on an array comprising a plurality of data items, the processing program logic supporting the SIMD processing for a plurality of array sizes; wherein the processing program logic is configured to select an array size for performing the SIMD processing based at least in part on SIMD processing configuration information; diagnostic information collector logic for collecting diagnostic information about software executing on the processing program logic; The diagnostic information collection program logic is configured to filter the collection of the diagnostic information based on the SIMD processing configuration information.

24. A computer-readable medium storing the computer program according to claim 23.