Distributed Hardware Tracing
The method addresses the complexity of analyzing distributed software performance by monitoring and correlating hardware events across multiple processor cores, enhancing computational efficiency and aiding in performance analysis.
Patent Information
- Application Number
- JP2023045549
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-03-29
- Filing Date
- 2023-03-22
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2037-10-20
AI Technical Summary
Analyzing the performance of distributed software executed across multiple processor cores is complex due to the lack of efficient methods for correlating and tracing hardware events across distributed hardware components.
A computer-implemented method that monitors the execution of program code by multiple processor components, stores data identifying hardware events, and generates a data structure to arrange these events chronologically for performance analysis.
Enables efficient correlation of hardware events, improving computational efficiency and aiding in debugging and performance analysis by providing a chronological arrangement of events across distributed processor units.
Smart Images

Figure 0007700167000001 
Figure 0007700167000002 
Figure 0007700167000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Patent Application No. 15 / 472, 932, filed on March 29, 2017, entitled "Synchronous Hardware Event Collection", and Attorney Docket No. 16113 - 8129001. The entire disclosure of U.S. Patent Application No. 15 / 472,932 is hereby incorporated by reference in its entirety.
Background Art
[0002] Background This specification relates to analyzing the execution of program code.
[0003] Effective performance analysis of distributed software executed within distributed hardware components can be a complex task. Distributed hardware components may be the respective processor cores of two or more central processing units (CPUs) (or graphics processing units (GPUs)) that cooperate and interact to execute a larger software program or a part of a program code. Device (Graphics Processing Unit: GPU)).
[0004] From a hardware perspective (e.g., within a CPU or GPU), the information or functions available for performance analysis generally include 1) hardware performance counters and 2) hardware event traces.
Summary of the Invention
Means for Solving the Problems
[0005] Summary Generally, one aspect of the subject matter described in this specification can be embodied in a computer-implemented method executed by one or more processors. The method includes monitoring the execution of program code by a first processor component, the first processor component being configured to execute at least a first portion of the program code, and the method further includes monitoring the execution of program code by a second processor component, the second processor component being configured to execute at least a second portion of the program code.
[0006] The method further includes storing, in at least one memory buffer, data that identifies one or more hardware events that occur across a processor unit that includes the first processor component and the second processor component by a computing system. Each hardware event represents at least one of a memory access operation of the program code, an issued instruction of the program code, or a data communication associated with an executed instruction of the program code. The data that identifies each of the one or more hardware events includes metadata that characterizes the hardware event and a hardware event timestamp. The method includes generating, by the computing system, a data structure that identifies the one or more hardware events, the data structure being configured to arrange the one or more hardware events in a chronological order of events associated with at least the first processor component and the second processor component.
[0007] The method further includes storing, in a memory bank of a host device, the generated data structure for use in analyzing the performance of program code executed by at least the first processor component or the second processor component.
[0008] These and other implementations are each optional and may include one or more of the following features. For example, in some implementations, the method further includes the computing system detecting a trigger function associated with a portion of program code executed by at least one of the first processor component or the second processor component, and in response to detecting the trigger function, the computing system initiating at least one trace event that causes data associated with one or more hardware events to be stored in at least one memory buffer.
[0009] In some implementations, the trigger function corresponds to at least one of a specific sequence step in the program code or a specific time parameter indicated by a global time clock used by the processor unit, and the step of initiating at least one trace event includes determining that a trace bit is set to a specific value, and at least one trace event is associated with a memory access operation including a plurality of intermediate operations occurring across processor units, and in response to determining that the trace bit is set to a specific value, data associated with the plurality of intermediate operations is stored in one or more memory buffers.
[0010] In some implementations, the step of storing data identifying one or more hardware events includes storing a first subset of data identifying the hardware events of the one or more hardware events in a first memory buffer of the first processor component. The storing step occurs in response to the first processor component executing a hardware trace instruction associated with at least a first portion of the program code.
[0011] In some implementation examples, the step of storing data for identifying one or more hardware events further includes storing a second subset of data for identifying the hardware events of the one or more hardware events in a second memory buffer of a second processor component. The storing step occurs in response to the second processor component executing hardware trace instructions associated with at least a second portion of the program code.
[0012] In some implementation examples, the step of generating a data structure further includes the computing system comparing at least the hardware event timestamps of each event in a first subset of data for identifying hardware events with at least the hardware event timestamps of each event in a second subset of data for identifying hardware events, and the computing system providing a correlated set of hardware events for presentation in the data structure, based at least in part on the comparison of each event in the first subset with each event in the second subset.
[0013] In some implementation examples, the generated data structure identifies at least one parameter indicating the latency attribute of a particular hardware event, and the latency attribute indicates at least the duration of the particular hardware event. In some implementation examples, at least one processor of the computing system is a multi-core multi-node processor having one or more processor components, and the one or more hardware events correspond at least in part to data transfers occurring between a first processor component of at least a first node and a second processor component of a second node. having one or more processor components, and the one or more hardware events correspond at least in part to data transfers occurring between a first processor component of at least a first node and a second processor component of a second node.
[0014] In some implementation examples, the first processor component and the second processor component are one of the processors, processor cores, memory access engines, or hardware functions of a computing system, and one or more hardware events partially correspond to the movement of data packets between a source and a destination, and the metadata characterizing the hardware events corresponds to at least one of a source memory address, a destination memory address, a unique trace identification number, or a size parameter associated with a direct memory access (DMA) trace.
[0015] In some implementation examples, a specific trace ID number is associated with a plurality of hardware events occurring between processor units, the plurality of hardware events correspond to a specific memory access operation, the specific trace ID number is used to correlate one or more of the plurality of hardware events, and is used to determine the latency attribute of the memory access operation based on the correlation.
[0016] Other aspects of the subject matter described in this specification can be embodied in a distributed hardware tracing system that includes one or more processors including one or more processor cores and one or more machine-readable storage units for storing instructions. The instructions are executable by one or more processors to perform operations, the operations include monitoring the execution of program code by a first processor component, the first processor component is configured to execute at least a first portion of the program code, the operations further include monitoring the execution of program code by a second processor component, and the second processor component is configured to execute at least a second portion of the program code.
[0017] The method further includes storing, in at least one memory buffer, data that identifies one or more hardware events occurring across a processor unit that includes a first processor component and a second processor component. Each hardware event represents at least one of a memory access operation of program code, an issued instruction of program code, or data communication associated with an executed instruction of program code. The data that identifies each of the one or more hardware events includes metadata that characterizes the hardware event and a hardware event timestamp. The method includes the computing system generating a data structure that arranges the one or more hardware events in a chronological order of events associated with at least the first processor component and the second processor component.
[0018] The method further includes storing the generated data structure in a memory bank of a host device for use in analyzing the performance of program code executed by at least the first processor component or the second processor component.
[0019] Other realizations of this and other aspects include corresponding systems, apparatuses, and computer programs configured to perform the actions of the method encoded on a computer storage device. One or more computer systems can be so configured by software, firmware, hardware, or combinations thereof installed in the system to cause the system to perform actions during operation. One or more computer programs can be so configured by having instructions that, when executed by a data processing device, cause the device to perform actions. The one or more computer programs can be so configured by having instructions that, when executed by a data processing device, cause the device to perform actions.
[0020] The subject matter described in this specification can be realized in certain embodiments so as to achieve one or more of the following advantages. The hardware tracing system described enables efficient correlation of hardware events that occur during the execution of a distributed software program by a distributed processing unit that includes a multi-node multi-core processor. The hardware tracing system described further includes mechanisms that enable collection and correlation of hardware events / trace data in multiple cross-node configurations.
[0021] The hardware tracing system increases computational efficiency by using dynamic triggers that are executed through hardware knobs / functions. Further, hardware events can be serially time-stamped using event descriptors such as unique trace identifiers, event timestamps, event source addresses, and event destination addresses. Such descriptors assist software programmers and processor design engineers in effectively debugging and analyzing software and hardware performance issues that may occur during source code execution.
[0022] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other potential features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
Brief Description of the Drawings
[0023]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
[0024] Like reference numerals and names in the various drawings indicate like elements.
DETAILED DESCRIPTION
[0025] Detailed Description The subject matter described in this specification generally relates to distributed hardware tracing. In particular, a computing system monitors the execution of program code executed by one or more processor cores. For example, the computing system can monitor the execution of program code executed by a first processor core and the execution of program code executed by at least a second processor core. The computing system stores data identifying one or more hardware events in a memory buffer. The stored data identifying the events corresponds to events occurring across a distributed processor unit including at least the first and second processor cores.
[0026] For each hardware event, the stored data includes metadata characterizing the hardware event and an event timestamp. The system generates a data structure identifying the hardware events. The data structure arranges the events in chronological order and associates the events with at least the first or second processor core. The system stores the data structure in a memory bank of the host device and uses the data structure to analyze the performance of the program code executed by the first or second processor core.
[0027] FIG. 1 shows a block diagram of an exemplary computing system 100 for distributed hardware tracing. As used in this specification, distributed hardware system tracing corresponds to the storage of data that identifies events occurring within components and sub-components of an exemplary processor microchip. Further, as used herein, a distributed hardware system (or tracing system) corresponds to a set of processor microchips or processing units that cooperate to execute respective portions of software / program code configured for distributed execution among a set of processor microchips or distributed processing units.
[0028] System 100 may be a distributed processing system having one or more processors or processing units that execute a software program distributively, i.e., by executing different portions of the program code on different processing units of System 100. The processing units may include two or more processors, processor microchips, or processing units, for example, at least a first processing unit and a second processing unit.
[0029] In some implementations, when a first processing unit receives and executes a first portion of the program code of a distributed software program, and a second processing unit receives and executes a second portion of the program code of the same distributed software program, the two or more processing units may be distributed processing units.
[0030] In some implementations, different processor chips of System 100 can form respective nodes of the distributed hardware system. In an alternative implementation, a single processor chip may include one or more processor cores and hardware functions each capable of forming a respective node of the processor chip.
[0031] For example, in the context of a central processing unit (CPU), the processor chip may include at least two nodes, and each node may be a respective core of the CPU. Alternatively, in the context of a graphics processing unit (GPU), the processor chip may include at least two nodes, and each node may be a respective streaming multiprocessor of the GPU. Computing system 100 may include a plurality of processor components. In some implementations, the processor components may be at least one of a processor chip, a processor core, a memory access engine, or at least one hardware component of the entire computing system 100.
[0032] In some examples, a processor component such as a processor core may be a fixed-function component configured to execute at least one specific operation based on at least one issued instruction of the program code being executed. In other examples, a processor component such as a memory access engine (MAE) may be configured to execute program code at a lower level of detail or granularity than the program code executed by other processor components of system 100.
[0033] For example, the program code executed by the processor core can cause an MAE descriptor to be generated and sent to the MAE. After receiving the descriptor, the MAE can execute a data transfer operation based on the MAE descriptor. In some implementations, the data transfer executed by the MAE may include, for example, moving data between components of system 100 via a certain data path or interface component of the system, or issuing a data request to an exemplary configuration bus of system 100.
[0034] In some implementation examples, each tensor node of the exemplary processor chip of system 100 may have at least two "front ends" that can be hardware blocks / functions for processing program instructions. As will be described in more detail below, the first front end can correspond to the first processor core 104, while the second front end can correspond to the second processor core 106. Thus, the first and second processor cores may also be described herein as the first front end 104 and the second front end 106, respectively.
[0035] As used in this specification, a trace chain may be a particular physical data communication bus where trace entries can be placed for transmission to an exemplary chip manager within system 100. The received trace entry may be a data word / structure that includes a plurality of bytes and a plurality of binary values or binary numbers. For this reason, the descriptor "word" refers to a fixed-size piece of binary data that can be treated as one unit by the hardware devices of an exemplary processor core.
[0036] In some implementation examples, the processor chip of the distributed hardware tracing system is a multi-core (i.e., having multiple cores) processor that executes a part of the program code on each core of the chip. In some implementation examples, a part of the program code can correspond to vectorized calculations for the inference workload of an exemplary multi-layer neural network. On the other hand, in alternative implementation examples, a part of the program code can generally correspond to software modules associated with conventional programming languages.
[0037] Computing system 100 generally includes a node manager 102, a first processor core (FPC) 104, and a second processor core (second The system 100 includes a processor core (SPC) 106, a node fabric (NF) 110, a data router 112, and a host interface block (HIB) 114. In some implementations, system 100 may include a memory mux 108 configured to perform signal switching, multiplexing, and demultiplexing functions. System 100 further includes a tensor core 116 within which the FPC 104 is disposed. The tensor core 116 may be an exemplary computing device configured to perform vectorized computations on multi-dimensional data arrays. The tensor core 116 may include a vector processing unit (VPU) 118 which may interact with a matrix unit (MXU) 120, a transpose unit (XU) 122, and a reduction and permutation unit (RPU) 124. In some implementations, the computing system 100 may include one or more execution units of a conventional CPU or GPU, such as a load / store unit, an arithmetic logic unit (ALU), and a vector unit. processing unit:VPU) 118, which may interact with a matrix unit (matrix unit:MXU) 120, a transpose unit (transpose unit:XU) 122, and a reduction and permutation unit (reduction and permutation unit:RPU) 124. In some implementations, the computing system 100 may include one or more execution units of a conventional CPU or GPU, such as a load / store unit, an arithmetic logic unit (ALU), and a vector unit. In some implementations, the computing system 100 may include one or more execution units of a conventional CPU or GPU, such as a load / store unit, an arithmetic logic unit (ALU), and a vector unit.
[0038] The components of system 100 collectively include a large set of hardware performance counters and support hardware that facilitates completion of trace activity within the components. As will be described in more detail below, by each processor core of system 100 The program code to be executed may include an embedded trigger used to enable multiple performance counters simultaneously during code execution. Generally, the detected trigger causes trace data to be generated for one or more trace events. The trace data can correspond to incremental parameter counts that are stored in counters and analyzed to identify performance characteristics of the program code. Data for each trace event can be stored in an exemplary storage medium (e.g., a hardware buffer) and may include a timestamp generated in response to detection of the trigger.
[0039] Furthermore, trace data can be generated for various events occurring within the hardware components of system 100. Exemplary events can include inter-node and cross-node communication operations such as direct memory access (DMA) operations and synchronization flag updates (each described in more detail below). In some implementations, system 100 can include a globally synchronized timestamp counter, generally referred to as a Global Time Counter (GTC). In other implementations, system 100 can include other types of global clocks such as a Lamport clock.
[0040] The GTC can be used for accurate correlation of program code execution with the performance of software / program code executed in a distributed processing environment. Additionally, and somewhat related to the GTC, in some implementations, system 100 can include one or more trigger mechanisms used by distributed software programs to start and stop data tracing in a very coordinated manner in a distributed system.
[0041] In some implementations, host system 126 compiles program code that may include embedded operands. The operands, when detected, trigger the capture and storage of trace data associated with the hardware event. In some implementations, host system 126 provides the compiled program code to one or more processor chips of system 100. In an alternative implementation, the program code may be compiled by an exemplary external compiler (using the embedded trigger) and loaded onto one or more processor chips of system 100. In some examples, the compiler can set one or more trace bits (described below) associated with a certain trigger embedded in a portion of the software instructions. The compiled program code may be a distributed software program executed by one or more components of system 100.
[0042] Host system 126 may include a monitoring engine 128 configured to monitor the execution of the program code by one or more components of system 100. In some implementations, monitoring engine 128 enables host system 126 to monitor the execution of program code executed by at least FPC 104 and SPC 106. For example, during code execution, host system 126 can monitor the performance of the executing code via monitoring engine 128 by receiving, at least, a periodic timeline of hardware events based on the generated trace data. Although a single block is shown for host system 126, in some implementations, system 126 may include multiple hosts (or host subsystems) associated with multiple processor chips or chip cores of system 100.
[0043] In another implementation example, when data traffic crosses the communication path between the FPC 104 and an exemplary third processor core / node, cross-node communication involving at least three processor cores may cause the host system 126 to monitor the data traffic with one or more intermediate "hops". For example, the FPC 104 and the third processor core may be the only cores that execute program code during a given period. Thus, when data is transferred from the FPC 104 to the third processor core, the SPC 106 can generate trace data for the intermediate hop as the data is transferred from the FPC 104 to the third processor core. Stated another way, during data routing in the system 100, data traveling from the first processor chip to the third processor chip may need to cross the second processor chip, and thus the execution of the data routing operation may cause trace entries to be generated for routing activity at the second chip.
[0044] When the compiled program code is executed, the components of the system 100 can interact to generate a timeline of hardware events that occur in a distributed computer system. The hardware events can include in-node and cross-node communication events. Exemplary nodes of the distributed hardware system and their associated communications are described in more detail below with reference to FIG. 2. In some implementation examples, a data structure is generated that identifies a set of hardware events for at least one hardware event timeline. The timeline enables the reconstruction of events that occur in the distributed system. In some implementation examples, event reconstruction can include correct event ordering based on the analysis of timestamps generated during the occurrence of a particular event.
[0045] Generally, an exemplary distributed hardware tracing system may include the above-described components of system 100 and at least one host controller associated with host system 126. The performance or debugging of data obtained from the distributed tracing system can be useful when event data is correlated, for example, in a time series or in an ordered manner. In some implementations, multiple stored hardware events corresponding to connected software modules are stored, and then data correlation can occur when they are ordered for structured analysis by host system 126. For implementations that include multiple host systems, the correlation of data obtained via different hosts may be performed, for example, by a host controller.
[0046] In some implementations, FPC 104 and SP 106 are each a separate core of one multi-core processor chip. On the other hand, in other implementations, FPC 104 and SP 106 are the respective cores of separate multi-core processor chips. As described above, system 100 may include a distributed processor unit having at least FPC 104 and SPC 106. In some implementations, the distributed processor unit of system 100 may include one or more hardware or software components configured to execute at least a portion of a larger distributed software program or program code.
[0047] Data router 112 is an inter-chip interconnect (ICI) that provides a data communication path between the components of system 100. In particular, router 112 can provide a communication coupling or connection between FPC 104 and SPC 106 and between the respective components associated with cores 104, 106. Node fabric 110 interacts with data router 112 to move data packets within the distributed hardware components and sub-components of system 100.
[0048] The node manager 102 is a high-level device that manages the low-level node functions in a multi-node processor chip. As will be described in more detail below, one or more nodes of the processor chip include a chip manager controlled by the node manager 102 to manage hardware event data and store it in a local entry log. It may include. The memory mux 108 is a multiplexing device that can perform switching, multiplexing, and demultiplexing operations on data signals provided to or received from an exemplary external high bandwidth memory (HBM). In some implementations, when the mux 108 switches between the FPC 104 and the SPC 106, exemplary trace entries (described below) may be generated by the mux 108. The memory mux 108 may affect the performance of certain processor cores 104, 106 that do not have access to the mux 108. For this reason, the trace entry data generated by the mux 108 can help understand the spikes that result in the latency of certain system activities associated with each core 104, 106. In some implementations, the hardware event data (e.g., trace points described below) that occurs within the mux 108 can be grouped with the event data for the node fabric 110 in an exemplary hardware event timeline. Event grouping can occur when a certain trace activity causes event data for multiple hardware components to be stored in an exemplary hardware buffer (e.g., the trace entry log 218 described below).
[0049]
[0050] In system 100, the performance analysis hardware includes FPC104, SPC106, mux108, node fabric 110, data router 112, and HIB114. Each of these hardware components or units includes a hardware performance counter and a hardware event trace mechanism and functionality. In some implementations, VPU118, MXU120, XU122, and RPU124 do not include their own dedicated performance hardware. Rather, in such implementations, FPC104 can be configured to provide the necessary counters for VPU118, MXU120, XU122, and RPU124.
[0051] VPU118 can include an internal design architecture that supports localized high-bandwidth data processing and arithmetic operations associated with the vector elements of an exemplary matrix-vector processor. MXU120 is a matrix multiplication unit configured to perform, for example, matrix multiplications of up to 128×128 on a vector data set of multiplicands.
[0052] XU122 is a transpose unit configured to perform, for example, matrix transpose operations of up to 128×128 on vector data associated with matrix multiplication operations. RPU124 can include a sigma unit and a replacement unit. The sigma unit performs sequential reduction on vector data associated with matrix multiplication operations. Reduction can include sums and various types of comparison operations. The replacement unit can completely replace or duplicate all elements of vector data associated with matrix multiplication operations.
[0053] In some implementation examples, the program code executed by the components of system 100 may represent machine learning, neural network inference calculations, and / or one or more direct memory access functions. The components of system 100 may be configured to execute one or more software programs that include instructions to cause one or more functions to be executed on the processing unit or device of the system. The term "component" is intended to include any data processing device, or storage device such as a control status register, or any other device capable of processing and storing data.
[0054] System 100 generally includes a plurality of processing units or devices that may include one or more processors (e.g., microprocessors or central processing units (CPUs)), graphics processing units (GPUs), application specific integrated circuits (ASICs), or combinations of different processors. In alternative embodiments, system 100 may each include other computing resources / devices (e.g., cloud-based servers) that provide additional processing options for performing calculations related to the hardware trace functions described herein.
[0055] The processing unit or device may further include one or more memory units or memory banks (e.g., registers / counters). In some implementations, the processing unit executes programmed instructions stored in memory for the system 100's devices to perform one or more of the functions described in this specification. The memory unit / bank may include one or more non-transitory machine-readable storage media. Non-transitory machine-readable storage media may include solid-state memory, magnetic disks, and optical disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (e.g., EPROM, EEPROM, or flash memory), or any other tangible medium capable of storing information.
[0056] Figure 2 shows a block diagram of a trace chain and exemplary nodes 200, 201 used for distributed hardware tracing executed by system 100. In some implementations, nodes 200, 201 of system 100 may be different nodes within a single multi-core processor. In other implementations, node 200 may be the first node in a first multi-core processor chip, and node 201 may be the second node in a second multi-core processor chip.
[0057] Although two nodes are illustrated in the implementation example of Figure 2, in alternative implementation examples, system 100 may include multiple nodes. For implementation examples with multiple nodes, cross-node data transfer can generate trace data at intermediate hops along an exemplary data path that traverses multiple nodes. For example, an intermediate hop can correspond to a data transfer passing through a separate node in a particular data transfer path. In some examples, trace data associated with ICI traces / hardware events can be generated for one or more intermediate hops that occur during cross-node data transfer passing through one or more nodes.
[0058] In some implementation examples, node 0 and node 1 are tensor nodes used for vectorized calculations associated with a part of the program code for the inference workload. As used in this specification, a tensor is a multi-dimensional geometric object, and exemplary multi-dimensional geometric objects include matrices and data arrays.
[0059] As shown in the implementation example of FIG. 2, node 200 includes a trace chain 203 that interacts with at least a subset of the components of system 100. Similarly, node 201 includes a trace chain 205 that interacts with at least a subset of the components of system 100. In some implementation examples, nodes 200, 201 are exemplary nodes of the same subset of components, while in other implementation examples, nodes 200, 201 are the respective nodes of separate component subsets. Data router / ICI 112 includes a trace chain 207 that generally converges with trace chains 203 and 205 to provide trace data to chip manager 216.
[0060] In the implementation example of FIG. 2, nodes 200, 201 may each include a respective subset of components having at least FPC 104, SPC 106, node fabric 110, and HIB 114. Each component of nodes 200, 201 includes one or more trace muxes configured to group trace points (described below) generated by a particular component of the node. FPC 104 is a trace mux It includes 204, the node fabric 110 includes trace muxes 210a / b, the SPC 106 includes trace muxes 206a / b / c / d, the HIB 214 includes trace mux 214, and the ICI 212 includes trace mux 212. In some implementations, the trace control registers for each trace mux enable individual trace points to be enabled and disabled. In some examples, for one or more trace muxes, their corresponding trace control registers may include individual enable bits and broader trace mux control.
[0061] Generally, the trace control register may be a conventional control status register (CSR) that receives and stores trace instruction data. Broader Regarding broader trace mux control, in some implementations, tracing can be enabled and disabled based on CSR writes executed by the system 100. In some implementations, tracing can be dynamically started and stopped by the system 100 based on the value of a global time counter (GTC), the value of an exemplary trace mark register in the FPC 104 (or core 116), or the value of a step mark in the SPC 106.
[0062] Details and explanations related to a computing system and a method implemented by a computer for dynamically starting and stopping trace activity and for synchronized hardware event collection are described in related U.S. Patent Application No. 15 / 472,932, titled "Synchronized Hardware Event Collection," filed on March 29, 2017, and Attorney Docket No. 16113 - 8129001. The entire disclosure of U.S. Patent Application No. 15 / 472,932 is hereby expressly incorporated by reference in its entirety.
[0063] In some implementation examples, for core 116, FPC 104 can use trace control parameters to define a trace window associated with event activities occurring within core 116. The trace control parameters enable the trace window to be defined by a lower limit and an upper limit for the GTC, as well as a lower limit and an upper limit for the trace mark register.
[0064] In some implementation examples, system 100 may include a function that enables a reduction in the number of generated trace entries, such as a trace event filtering function. For example, FPC 104 and SPC 106 may each include a filtering function that restricts the rate at which each core sets trace bits in an exemplary generated trace descriptor (described below). HIB 114 may include a similar filtering function, such as an exemplary DMA rate limiter that restricts the trace bits associated with the capture of a certain DMA trace event. In addition, HIB 114 may include control (e.g., via an enable bit) for restricting the queuing of source DMA trace entries.
[0065] In some implementation examples, descriptors for DMA operations may have trace bits set by an exemplary compiler of host system 126. When the trace bits are set, the hardware functions / knobs that determine and generate trace data are used to complete an exemplary trace event. In some examples, the last trace bit in the DMA may be a logical OR operation of the trace bits statically inserted by the compiler and the trace bits dynamically determined by a specific hardware component. Thus, in some examples, the trace bits generated by the compiler can provide a mechanism for reducing the overall amount of generated trace data, apart from filtering.
[0066] For example, the compiler of host system 126 may perform one or more remote DMA operations (e.g., For example, it may be determined to set only trace bits for DMA spanning between at least two nodes and clear trace bits for one or more local DMA operations (e.g., DMA within a specific tensor node such as node 200). In this way, the amount of trace data generated can be reduced based on trace activities limited to cross-node (i.e., remote) DMA operations rather than trace activities including both cross-node and local DMA operations.
[0067] In some implementations, at least one trace event initiated by system 100 may be associated with a memory access operation that includes a plurality of intermediate operations occurring across system 100. A descriptor for the memory access operation (e.g., a MAE descriptor) may include trace bits that cause data associated with the plurality of intermediate operations to be stored in one or more memory buffers. Thus, the trace bits can be used to "tag" intermediate memory operations at intermediate hops of the DMA operation as data packets traverse system 100, generating a plurality of trace events.
[0068] In some implementations, ICI 112 may include a set of enable bits and a set of packet filters that provide control functionality for each of the ingress and egress ports of specific components of nodes 200, 201. These enable bits and packet filters enable ICI 112 to enable and disable trace points associated with specific components of nodes 200, 201. In addition to enabling and disabling trace points, ICI 112 may be configured to filter trace data based on event source, event destination, and trace event packet type.
[0069] In some implementations, in addition to using step markers, GTCs, or trace markers, each trace control register for processor cores 104, 106, and HIB 114 may also include an "everyone" trace mode. This "everyone" trace mode may enable tracing across the entire processor chip to be controlled by either trace mux 204 or trace mux 206a. In the everyone trace mode, trace muxes 204 and 206a can send a "window-in" trace control signal that identifies whether that particular trace mux, i.e., either mux 204 or mux 206a, is within the trace window.
[0070] The window-in trace control signal can be sent in parallel or broadcast to all other trace muxes within, for example, one processor chip or across multiple processor chips. When either mux 204 or mux 206a is performing trace activity, all tracing can be enabled by the parallel send to the other trace muxes. In some implementations, each trace mux associated with processor cores 104, 106, and HIB 114 includes a trace window control register that identifies when and / or how the "everyone trace" control signal is generated.
[0071] In some implementations, trace activity in trace muxes 210a / b and trace mux 212 is generally enabled based on whether a trace bit is set in a data word for a DMA operation or control message crossing the ICI / data router 112. The DMA operation or control message may be a fixed-size binary data structure that may have a trace bit within a binary data packet set based on a certain situation or software state.
[0072] For example, if a DMA operation is a trace type DMA instruction to FPC 104 (or SP Starting with C106, if the initiator (processor core 104 or 106) is within the trace window, the trace bit will be set in that particular DMA. In another example, for FPC104, if FPC104 is within the trace window and a trace point that stores trace data is enabled, a control message for writing data to another component within system 100 will set the trace bit.
[0073] In some implementations, zero - length DMA operations provide an example of a broader DMA implementation within system 100. For example, some DMA operations can generate non - DMA activity within system 100. The execution of non - DMA activity can also be traced (e.g., generate trace data) as if the non - DMA activity were a DMA operation (e.g., including non - zero - length operations). For example, a DMA operation that starts at a source location but has no data to be transmitted or transferred (e.g., zero - length) may instead send a control message to the destination location. The control message will indicate that there is no data to be received or worked on at the destination. And the control message itself will be traced by system 100 as if it were a non - zero - length DMA operation.
[0074] In some examples, for SPC106, zero - length DMA operations can generate a control message, and the trace bit associated with that message will be set only if the DMA sets the trace bit, i.e., only if the control message does not have zero - length. Generally, if HIB114 is within the trace window, a DMA operation started from host system 126 will set the trace bit.
[0075] In the implementation example of FIG. 2, trace chain 203 receives trace entry data for a subset of components aligned with node 0, while trace chain 205 receives trace entry data for a subset of components aligned with node 1. Each of the trace chains 203, 205, 207 is a separate data communication path used by their respective nodes 200, 201, and ICI112 to provide trace entry data to the exemplary trace entry data log 218 of chip manager 216. For this reason, the endpoints of the trace chains 203, 205, 207 are the chip manager 216 where trace events can be stored in an exemplary memory unit.
[0076] In some implementation examples, at least one memory unit of the chip manager 216 can be 128 bits wide and have a memory depth of at least 20,000 trace entries. In alternative implementation examples, at least one memory unit may have a larger or smaller bit width and may have a memory depth that can store more or fewer entries.
[0077] In some implementation examples, the chip manager 216 may include at least one processing device that executes instructions for managing the received trace entry data. For example, the chip manager 216 can execute instructions to scan / analyze the timestamp data for each hardware event of the trace data received via the trace chains 203, 205, 207. Based on the analysis, the chip manager 216 can populate the trace entry log 218 to include data that can be used to identify (or generate) the chronological order of the hardware trace events. The hardware trace events can correspond to the movement of data packets that occur at the component and sub-component levels when the processing unit of the system 100 executes an exemplary distributed software program. It can correspond to the movement of data packets that occur at the component and sub-component levels when the processing unit of the system 100 executes an exemplary distributed software program.
[0078] In some implementation examples, the hardware unit of system 100 may generate trace entries (and corresponding timestamps) that non-sequentially (i.e., in any order) populate an exemplary hardware trace buffer. For example, chip manager 216 can cause a plurality of trace entries with generated timestamps to be inserted into entry log 218. Among the plurality of inserted trace entries, each trace entry may not be serialized with respect to each other. In this implementation example, the non-sequential trace entries can be received by an exemplary host buffer of host system 126. When received by the host buffer, host system 126 can execute instructions related to performance analysis / monitoring software to scan / analyze the timestamp data for each trace entry. The executed instructions can be used to sort the trace entries and to construct / generate a timeline of hardware trace events.
[0079] In some implementation examples, during a tracing session, trace entries can be removed from entry log 218 via a host DMA operation. In some examples, host system 126 may not remove DMA entries from trace entry log 218 as quickly as they are added to the log. In other implementation examples, entry log 218 may include a predefined memory depth. When the memory depth limit of entry log 218 is reached, additional trace entries may be lost. To control which trace entries are lost, entry log 218 can operate in a first-in-first-out (FIFO) mode or, alternatively, in an overwrite recording mode.
[0080] In some implementation examples, the overwrite recording mode may be used by system 100 to support performance analysis associated with postmortem debugging. For example, program code may be executed for a period of time with tracing activities enabled and the overwrite recording mode enabled. In response to a postmortem software event (such as program corruption) within system 100, monitoring software executed by host system 126 may analyze the data content of an exemplary hardware trace buffer to understand the hardware events that occurred prior to the program corruption. As used in this specification, postmortem debugging relates to the analysis or debugging of program code after the code has been corrupted or after it has generally become unable to execute / operate as intended.
[0081] In FIFO mode, if entry log 218 is full and host system 126 removes stored log entries within a certain time frame, new trace entries may not be stored in the memory unit of chip manager 216 to conserve memory resources. On the other hand, in overwrite recording mode, if entry log 218 is full and host system 126 removes stored log entries within a certain time frame, new trace entries can be overwritten on the oldest trace entry stored in entry log 218 to conserve memory resources. In some implementation examples, trace entries are moved to the memory of host system 126 in response to a DMA operation using the processing function of HIB114.
[0082] As used in this specification, a trace point is the origin of a trace entry and the data associated with that trace entry received by chip manager 216 and stored in trace entry log 218. In some implementation examples, a multi-core multi-node processor microchip includes three trace chains within the chip It may also be the case that the first trace chain receives a trace entry from chip node 0, the second trace chain receives a trace entry from chip node 1, and the third trace chain receives a trace entry from the ICI router of the chip.
[0083] Each trace point has a unique trace identification number within its trace chain that it inserts into the header of the trace entry. In some implementations, each trace entry identifies the trace chain on which it occurred in a header indicated by one or more byte / bit data words. For example, each trace entry may include a data structure having a defined field format (e.g., header, payload, etc.) that conveys information about a particular trace event. Each field in the trace entry corresponds to useful data applicable to the trace point that generated the trace entry.
[0084] As described above, each trace entry may be written or stored in the memory unit of the chip manager 216 associated with the trace entry log 218. In some implementations, the trace points may be individually enabled or disabled, and multiple trace points may generate trace entries having different trace point identifiers but of the same type.
[0085] In some implementations, each trace entry type may include a trace name, a trace description, and a header that identifies an encoding for specific fields and / or a set of fields within the trace entry. These names, descriptions, and headers collectively provide a description of what the trace entry represents. From the perspective of the chip manager 216, this description can also identify which specific trace chains 203, 205, 207 within a specific processor chip a particular trace entry belongs to. Thus, the fields within a trace entry represent pieces of data (e.g., in bytes / bits) related to the description and may be trace entry identifiers used to determine which trace point generated a particular trace entry.
[0086] In some implementations, the trace entry data associated with one or more of the stored hardware events can partially correspond to data communications that occur a) between at least node 0 and node 1, b) between components within at least node 0, and c) between components within at least node 1. For example, the stored hardware event can partially correspond to data communications that occur in at least one of 1) between the FPC 104 of node 0 and the FPC 104 of node 1, between the FPC 104 of node 0 and the SPC 106 of node 0, and 2) between the SPC 106 of node 0 and the SPC 106 of node 1.
[0087] FIG. 3 shows a block diagram of an exemplary trace multiplexing design architecture 300 and an exemplary data structure 320. The trace multiplexing design 300 generally includes a trace bus input 302, a bus arbiter 304, and a local trace point arbiter 306, a bus FIFO 308, at least one local trace event queue 310, a shared trace event FIFO 312, and a trace bus output 314.
[0088] The multiplexing design 300 corresponds to an exemplary trace mux disposed within a component of the system 100. The multiplexing design 300 may include the following functionality. The bus in 302 may be associated with local trace point data that is temporarily stored in the bus FIFO 308 until the trace data is placed on an exemplary trace chain by time arbitration logic (e.g., arbiter 304). One or more trace points for the component can insert trace event data into at least one local trace event queue 310. The arbiter 306 provides a first level of arbitration and enables selection of events from the local trace events stored in the queue 310. The selected events are placed in the shared trace event FIFO 312 which also functions as a storage queue.
[0089] The arbiter 304 receives local trace events from the FIFO queue 312 and provides a second level of arbitration that merges the local trace events onto specific trace chains 203, 205, 207 via the trace bus out 314. In some implementations, trace entries may be pushed into the local queue 310 faster than they can be merged into the shared FIFO 312. Or, alternatively, trace entries may be pushed into the shared FIFO 312 faster than they can be merged onto the trace bus 314. If these scenarios occur, each of the queues 310 and 312 will become full with trace data.
[0090] In some implementations, when any of the queues 310 or 312 becomes full of trace data, the system 100 may be configured such that the latest trace entry is dropped and not stored or merged into a particular queue. In other implementations, rather than dropping trace entries when a queue (e.g., queue 310, 312) becomes full, the system 100 may be configured to stall an exemplary processing pipeline until the once-full queue has available queue space to receive an entry.
[0091] For example, a processing pipeline using queues 310, 312 may be stalled until a sufficient or threshold number of trace entries are merged on the trace bus 314. The sufficient or threshold number can correspond to a particular number of merged trace entries that results in available queue space for one or more trace entries to be received by queues 310, 312. Implementations where the processing pipeline is stalled until downstream queue space becomes available can provide higher-fidelity trace data based on the fact that trace entries are not dropped but preserved.
[0092] In some implementations, the local trace queue is as wide as required by a trace entry so that each trace entry occupies only one location in the local queue 310. However, the shared trace FIFO queue 312 can use a unique trace entry line encoding such that some trace entries can occupy two locations in the shared queue 312. In some implementations, if any data of a trace packet is dropped, the entire packet is dropped so that partial packets do not appear in the trace entry log 218.
[0093] Generally, a trace is a timeline of activities or hardware events associated with specific components of system 100. Unlike performance counters, which are aggregate data (described below), a trace includes detailed event data that provides insight into the hardware activities that occur within a specified trace window. The hardware systems described enable extensive support for distributed hardware tracing, including generation of trace entries, temporary storage of trace entries in a hardware management buffer, static and dynamic enabling of one or more trace types, and streaming of trace entry data to host system 126.
[0094] In some implementations, a trace can be generated for hardware events that are executed by components of system 100, such as generation of a DMA operation, execution of a DMA operation, issue / execution of an instruction, or update of a synchronization flag. In some examples, trace activity can be used to track DMA through the system or to track instructions executed on a particular processor core.
[0095] System 100 can be configured to generate at least one data structure 320 that identifies one or more hardware events 322, 324 from a timeline of hardware events. In some implementations, data structure 320 arranges one or more hardware events 322, 324 in chronological order of events associated with at least FPC 104 and SPC 106. In some examples, system 100 can store data structure 320 in a memory bank of a host control device of host system 126. Data structure 320 can be used to evaluate the performance of program code executed by at least processor cores 104 and 106.
[0096] As indicated by hardware event 324, in some implementations, a particular trace identification (ID) number (e.g., trace ID ‘003) can be associated with multiple hardware events that occur across distributed processor units. The multiple hardware events can correspond to a particular memory access operation (e.g., DMA), and the particular trace ID number is used to correlate one or more hardware events.
[0097] For example, as indicated by event 324, a single trace ID for a DMA operation can include multiple timestamps corresponding to multiple different points in the DMA. In some examples, trace ID ‘003 can have “issued,” “executed,” and “completed” events that are identified as being somewhat temporally separated from each other. Thus, in this regard, the trace ID can further be used to determine a latency attribute of a memory access operation based on correlation and with reference to the timestamps.
[0098] In some implementations, generating data structure 320 can include, for example, the system 100 comparing the event timestamp of each event in a first subset of hardware events with the event timestamp of each event in a second subset of hardware events. Generating data structure 320 can further include the system 100 providing a correlated set of hardware events for presentation in the data structure, based in part on the comparison of the first subset of events and the second subset of events.
[0099] As shown in FIG. 3, the data structure 320 can identify at least one parameter indicating the latency attributes of specific hardware events 322, 324. The latency attributes can at least indicate the duration of a specific hardware event. In some implementations, the data structure 320 is generated by software instructions executed by a control device of the host system 126. In some examples, the structure 320 can be generated in response to the control device storing trace entry data in a memory disk / unit of the host system 126.
[0100] FIG. 4 is a block diagram 400 showing exemplary trace activities for direct memory access (DMA) trace events executed by the system 100. For DMA tracing, data about an exemplary DMA operation occurring from a first processor node to a second processor node can proceed through the ICI 112 and can generate intermediate ICI / router hops along the data path. When the DMA operation crosses the ICI 112, the DMA operation will generate trace entries at each node within the processor chip and at each hop. To reconstruct the time progression of the DMA operation along the nodes and hops, information is captured by each of these generated trace entries. The exemplary DMA operation can be associated with the process steps shown in the implementation example of FIG. 4. For this operation, local DMA transfers data from the virtual memory 402 (vmem402) associated with at least one of the processor cores 104, 106 to the HBM 108. The numbering shown in the block diagram 400 corresponds to the steps in the table 404 and generally represents activities in the node fabric 110 or activities initiated by the node fabric 110.
[0101]
[0102] The steps of Table 404 generally describe the associated trace points. The exemplary operation will generate six trace entries for this DMA. Step 1 includes the first DMA request from the processor core to the node fabric 110, which generates a trace point in the node fabric. Step 2 includes a read command that requests the node fabric 110 to transfer data to the processor core, which generates another trace point in the node fabric 110. When vmem402 completes the read of the node fabric 110, the exemplary operation has no trace entry for Step 3.
[0103] Step 4 includes the node fabric 110 performing a read resource update to cause a synchronization flag update in the processor core, which generates a trace point in the processor core. Step 5 includes a write command that the node fabric 110 notifies the memory mux108 that the next data is to be written to the HBM. The notification via the write command generates a trace point in the node fabric 110, while in Step 6, the completion of the write to the HBM also generates a trace point in the node fabric 110. In Step 7, the node fabric 110 performs a write resource update to cause a synchronization flag update in the processor core, which generates a trace point in the processor core (e.g., in FPC104). In addition to the write resource update, the node fabric 110 can perform a receive confirmation update (ack update) where the data completion for the DMA operation is signaled back to the processor core. The ack update can generate a trace entry similar to the trace entry generated by the write resource update.
[0104] In another exemplary DMA operation, when a DMA command is issued at the node fabric 110 of the source node, a first trace entry is generated. Additional trace entries may be generated at the node fabric 110 to capture the time used to read data for the DMA and write that data to the transmit queue. In some implementations, the node fabric 110 can packetize the DMA data into smaller chunks of data. For the data packetized into smaller chunks, read and write trace entries may be generated for the first and last data chunks. Optionally, in addition to the first and last data chunks, all data chunks may be set to generate trace entries.
[0105] For remote / non-local DMA operations that may require ICI hops, the first and last data chunks can generate additional trace entries at the ingress and egress points at each intermediate hop along the ICI / router 112. When the DMA data arrives at the destination node, trace entries similar to the previous node fabric 110 entries are generated at the destination node (e.g., read / write of the first and last data chunks). In some implementations, the last step of the DMA operation may include causing the executed command associated with the DMA to trigger an update of the synchronization flag at the destination node. When the synchronization flag is updated, a trace entry indicating the completion of the DMA operation may be generated.
[0106] In some implementation examples, to enable trace points to be executed, DMA tracing is initiated by FPC104, SPC106, or HIB114 when each component is in trace mode. Components of system 100 can enter trace mode based on global control in FPC104 or SPC106 via a trigger mechanism. In response to the occurrence of a specific action or state associated with the execution of program code by a component of system 100, a trace point is triggered. For example, a part of the program code may include an embedded trigger function that is detectable by at least one hardware component of system 100.
[0107] Components of system 100 can be configured to detect a trigger function associated with a part of the program code executed by at least one of FPC104 or SPC106. In some examples, the trigger function can correspond to at least one of 1) a specific sequence step in a part or module of the executed program code, or 2) a specific time parameter indicated by the GTC used by the distributed processor units of system 100.
[0108] In response to the detection of the trigger function, a specific component of system 100 can initiate, trigger, or execute at least one trace point (e.g., a trace event) that causes trace entry data associated with one or more hardware events to be stored in at least one memory buffer of a hardware component. As described above, the stored trace data can then be provided to chip manager 216 via at least one of trace chains 203, 205, 207.
[0109] FIG. 5 is a process flow diagram of an exemplary process 500 for distributed hardware tracing using the component functions of system 100 and one or more nodes 200, 201 of system 100. Thus, process 500 can be implemented using one or more of the computing resources of system 100 described above, including nodes 200, 201.
[0110] Process 500 begins at block 502 and includes the step of the computing system 100 monitoring the execution of program code executed by one or more processor components (including at least FPC 104 and SPC 106). In some implementations, the execution of the program code that generates the trace activity can be monitored at least in part by multiple host systems, or subsystems of a single host system. Thus, in these implementations, system 100 can perform multiple processes 500 related to the analysis of trace activity for hardware events occurring across distributed processing units.
[0111] In some implementations, the first processor component is configured to execute at least a first portion of the program code being monitored. At block 504, process 500 includes the step of the computing system 100 monitoring the execution of program code executed by a second processor component. In some implementations, the second processor component is configured to execute at least a second portion of the program code being monitored.
[0112] Each component of computing system 100 includes at least one memory It may include a buffer. Block 506 of process 500 includes the step of the system 100 storing data identifying one or more hardware events in at least one memory buffer of a particular component. In some implementations, the hardware events occur across a distributed processor unit including at least a first processor component and a second processor component. Each of the stored data identifying the hardware events may include metadata characterizing the hardware event and a hardware event timestamp. In some implementations, the set of hardware events corresponds to timeline events.
[0113] For example, the system 100 can store data identifying one or more hardware events that partially correspond to the movement of data packets between a source hardware component within the system 100 and a destination hardware component within the system 100. In some implementations, the stored metadata characterizing the hardware event can correspond to at least one of 1) a source memory address, 2) a destination memory address, 3) a unique trace identification number associated with a trace entry that stores the hardware event, or 4) a size parameter associated with a direct memory access (DMA) trace entry.
[0114] In some implementations, the step of storing data identifying the set of hardware events includes, for example, storing event data in a memory buffer of the FPC 104 and / or the SPC 106 corresponding to at least one local trace event queue 310. The stored event data may indicate a subset of the hardware event data that can be used to generate a larger timeline of the hardware events. In some implementations, the storing of the event data occurs in response to at least one of the FPC 104 or the SPC 106 executing a hardware trace instruction associated with a portion of the program code executed by a component of the system 100.
[0115] In block 508 of process 500, system 100 generates a data structure, such as structure 320, that identifies one or more hardware events from a set of hardware events. The data structure arranges the one or more hardware events in a chronological order of events associated with at least a first processor component and a second processor component. In some implementations, the data structure identifies a hardware event timestamp for a particular trace event, a source address associated with that trace event, or a memory address associated with that trace event.
[0116] In block 510 of process 500, system 100 stores the generated data structure in a memory bank of a host device associated with host system 126. In some implementations, the stored data structure can be used by host system 126 to analyze the performance of program code executed by at least the first processor component or the second processor component. Similarly, the stored data structure can be used by host system 126 to analyze the performance of at least one component of system 100.
[0117] For example, a user or host system 126 can analyze the data structure to detect or determine whether there are performance issues associated with the execution of a particular software module within program code. Exemplary problems can include the software module not completing execution within an allotted execution time window.
[0118] Further, the user or host device 126 can detect or determine whether a particular component of the system 100 is operating above or below a threshold performance level. Exemplary problems related to component performance can include a particular hardware component performing an event but generating result data that is outside the acceptable parameter range for the result data. In some implementations, this result data may not match the result data generated by other relevant components of the system 100 performing substantially similar operations.
[0119] For example, during the execution of program code, a first component of the system 100 may be required to complete an operation and to generate a result. Similarly, a second component of the system 100 may be required to complete a substantially similar operation and to generate a substantially similar result. Analysis of the generated data structure may indicate that the second component generated a result that is significantly different from the result generated by the first component. Similarly, the data structure may indicate result parameter values for the second component that are significantly outside the range of acceptable result parameters. These results may potentially indicate a performance problem with the second component of the system 100.
[0120] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more combinations of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier for execution by, or to control the operation of, a data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a receiver apparatus suitable for execution by a data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or one or more combinations of them.
[0121] The processes and logical flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, and the apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array), an ASIC (application specific integrated circuit ), or a GPGPU (general purpose graphics processing unit).
[0122] Computers suitable for the execution of a computer program include, by way of example, general purpose or special purpose microprocessors or both, or any other kind of central processing unit, and may be based thereon. Generally, the central processing unit will receive instructions and data from read-only memory or random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions, and one or more memory devices for storing instructions and data. Generally, a computer also includes, or is operatively coupled to receive data from, or transfer data to, one or more mass storage devices, such as magnetic disks, magneto-optical disks, or optical disks, for storing data, although a computer need not have such devices. However, a computer need not have such devices.
[0123] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks. The processor and the memory may be supplemented or incorporated by special purpose logic circuitry.
[0124] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of the invention or the claims, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Features described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable partial combination. Furthermore, features may be described as acting in a certain combination and as such may initially be claimed, but in some cases, one or more features from the claimed combination may be deleted from that combination, and the claimed combination may be directed to a partial combination or a variation of a partial combination.
[0125] Similarly, operations are shown in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in the particular order or in a sequential order shown for achieving desirable results, or that all of the illustrated operations be performed. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0126] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes shown in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve desirable results. In some implementations, multitasking and parallel processing may be advantageous.
Claims
1. A method performed by a hardware tracing system for capturing data that describes hardware events, the method comprising: detecting, in a multi-core processor including a plurality of cores, a trigger that causes capture of the data that describes the hardware events, the trigger being i) satisfied as a result of executing program code using respective processor components within each of the plurality of cores, ii) configured to cause storage of event data regarding an operation performed by the multi-core processor, the method further comprising storing, in the multi-core processor, event data that describes trace activity in response to detecting the trigger, the event data being captured and stored during execution of the program code in the multi-core processor, the event data being i) specific to operations involving transfer of data between respective processor components within at least one of the plurality of cores, and ii) transfer of data between respective cores of the multi-core processor, the method further comprising providing the event data stored in response to detecting the trigger to host software, the event data provided to the host software representing a timeline of operations and enabling analysis of the performance of the program code using the host software, generating, in the multi-core processor, a plurality of respective trace points during the transfer of the data between the respective processor components, wherein a subset of the trace points is generated in one or more of the respective cores of the multi-core processor.
2. The method of claim 1, wherein the trigger is encoded as an embedded operand in the program code using a parameter value that defines one step in the computing operation of the program code.
3. The step of storing event data in response to detecting the trigger includes storing event data in respective performance counters of the multi-core processor. The method according to claim 1 or 2, wherein the event data stored in each of the performance counters includes an incremental parameter count that describes trace activities for a specific operation. **Claim 4**: A method performed by a hardware tracing system for capturing data describing hardware events, the method comprising: detecting, in a multi-core processor including a plurality of cores, a trigger that causes capture of the data describing the hardware events, the trigger being i) embedded as an operand in program code executed using each processor component within each of the plurality of cores, the trigger being satisfied as a result of execution of the operand, ii) configured to cause storage of event data regarding an operation performed in the multi-core processor, the method further comprising: storing, in response to detecting the trigger, event data describing trace activities in the multi-core processor, the event data being captured and stored during execution of the program code in the multi-core processor, the event data being i) specific to operations involving transfer of data between respective processor components within at least one of the plurality of cores, and ii) transfer of data between respective cores of the multi-core processor, the method further comprising: providing the event data stored in response to detecting the trigger to host software, the event data provided to the host software representing a timeline of operations and enabling analysis of the performance of the program code using the host software, generating, in the multi-core processor, a plurality of respective trace points during the transfer of the data between the respective processor components, wherein a subset of the trace points is generated in one or more of the respective cores of the multi-core processor. **Claim 5** The method according to claim 1 or 4, wherein a trace point in the first core of the multi-core processor is used to store event data for a processor component of the first core based on a trace bit that controls the trace activity of the processor component.
6. Detecting that an instruction for one operation includes the trace bit for controlling the trace activity of the processor component of the first core; Determining whether the trace bit is set to a value that enables the trace activity of the processor component of the first core; Storing event data in a performance counter of the first core for a trace point generated by the multi-core processor in response to a determination that the trace bit is set to a value that enables trace activity, the method according to claim 5.
7. The step of storing the event data occurs based on an operation in which data being transferred crosses a plurality of intermediate positions along a data path, the data path Communicatively coupling the respective cores of the multi-core processor, or Communicatively coupling the respective multi-core processors of the hardware tracing system, the method according to claim 1 or 4.
8. Each of the plurality of intermediate positions corresponds to one or more trace points managed by a trace mux of a respective core of the multi-core processor, the method according to claim 7.
9. The operation involving the transfer of data includes a memory access operation, The memory access operation Accessing data stored in a first core of the multi-core processor, and Transferring the data using an inter-chip interconnect of the hardware tracing system for storage in a second different core of the multi-core processor or for storage in another multi-core processor, the method according to any one of claims 1 to 8.
10. The step of providing the event data to the host software includes providing the event data to host software executed on a host controller of the hardware tracing system. The method according to any one of claims 1 to 9, wherein the event data is provided for analysis by the host controller in order to debug hardware problems occurring during the execution of the program code.
11. A hardware tracing system for capturing data describing hardware events, the hardware tracing system comprising: one or more processing devices; and one or more storage devices for storing instructions executable by the one or more processing devices to cause the execution of an operation, the operation comprising: detecting a trigger in a multi-core processor including a plurality of cores that causes the capture of the data describing the hardware event, the trigger being: i) satisfied as a result of executing program code using respective processor components within each of the plurality of cores; ii) configured to cause the storage of event data regarding the operation performed on the multi-core processor, the operation further comprising: in response to detecting the trigger, storing, in the multi-core processor, event data describing trace activity, the event data being: captured and stored during the execution of the program code on the multi-core processor; i) specific to operations involving the transfer of data between respective processor components within at least one of the plurality of cores; and ii) the transfer of data between respective cores of the multi-core processor, the operation further comprising: in response to detecting the trigger, providing the stored event data to host software, the event data provided to the host software representing a timeline of the operation and enabling the analysis of the performance of the program code using the host software; the operation includes generating a plurality of respective trace points in the multi-core processor during the transfer of the data between the respective processor components; A hardware tracing system, wherein a subset of the trace points is generated in one or more of the respective cores of the multi-core processor.
12. The hardware tracing system according to claim 11, wherein the trigger is encoded as an embedded operand in the program code using a parameter value that defines one step in the computing operation of the program code.
13. Storing event data in response to detecting the trigger includes storing event data in respective performance counters of the multi-core processor, The hardware tracing system according to claim 11 or 12, wherein the event data stored in each of the respective performance counters includes an incremental parameter count that describes trace activity for a specific operation.
14. A hardware tracing system for capturing data that describes a hardware event, the hardware tracing system comprising: One or more processing devices; One or more storage devices for storing instructions executable by the one or more processing devices to cause execution of an operation, the operation including: Detecting a trigger in a multi-core processor including a plurality of cores that causes capture of the data that describes the hardware event, the trigger being: i) Embedded as an operand in program code executed using respective processor components within each of the plurality of cores, the trigger being satisfied as a result of execution of the operand; ii) Configured to cause storage of event data regarding the operation performed on the multi-core processor, the operation further including: Storing, in the multi-core processor, event data that describes trace activity in response to detecting the trigger, the event data being: Captured and stored during execution of the program code on the multi-core processor, i) Transfer of data between respective processor components within at least one of the plurality of cores; and ii) Transfer of data between respective cores of the multi-core processor, and is specific to an operation involving, the operation further including: including providing the event data stored in response to detecting the trigger to host software, the event data provided to the host software representing a timeline of operations and enabling analysis of the performance of the program code using the host software, the operations including generating a plurality of respective trace points in the multi-core processor during the transfer of data between the respective processor components, a hardware tracing system in which a subset of the trace points is generated in one or more of the respective cores of the multi-core processor. **Claim 15** The hardware tracing system according to claim 11 or 14, wherein a trace point in a first core of the multi-core processor is used to store event data for the processor component of the first core based on a trace bit that controls the trace activity of the processor component. **Claim 16** The operations are detecting that an instruction for one operation includes the trace bit for controlling the trace activity of the processor component of the first core, determining whether the trace bit is set to a value that enables trace activity of the processor component of the first core, and in response to determining that the trace bit is set to a value that enables trace activity, storing event data in a performance counter of the first core for a trace point generated by the multi-core processor. The hardware tracing system according to claim 15. **Claim 17** The storing of the event data occurs based on an operation in which data being transferred traverses a plurality of intermediate positions along a data path, the data path communicatively coupling the respective cores of the multi-core processor or communicatively coupling the respective multi-core processors of the hardware tracing system. The hardware tracing system according to claim 16. **Claim 18** The operation involving the transfer of data is accessing data stored in a first core of the multi-core processor, transferring the data using an inter-chip interconnect of the hardware tracing system for storage in a second different core of the multi-core processor or for storage in another multi-core processor, the memory access operation including the hardware tracing system according to any of claims 11 to 17.
19. providing the event data to the host software includes providing the event data to host software executed on a host controller of the hardware tracing system, the event data is provided for analysis by the host controller to debug hardware problems occurring during execution of the program code, the hardware tracing system according to any of claims 11 to 18.
20. A program having instructions that, when executed by a computer system, cause the computer system to perform the method according to any of claims 1 to 10.
Citation Information
Patent Citations
Tracking of program
JP1985010354A
Internal interruption control system for microprocessor
JP1990083749A
Tracing system for information processor
JP1992148439A
Data processor, program translation method and debug tool
JP1995200352A
Debugging on multicore architectures
JP2008513853A